Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

fideslib_py — Python bindings for FIDESlib (CKKS on GPU)

Minimal pybind11 wrapper around FIDESlib v2.1.2, the CKKS GPU library interoperable with OpenFHE. Exposes the subset of the API needed for encrypted inference pipelines (HerMiniRocket/PolyMiniRocket): context setup, key generation, encoding, encrypt/decrypt, leveled arithmetic, rotations, EvalChebyshevSeries, AccumulateSum and bootstrapping.

The compiled module statically embeds FIDESlib and a patched OpenFHE 1.5.1. CMake clones and compiles its own pinned copy of FIDESlib (and FIDESlib's own vendored OpenFHE) entirely inside build/ -- no git submodule, no system install, no path outside this repo. The pin lives in CMakeLists.txt (FIDESLIB_REPOSITORY/FIDESLIB_GIT_TAG cache variables). Its only runtime dependencies are the CUDA runtime (RPATH'd to /usr/local/cuda/lib64) and an NVIDIA GPU.

Build

./build.sh # clones + builds FIDESlib (and its OpenFHE) + the wrapper
./build.sh /usr/bin/python3 # against a specific interpreter

First run compiles the vendored OpenFHE from scratch (~10-30 min); later runs are fast since it's only rebuilt when missing. To pin a different FIDESlib commit or fork, pass -DFIDESLIB_REPOSITORY=... -DFIDESLIB_GIT_TAG=... to the cmake -B build step in build.sh, or edit the defaults in CMakeLists.txt.

Prereqs: CUDA toolkit ≥ 12.4, gcc ≥ 11, CMake ≥ 3.25, network access (FIDESlib/OpenFHE/pybind11 fetches), and the Python dev headers for the chosen interpreter.

The module is bound to the Python minor version it was built against (e.g. _core.cpython-312-x86_64-linux-gnu.so ⇒ Python 3.12). Rebuild to switch.

Use

export PYTHONPATH=/path/to/PyFIDESlib # or sys.path.insert / pip install -e
python examples/00_onboarding.py
importfideslib_pyasfheparams=fhe.CCParams()
params.SetSecurityLevel(fhe.HEStd_128_classic)
params.SetMultiplicativeDepth(32)
params.SetScalingModSize(50)
params.SetScalingTechnique(fhe.FLEXIBLEAUTO)
params.SetKeySwitchTechnique(fhe.HYBRID)
params.SetDevices([0]) # GPU id(s)cc=fhe.GenCryptoContext(params)
forfin (fhe.PKE, fhe.KEYSWITCH, fhe.LEVELEDSHE, fhe.ADVANCEDSHE, fhe.FHE):
cc.Enable(f)
keys=cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
cc.EvalRotateKeyGen(keys.secretKey, [1, 2, 3, 4])
cc.LoadContext(keys.publicKey) # pushes keys to GPU; AFTER keygenct=cc.Encrypt(keys.publicKey, cc.MakeCKKSPackedPlaintext([1.0, 2.0, 3.0]))
ct=cc.EvalChebyshevSeries(ct, coeffs, -1.0, 1.0) # GPUpt=cc.Decrypt(keys.secretKey, ct)
pt.SetLength(3)
print(pt.GetRealPackedValue())

Offloading ciphertexts / reclaiming VRAM

When you hold more ciphertexts than fit in VRAM, or want to free GPU memory before other work, evict GPU-resident ciphertexts to host RAM and bring them back on demand:

ct.Offload() # limbs -> host RAM, VRAM freed into FIDESlib's poolct.IsOffloaded() # -> Truect.Reload() # limbs -> GPU (also happens automatically on first use)cc.TrimGPUMemoryPool()# return the freed VRAM to the OS

Offload()/Reload() are a bit-exact round trip (no decrypt/rescale/NTT). Offload() on its own only pools the memory for cheap reuse by later FIDESlib ops — it does not shrink the process's VRAM footprint (nvidia-smi shows no drop). Call cc.TrimGPUMemoryPool() once, after offloading, to actually hand that memory back to the system; skip it if you only intend to reuse the memory for more FIDESlib work. See examples/03_offload.py.

FIDESlib also has a higher-level cache of destroyed ciphertext polynomials. It can help regular create/destroy workloads, but its upstream-unbounded behavior may strand many GiB that plaintext and other allocation paths cannot reuse. Set its maximum retained polynomial count before starting Python; 0 disables only this upper cache while the lower limb allocator continues to reuse GPU buffers:

export FIDESLIB_AUX_POLY_CACHE_LIMIT=0

If the variable is absent or invalid, the original unbounded behavior is retained.

Examples (in suggested order)

ScriptWhat it showsNeeds
examples/00_onboarding.pycontext → keys → encrypt → add/mult/rotate/sum → decrypt<1 GB VRAM, seconds
examples/01_chebyshev.pydeg-31 polynomial + X4 cleaning vs CPU Clenshaw reference~1 GB VRAM, seconds
examples/02_step_herminirocket.pyfull Step() (Lee α=8 + 2×X4) at logN=17, secure params~2–4 GB VRAM, ~1–2 min
examples/03_offload.pyoffload/reload ciphertexts to host RAM, reclaim VRAM with TrimGPUMemoryPool<1 GB VRAM, seconds

Coming from openfhe-python

openfhe-pythonfideslib_py
cc.EvalSum(ct, n)cc.AccumulateSum(ct, n, stride=1)
cc.EvalSumKeyGen(sk)cc.EvalRotateKeyGen(sk, fhe.accumulate_rotation_indices(n, stride))
cc.LoadContext(keys.publicKey) — required once, after all keygen
cc.EvalChebyshevSeries(ct, coeffs, a, b)same
GetSchemeSwitchingData, FHEW comparisons, BFV/BGVnot available (CKKS only)

All Eval* calls release the GIL, so a Python timing/monitoring thread stays responsive. A single dispatch thread is enough — the GPU serializes the work.

Caveats

  • Never import openfhe and import fideslib_py in the same process — they embed different OpenFHE versions (1.5.0 vs patched 1.5.1). Run CPU/GPU comparisons as separate processes.
  • Rotation keys must exist for every index used by EvalRotate / AccumulateSum (helper: fhe.accumulate_rotation_indices).
  • GPU out-of-memory aborts the process (FIDESlib behavior) — check nvidia-smi before logN=17 runs on a shared GPU.
  • numpy arrays are accepted wherever a list of floats is (converted on the way in); returned values are Python lists.

About

Python wrapper for the FIDESlib HE library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages