Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance by trvachov · Pull Request #2 · NVIDIA-BioNeMo/ReaSyn · GitHub
Skip to content

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance - #2

Open
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested
Open

perf(fpindex): GPU fingerprint queries + RapidFuzz edit-distance#2
trvachov wants to merge 1 commit into
reasyn_v2from
trvachov/fpindex-gpu-rapidfuzz-tested

Conversation

@trvachov

@trvachovtrvachov commented Jul 18, 2026

Copy link
Copy Markdown

Summary

Accelerates the building-block fingerprint lookup in get_reactants — the single largest cost in ReaSyn inference (~60% of wall time in profiling).

Two changes to FingerprintIndex.query_cuda:

  1. Run the fingerprint cdist on the model's GPU. The query fingerprint arrives as a CPU tensor, so despite the query_cuda name the distance search was executing on the CPU over the 211k×2048 matrix (~53 ms/query). Routed to the model's assigned cuda:<id>~53 → 1.4 ms.
  2. RapidFuzz for the invalid-SMILES fallback. Invalid generated SMILES fall back to Levenshtein matching over all 211k building-block strings, previously a Python editdistance loop (~335–460 ms/call — the true dominant cost). Replaced with rapidfuzz.process.cdist (batched C++/SIMD) → ~9–15×, identical distances.

Correctness / safety

  • Results preserved exactly. GPU cdist distances are bit-identical to CPU; to also preserve the selected indices at tied distances, the distance vector is copied back to CPU for topk (GPU/CPU topk break ties differently). RapidFuzz distances are identical to editdistance.
  • Multi-GPU safe. Query is routed to the model's explicit cuda:<gpu_id> (not the process-default GPU 0); fp_cuda cache key canonicalized.
  • RapidFuzz pinned in both pyproject.toml and env.yml; workers=1.

Tests

tests/test_fpindex.py — GPU/CPU parity (incl. the real 211k index), device routing, cache behavior, fallback paths, RapidFuzz↔editdistance identity. 9 passed on A100.

Toggles

REASYN_GPU_QUERY (default on), REASYN_RAPIDFUZZ (default on); set =0 to restore legacy behavior.

Independently reviewed and fixed (multi-GPU device routing, topk tie identity, manifest pinning) before opening.

Testing environment ⚠️

All test/benchmark results above were produced on the review container: Python 3.12 / PyTorch 2.11 / CUDA 13.2not the repo-pinned stack (Python 3.10 / PyTorch 2.7 / CUDA 11.8 per env.yml / pyproject.toml). Kernel selection, SDPA backends, and tie-breaking are version-sensitive, so these results should be re-confirmed on the pinned stack / CI before merge.

Run fingerprint cdist on the model-assigned GPU and accelerate invalid-SMILES Levenshtein lookup with RapidFuzz.
Thread explicit multi-GPU device routing, canonicalize CUDA cache keys, restore legacy CPU topk tie identity, pin RapidFuzz in both manifests, and use workers=1. Add tests/test_fpindex.py for GPU/CPU parity, device routing, cache behavior, fallback paths, and the real index.
Reviewed-and-tested: adversarially verified on the 211k fingerprint index with an A100, on CPU-only paths, and against both model checkpoints.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@trvachov