Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Neofuzz


Blazing fast, lightweight and customizable fuzzy and semantic text search in Python.

Introduction (Documentation)

Neofuzz is a fuzzy search library based on vectorization and approximate nearest neighbour search techniques.

New in version 0.3.0

Now you can reorder your search results using Levenshtein distance! Sometimes n-gram processes or vectorized processes don't quite order the results correctly. In these cases you can retrieve a higher number of examples from the indexed corpus, then refine those results with Levenshtein distance.

fromneofuzzimportchar_ngram_processprocess=char_ngram_process()
process.index(corpus)
process.extract("your query", limit=30, refine_levenshtein=True)

Why is Neofuzz fast?

Most fuzzy search libraries rely on optimizing the hell out of the same couple of fuzzy search algorithms (Hamming distance, Levenshtein distance). Sometimes unfortunately due to the complexity of these algorithms, no amount of optimization will get you the speed, that you want.

Neofuzz makes the realization, that you can’t go above a certain speed limit by relying on traditional algorithms, and uses text vectorization and approximate nearest neighbour search in the vector space to speed up this process.

When it comes to the dilemma of speed versus accuracy, Neofuzz goes full-on speed.

When should I choose Neofuzz?

  • You need to do repeated searches in the same corpus.
  • Levenshtein and Hamming distance is simply not fast enough.
  • You are willing to sacrifice the quality of the results for speed.
  • You don’t mind that the up-front computation to index a corpus might take time.
  • You have very long strings, where other methods would be impractical.
  • You want to rely on semantic content.
  • You need a drop-in replacement for TheFuzz.

When should I NOT choose Neofuzz?

  • The corpus changes all the time, or you only want to do one search in a corpus. (It might still give speed-up in that case though.)
  • You value the quality of the results over speed.
  • You don’t mind slower searches in favor of no indexing.
  • You have a small corpus with short strings.

You can install Neofuzz from PyPI:

pip install neofuzz

If you want a plug-and play experience you can create a generally good quick and dirty process with the char_ngram_process() process.

fromneofuzzimportchar_ngram_process# We create a process that takes character 1 to 5-grams as features for# vectorization and uses a tf-idf weighting scheme.# We will use cosine distance for the nearest neighbour search.process=char_ngram_process(ngram_range=(1,5), metric="angular", tf_idf=True)
# We index the options that we are going to search inprocess.index(options)
# Then we can extract the ten most similar items the same way as in# thefuzzprocess.extract("fuzz", limit=10)
---------------------------------
[('fuzzer', 67),
('Januzzi', 30),
('Figliuzzi', 25),
('Fun', 20),
('Erika_Petruzzi', 20),
('zu', 20),
('Zo', 18),
('blog_BuzzMachine', 18),
('LW_Todd_Bertuzzi', 18),
('OFU', 17)]

You can customize Neofuzz’s behaviour by making a custom process. Under the hood every Neofuzz Process relies on the same two components:

  • A vectorizer, which turns texts into a vectorized form, and can be fully customized.
  • Approximate Nearest Neighbour search, which indexes the vector space and can find neighbours of a given vector very quickly.

Words as Features

If you’re more interested in the words/semantic content of the text you can also use them as features. This can be very useful especially with longer texts, such as literary works.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizer# Vectorization with words is the default in sklearn.vectorizer=TfidfVectorizer()
# We use cosine distance because it's waay better for high-dimentional spaces.process=Process(vectorizer, metric="angular")

Dimensionality Reduction

You might find that the speed of your fuzzy search process is not sufficient. In this case it might be desirable to reduce the dimentionality of the produced vectors with some matrix decomposition method or topic model.

Here for example I use NMF (excellent topic model and incredibly fast one too) too speed up my fuzzy search pipeline.

fromneofuzzimportProcessfromsklearn.feature_extraction.textimportTfidfVectorizerfromsklearn.decompositionimportNMFfromsklearn.pipelineimportmake_pipeline# Vectorization with tokens againvectorizer=TfidfVectorizer()
# Dimensionality reduction method to 20 dimensionsnmf=NMF(n_components=20)
# Create a pipeline of the twopipeline=make_pipeline(vectorizer, nmf)
process=Process(pipeline, metric="angular")

Semantic Search/Large Language Models

With Neofuzz you can easily use semantic embeddings to your advantage, and can use both attention-based language models (Bert), just simple neural word or document embeddings (Word2Vec, Doc2Vec, FastText, etc.) or even OpenAI’s LLMs.

We recommend you try embetter, which has a lot of built-in sklearn compatible vectorizers.

pip install embetter
fromembetter.textimportSentenceEncoderfromneofuzzimportProcess# Here we will use a pretrained Bert sentence encoder as vectorizervectorizer=SentenceEncoder("all-distilroberta-v1")
# Then we make a process with the language modelprocess=Process(vectorizer, metric="angular")
# Remember that the options STILL have to be indexed even though you have a pretrained vectorizerprocess.index(options)

About

Blazing fast fuzzy text search for Python.

Topics

Resources

Stars

52 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages