Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

FuzzySharp

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Release Notes:

v.2.0.0

As of 2.0.0, all empty strings will return a score of 0. Prior, the partial scoring system would return a score of 100, regardless if the other input had correct value or not. This was a result of the partial scoring system returning an empty set for the matching blocks As a result, this led to incorrrect values in the composite scores; several of them (token set, token sort), relied on the prior value of empty strings.

As a result, many 1.X.X unit test may be broken with the 2.X.X upgrade, but it is within the expertise fo all the 1.X.X developers to recommednd the upgrade to the 2.X.X series regardless, should their version accommodate it or not, as it is closer to the ideal behavior of the library.

Usage

Install-Package FuzzySharp

Simple Ratio

Fuzz.Ratio("mysmilarstring","myawfullysimilarstirng")72
Fuzz.Ratio("mysmilarstring","mysimilarstring")97

Partial Ratio

Fuzz.PartialRatio("similar","somewhresimlrbetweenthisstring")71

Token Sort Ratio

Fuzz.TokenSortRatio("order words out of"," words out of order")100
Fuzz.PartialTokenSortRatio("order words out of"," words out of order")100

Token Set Ratio

Fuzz.TokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100
Fuzz.PartialTokenSetRatio("fuzzy was a bear","fuzzy fuzzy fuzzy bear")100

Token Initialism Ratio

Fuzz.TokenInitialismRatio("NASA","National Aeronautics and Space Administration");89
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration");100
Fuzz.TokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");53
Fuzz.PartialTokenInitialismRatio("NASA","National Aeronautics Space Administration, Kennedy Space Center, Cape Canaveral, Florida 32899");100

Token Abbreviation Ratio

Fuzz.TokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);40
Fuzz.PartialTokenAbbreviationRatio("bl 420","Baseline section 420",PreprocessMode.Full);50

Weighted Ratio

Fuzz.WeightedRatio("The quick brown fox jimps ofver the small lazy dog","the quick brown fox jumps over the small lazy dog")95

Process

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"})(string:DallasCowboys,score:90,index:3)
Process.ExtractTop("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},limit:3);[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:bing,score:22,index:1),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:twitter,score:15,index:4),(string:googleplus,score:75,index:5),(string:bingnews,score:29,index:6),(string:plexoogl,score:43,index:7)]// score cutoff
Process.ExtractAll("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"},cutoff:40)[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7)]
Process.ExtractSorted("goolge",new[]{"google","bing","facebook","linkedin","twitter","googleplus","bingnews","plexoogl"})[(string:google,score:83,index:0),(string:googleplus,score:75,index:5),(string:plexoogl,score:43,index:7),(string:facebook,score:29,index:2),(string:linkedin,score:29,index:3),(string:bingnews,score:29,index:6),(string:bing,score:22,index:1),(string:twitter,score:15,index:4)]

Extraction will use WeightedRatio and full process by default. Override these in the method parameters to use different scorers and processing. Here we use the Fuzz.Ratio scorer and keep the strings as is, instead of Full Process (which will .ToLowercase() before comparing)

Process.ExtractOne("cowboys",new[]{"Atlanta Falcons","New York Jets","New York Giants","Dallas Cowboys"}, s =>s,ScorerCache.Get<DefaultRatioScorer>());(string:DallasCowboys,score:57,index:3)

Extraction can operate on objects of similar type. Use the "process" parameter to reduce the object to the string which it should be compared on. In the following example, the object is an array that contains the matchup, the arena, the date, and the time. We are matching on the first (0 index) parameter, the matchup.

varevents=new[]{new[]{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},new[]{"new york yankees vs boston red sox","Fenway Park","2011-05-11","8pm"},new[]{"atlanta braves vs pittsburgh pirates","PNC Park","2011-05-11","8pm"},};varquery=new[]{"new york mets vs chicago cubs","CitiField","2017-03-19","8pm"};varbest=Process.ExtractOne(query,events, strings =>strings[0]);best:(value:{"chicago cubs vs new york mets","CitiField","2011-05-11","8pm"},score:95,index:0)

FuzzySharp in Different Languages

FuzzySharp was written with English in mind, and as such the Default string preprocessor only looks at English alphanumeric characters in the input strings, and will strip all others out. However, the Extract methods in the Process class do provide the option to specify your own string preprocessor. If this parameter is omitted, the Default will be used. However if you provide your own, the provided one will be used, so you are free to provide your own criteria for whatever character set you want to admit. For instance, using the parameter (s) => s will prevent the string from being altered at all before being run through the similarity algorithms.

E.g.,

varquery="strng";varchoices=new[]{"stríng","stráng","stréng"};varresults=Process.ExtractAll(query,choices,(s)=>s);

The above will run the similarity algorithm on all the choices without stripping out the accented characters.

Using Different Scorers

Scoring strategies are stateless, and as such should be static. However, in order to get them to share all the code they have in common via inheritance, making them static was not possible. Currently one way around having to new up an instance everytime you want to use one is to use the cache. This will ensure only one instance of each scorer ever exists.

varratio=ScorerCache.Get<DefaultRatioScorer>();varpartialRatio=ScorerCache.Get<PartialRatioScorer>();vartokenSet=ScorerCache.Get<TokenSetScorer>();varpartialTokenSet=ScorerCache.Get<PartialTokenSetScorer>();vartokenSort=ScorerCache.Get<TokenSortScorer>();varpartialTokenSort=ScorerCache.Get<PartialTokenSortScorer>();vartokenAbbreviation=ScorerCache.Get<TokenAbbreviationScorer>();varpartialTokenAbbreviation=ScorerCache.Get<PartialTokenAbbreviationScorer>();varweighted=ScorerCache.Get<WeightedRatioScorer>();

Credits

  • SeatGeek
  • Adam Cohen
  • David Necas (python-Levenshtein)
  • Mikko Ohtamaa (python-Levenshtein)
  • Antti Haapala (python-Levenshtein)
  • Panayiotis (Java implementation I heavily borrowed from)

About

C# .NET fuzzy string matching implementation of Seat Geek's well known python FuzzyWuzzy algorithm.

Resources

Stars

797 stars

Watchers

18 watching

Forks

Releases

Packages

Used by

Contributors

Languages