Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

NAME

librx - Regex library inspired by Thompson NFAs, Perl 6, and PCRE

SYNOPSIS

#include <rx.h>
int main {
Rx *rx = rx_new(
"([chapter|page|line] - <digit>+) [',' \\s* <~~0>] ** 1..2"
);
if (rx_match(rx, "chapter-55, page-44, line-33"))
printf("it matches!\n");
rx_free(rx);
}

DESCRIPTION

This regular expression library is based on a Thompson NFA rather than a backtracking NFA. I originally read about Thompson's NFA in Russ Cox's article, "Regular Expression Matching Can Be Simple And Fast". It also uses some syntactic features found in Perl 6 Synopse 05.

During a match all possible paths are explored until there are no paths left or no characters left in the string.

FUNCTIONS

  • Rx *rx_new(const char *rx_str)

    Allocate a new Rx object from a string containing the regular expression.

  • int rx_match(Rx *rx, const char *str)

    Match the regex against a string. Returns whether it matched. Eventually this should fill in a match object which will allow one to find out what matched and the groups that matched in it.

  • void rx_free(Rx *rx)

    Frees the memory of a regex previously created by rx_new().

  • int rx_debug

    You may set this global variable to cause rx_new() to print out a representation of the regex to stdout and rx_match() will print out its list of paths and matches after each character of the string is read.

SYNTAX

Many regex features that one may expect are supported.

An alphanumeric character, _, or - will match itself. All other characters need to be escaped with a backslash or enclosed in quotes or character classes.

Note that in C, double quoted strings interpolate escapes, so you have to escape all backslashes before sending them to rx_new().

You may quote a string of characters with single (') or double (") quotes and its contents will match unaltered. There is no difference between single and double quotes except double quotes allow for escapes. For example, '*runs away*' will match the string "*runs away*".

All whitespace is insignificant except in quoted forms.

A | separates alternate matches.

Each atom may have a quantifier after it.

  • * matches 0 or more times
  • + matches 1 or more times
  • ? matches 0 or 1 times
  • ** n matches n times
  • ** n..m matches at least n times and at most m times
  • ** n..* matches n or more times

You may group a portion of the regex in parentheses ( which may be used as any other atom and referenced later either with <~~#> or through the Match object. There is also the ability to group without capturing with square brackets [.

An extensible meta-syntax of the form <...> has been added to implement special features much like the Perl 5 construct of (?...).

You can refer to the pattern in previous groups by referencing them as a number in the extensible meta syntax. /(cool)<~~0>/. These can even refer to its own group recursively. You can refer to the whole pattern by using <~~>.

The . character really matches any character. If you want everything but a newline, use \N. Also, there are escapes \T and \R for anything but \t and \r.

Escaped character classes \w matches a word char, \s matches a space char, and \d matches a digit. They may be negated with \W, \S, and \D which will match anything but what their lower case version would match.

A character class is specified with <[...]>. For example, <[a..z_]>, specifies any character from a to z or _. Whitespace is ignored in this construct, and you can combine character classes by adding and subtracting them like this <[a..z] + ['] - [m..q]>. Negated character classes start with a -, so <-[aeiou]> matches anything but a vowel.

The following named character classes are allowed as well: upper, lower, alpha, digit, xdigit, print, graph, cntrl, punct, alnum, space, blank, and word. They may be combined with + and - just as the bracketed char classes can. <[_] + alpha + punct> or used on their own like <print>.

Assertions ^ matches the beginning of the string, ^^ matches the beginning of a line, $ matches the end of the string, $$ matches the end of a line, << matches a left word boundary, >> matches a right word boundary, \b matches a word boundary regardless of being on the left or right side, and \B matches a non-word boundary.

About

Regex library inspired by Thompson NFAs, Perl 6, and PCRE

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages