Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

diffract

An OCaml library and CLI tool for parsing source files using tree-sitter and pattern matching with concrete syntax.

Features

  • Parse source using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Generic patch transforms in the style of Coccinelle (not-really semantic patches)
  • Expansion transforms: join or restructure each element of a matched sequence
  • Change summaries (summarize): infer the spatch rules behind a changeset — cluster the systematic edits in a before/after directory pair into rules, with per-file residuals for everything the rules don't explain
  • Support for TypeScript, TSX, Kotlin, PHP, Scala (extensible)

Building

Prerequisites

  • OCaml 5.4+
  • opam
  • tree-sitter library (libtree-sitter) - see below
  • npm (to fetch grammar sources)

Installing tree-sitter

macOS:

brew install tree-sitter

Ubuntu/Debian:

sudo apt install libtree-sitter-dev

Arch Linux:

sudo pacman -S tree-sitter

Build Steps

# Install OCaml dependencies (add --with-test to include test deps)
opam install . --deps-only --with-test
# Build grammar libraries (TypeScript, Kotlin)cd grammars && ./build-grammars.sh &&cd ..
# Build the project
dune build
# Run tests
dune test# Format code (requires ocamlformat: opam install ocamlformat)
dune fmt
# Enable the pre-push hook (builds & tests the pushed commit in a clean# worktree, so untracked files can't mask a broken commit)
git config core.hooksPath .githooks

Formatting is managed via ocamlformat. Run dune fmt to reformat all OCaml and dune files before committing.

Because of caching, rerunning dune fmt might not produce warnings for things like stray @ in doc comments. One can just run ocamlformat manually:

ocamlformat --check $(find lib -name '*.ml' -o -name '*.mli')

Usage

# Parse and print the syntax tree
diffract parse example.ts
diffract parse --language kotlin example.kt
# Search for a pattern in a single file
diffract search pattern.txt source.ts
# Scan a directory for pattern matches
diffract search --include '*.ts' pattern.txt src/
# Scan with custom directory exclusions
diffract search --include '*.ts' -e vendor -e dist pattern.txt src/
# Show how a pattern tokenizes (diagnose a pattern that matches nothing)
diffract search --debug-tokens pattern.txt source.ts
# Apply a semantic patch (preview diff)
diffract apply patch.txt source.ts
# Apply a semantic patch in place
diffract apply --in-place patch.txt source.ts
# Apply across a directory
diffract apply --include '*.ts' patch.txt src/
# Show AST-level changes between two file versions
diffract diff before.ts after.ts
# Summarize a changeset: infer the rules behind a before/after directory pair
diffract summarize -l typescript -i '*.ts' before/ after/
# List available languages
diffract languages

Transforms (Semantic Patches)

Patterns can include -/+ prefixed lines to describe code transformations. For example, to rename console.log to logger.info:

patch.txt:

@@
match: strict
metavar $MSG: single
@@
- console.log($MSG)
+ logger.info($MSG)
$ diffract apply patch.txt source.ts
--- a/source.ts
+++ b/source.ts
@@ -1,3 +1,3 @@
functiongreet(name: string) {
- console.log(name);
+ logger.info(name);
}

Lines prefixed with - are matched and removed; lines with + are inserted. Unprefixed (or space-prefixed) lines are context that appears in both match and replace. Metavariables carry values from the match side to the replace side.

Sequence rendering: join and foreach

A sequence metavar referenced in a + template is rendered and substituted in place. Two knobs: a join $VAR by "<sep>" preamble directive sets the separator between elements (default: empty), and a following foreach $VAR section applies a per-element transform.

For per-element transforms (e.g. converting a match expression to a method chain), use a two-section pattern with foreach:

@@
match: strict
metavar $TAG: single
metavar $CASES: sequence
@@
- matchExhaustive($TAG, { $CASES });
+ match($TAG)$CASES.exhaustive();
@@
match: strict
foreach $CASES
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

Applied to:

matchExhaustive(tag,{A: ()=>1,B: ()=>2});

Produces:

match(tag).with("A",()=>1).with("B",()=>2).exhaustive();

See Transform documentation for partial-mode, field-mode, and sequence transforms.

Change Summaries (summarize)

summarize runs the transform machinery in reverse: given a before/after pair of directory trees, it infers the semantic-patch rules behind the changeset. Instead of reading the same edit repeated across a large diff, a reviewer reads each rule once and then inspects only the residuals — per-file diffs of whatever no rule explains. For example, a changeset that converts Kotlin block bodies to expression bodies:

// before/a.kt // after/a.ktfunlabel(): String { funlabel(): String="repo"return"repo"
}
overridefuncount(): Int { overridefuncount(): Int= items.size
return items.size
}
$ diffract summarize -l kotlin -i '*.kt' before/ after/
# rule R1 support=6 language=kotlin
@@
match: field
metavar _H0: single
metavar _H1: single
@@
fun _H1(...)
- {
- return _H0
- }
+ = _H0
# sites R1
a.kt
b.kt

One match: field rule covers every site even though the signatures differ — visibility modifiers, override, return types, parameter lists — because field mode aligns only the parts the pattern addresses. And the rules are safe: a rule claims a file only when applying it there is verified to reproduce part of that file's actual change, so look-alike sites that did something subtly different are reported as residuals instead of being claimed wrongly.

On a real-world changeset this example is distilled from -- a mechanical refactor touching ~100 Kotlin files, ~2,700 changed diff lines -— summarize reduces the whole diff to two rules plus two small residuals.

See Change summaries for the output format, more worked examples, and --ignore-formatting.

Directory Scanning

When the target is a directory, use --include to specify which files to scan:

OptionDescription
--include GLOB / -iGlob pattern for files (e.g., *.ts, *.py). Required for directories.
--exclude DIR / -eDirectory names to skip (repeatable). Defaults: node_modules, .git, _build, target, __pycache__, .hg, .svn

Supported glob patterns:

  • *.ts - files ending with .ts
  • prefix* - files starting with prefix
  • *suffix - files ending with suffix

Example output:

src/api/auth.ts:15: console.log("login")
$msg = "login"
src/utils/logger.ts:8: console.log("initialized")
$msg = "initialized"
Found 2 match(es) in 2 file(s) (scanned 47 files)

Documentation

User guides:

Architecture and design:

Matcher Architecture (Quick Overview)

The matching pipeline is split into focused modules:

  • tokenize parses a pattern body with tree-sitter as a lexer, keeping leaves as a (text, node_type) token stream (sigil-free metavars, ellipsis, fragments).
  • cursor / tree_sitter_cursor define and implement the tree-cursor interface the engine runs over.
  • stmatch is the matching engine: leaf-level strict / partial / field matching with backtracking, over any Cursor.S.
  • matcher ties it together — preamble parse → tokenize → match → transform — and exposes the public find / transform / debug_tokens / pattern_warnings API.

License

GPL-3.0-or-later

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages