Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

BisectRun — git-bisect for coding agents

BisectRun

git-bisect, but for Coding Agents. Find the exact span where two agent runs diverge.

npm versionlicensenodeteststypecheck

An Agent that fixed the bug last week shipped clean. Today's run blew up somewhere in 200 tool calls. git bisect finds the bad commit — but what finds the bad step? BisectRun replays two Claude Code session transcripts, aligns their span trees semantically, and pinpoints the first divergence: the single tool call where a passing run and a failing run parted ways.

BisectRun architecture atlas

Quickstart

Install and bisect two sessions in one command:

npx bisect-run@latest failed-session.jsonl good-session.jsonl

Boot the side-by-side diff UI:

npx bisect-run@latest failed-session.jsonl good-session.jsonl --ui
# → BisectRun UI → http://localhost:7331

Point it straight at a Claude Code project transcript by session id:

bisect-run ~/.claude/projects/-home-alice-myapp/abc123.jsonl \
~/.claude/projects/-home-alice-myapp/def456.jsonl --ui

Headless output

────────────────────────────────────────────────────
BisectRun — first divergence
────────────────────────────────────────────────────
matched depth: 1 span(s)
FAILED divergence: Edit [2026-07-22T09:00:12.000Z] — src/math.ts
GOOD divergence: Edit [2026-07-22T08:00:12.000Z] — src/math.ts
First divergence: same tool "Edit" but with different input —
failed: Edit:src/math.ts|return a + b;
good: Edit:src/math.ts|return a - b;

Demo

BisectRun demo

The tape that renders this gif lives in docs/demo.tape.


Why

Coding Agents are non-deterministic. The same prompt can produce a passing run today and a failing run tomorrow, and the diff between them is rarely a single line of code — it is a sequence of tool calls. Reviewing two 500-step transcripts by eye to find where they split is miserable.

BisectRun gives that workflow a name and a button:

  • Span-semantic alignment — two spans match when they invoke the same tool with the same intent (a stable fingerprint of the tool_use input), not when their byte-identical. Whitespace, uuids, and timestamps are ignored so "the same step" still matches across runs.
  • First divergence — the analogue of git bisect's first bad commit, but at the granularity of an agent step. That single span is almost always the thing worth staring at.
  • Signature check — the run signature (cwd + git branch + prompt) is compared up front so you don't bisect two unrelated tasks by accident.
  • Side-by-side UI — two span trees rendered next to each other, the divergence highlighted, so a human can see the fork instantly.

BisectRun is the missing observability layer for Coding Agents: turn "the agent regressed" into "the agent regressed at this span."


How it works

  1. Parse — read a Claude Code .jsonl transcript and collapse each tool_use + its tool_result into a Span, stitching them into a tree via the records' parentUuid/uuid linkage. The first user message, cwd, and gitBranch become the run's signature.
  2. Align — flatten both span trees to chronological order and walk them index by index, comparing per-tool fingerprints (Bash → command, Edit → file + old_string head, Read → path, …). The first index where fingerprints disagree is the divergence.
  3. Report — headless mode prints the divergence to stdout; --ui serves a Hono page with both trees side by side and the divergence highlighted.

Project layout

src/
types.ts Span, RunSignature, AlignmentResult
parse-jsonl.ts transcript → span tree + signature
align.ts span-semantic alignment + first divergence
cli.ts `bisect-run` commander entry
server.ts Hono diff UI
test/
fixtures/ failed/good session JSONL
parse.test.ts align.test.ts
assets/ animated hero + atlas SVG (light/dark)
docs/demo.tape vhs tape rendering the demo gif

Tests

npm test# node:test, 10 cases
npm run typecheck
npm run build # tsc → dist/

Roadmap

  • Auto-locate transcripts from (cwd, sessionId) without manual paths
  • Subtree alignment for parallel / branching tool calls
  • LLM-assisted semantic diff of the divergent span's input
  • CI-integrated: bisect a PR's agent run against main and fail on regression
  • Export divergence as a shareable, replayable link
  • Support for non-Claude-Code transcript formats

License

MIT © 2026 SuperMarioYL