Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

verd

Five minds enter. They argue, challenge, cross-examine. Only the truth walks out.

verd spawns multiple AI models from different families — each with a specialized role, has them debate your question across rounds, then a stronger judge delivers the final verdict with strengths, issues, and actionable fixes.

Use it everywhere: CLI for code reviews, MCP inside Claude Code and Cursor, and Slack as @verd in any conversation.

Getting Started

Requires Python 3.11+.

pip install verd
verd setup

The setup wizard walks you through provider selection (OpenRouter, LiteLLM, or other) and outputs the exact config you need — for both CLI (.env) and MCP (JSON to paste into your editor config).

verd runs multiple models in parallel (Claude, Gemini, GPT, DeepSeek) so it needs a multi-provider router. OpenRouter is the easiest — one key, all models. LiteLLM proxy works too.

Usage

CLI

verd "can this auth middleware be bypassed?" -f auth.py middleware.py
verdh "should we merge this?" -gb main # deep mode — 5 models, 3 rounds

MCP (Claude Code / Cursor) — use verd, verdl, verdh as tools directly in chat:

verdh based on the context above and this file, do you think we can proceed?
is this approach correct given what we discussed, use verdl

Slack — mention @verd in any channel or thread:

@verd what do you think? — reads thread context, debates, replies
@verd deep is this secure? — uses verdh (5 models, 3 rounds)
/verd Kafka, SQS, or RabbitMQ for our event pipeline? — slash command with live progress

verd is for critical decisions and deep analysis — not simple lookups. If a single model can answer it, verd is overkill.

Output

FAIL 77% In-memory rate limiter is unsafe for production
claude:FAIL gpt:FAIL gemini:FAIL gpt:FAIL (FULL)
+ Conceptually correct sliding-window logic
- Global dict is unsynchronized — race conditions in multi-thread servers
- Per-user lists grow without bounds — memory leak / DoS vector
! gpt-5-mini caught the risk of system clock jumps with time.time()
→ Move state to Redis with atomic operations
completed in 69.3s • 22,449 tokens • ~$0.07

Vote breakdown, unique catches (!), dissent, strengths, issues, and actionable fixes — all in one view.

Modes

CommandDebatersRolesRoundsSpeedCost
verdl2 + judgeanalyst, devils_advocate1~15s+~$0.01
verd4 + judgeanalyst, devils_advocate, logic_checker, pragmatist2~30s+~$0.05+
verdh5 + judgeanalyst, devils_advocate, logic_checker, fact_checker, pragmatist3~60s+~$0.25+

Benchmark

Tested on the Martian Code Review Benchmark — 50 real PRs from Cal.com, Discourse, Grafana, Keycloak, and Sentry with expert-labeled golden comments. No code-review-specific tuning.

ModePrecisionRecallF1 ScoreAvg Issues
GPT-5.4 (alone)13.0%70.6%21.9%14.6
Claude Opus 4.6 (alone)18.5%69.9%29.2%10.1
verdh (5-model debate)29.1%64.0%40.0%5.9

+37% F1 over Claude solo. 57% more precise. 42% fewer false positives.

How it works

  1. Your question + content gets sent to multiple AI models in parallel
  2. Each model has a specialized role (analyst, devils_advocate, logic_checker, fact_checker, pragmatist)
  3. Models see each other's responses and cross-examine for 1-3 rounds
  4. Anti-groupthink prompts ensure models hold their ground when they have evidence — consensus without new evidence is rejected
  5. A stronger judge model synthesizes the debate, weighting each reviewer by their role
  6. Confidence is calculated from vote distribution — a fact_checker's dissent lowers confidence more than a devils_advocate's expected pushback
  7. You get: verdict, vote breakdown, strengths, issues, unique catches, dissent, and actionable fixes

The key insight: different model families have different blind spots and training biases. Claude spots nuance GPT misses. Gemini catches logic errors DeepSeek overlooks. More importantly — if the same model writes the review and judges its quality, it's likely to agree with itself. Cross-model diversity means the judge is a genuine quality gate, not a model grading its own homework. The debate surfaces what each model uniquely caught and tells you exactly which model caught what.

Roles

RoleJobExample catch
analystBalanced initial assessment, main arguments for and against"The architecture is sound but the auth flow has a gap"
devils_advocateFind what others miss — edge cases, hidden assumptions, failure modes"What happens when the token expires mid-transaction?"
logic_checkerVerify reasoning quality — fallacies, off-by-one, race conditions"The pagination math is wrong: total_pages needs ceil division"
fact_checkerWeb-grounded verification — do these APIs/libraries actually work?"That library was deprecated in v3, use the new API"
pragmatistReal-world practicality — will this ship? What's the ops burden?"This works but needs 3 new infra dependencies your team doesn't know"

The judge weighs each reviewer's input by role — a fact_checker citing sources carries more weight than a devils_advocate pushing back.

Config

Override models via env vars or CLI flags. Per-tier env vars let you set different models for each mode:

VERDL_JUDGE=o4-mini VERDL_DEBATERS=gpt-4.1-mini,gemini-3.1-flash-lite-preview
VERD_JUDGE=o3 VERD_DEBATERS=claude-sonnet-4-6,gpt-4.1,gemini-3.1-pro-preview,gpt-4.1-mini
VERDH_JUDGE=o3 VERDH_DEBATERS=claude-opus-4-6,deepseek-r1,gemini-3.1-pro-preview,sonar-pro,gpt-4.1

Or use VERD_JUDGE / VERD_DEBATERS as a global override for all tiers. verd setup generates the right config for your provider.

Flags

-c TEXT inline content string
-f FILE [FILE ...] one or more files to evaluate
-d [DIR] read all files in a directory (default: current dir)
-g use unstaged git diff as content
-gs use staged git diff as content
-gb REF use git diff REF...HEAD as content (e.g. main)
-a / --all scan all files, skip smart selection (use with -d)
--ext EXT [EXT ...] filter by extension (use with -d)
--exclude PATTERN glob patterns to exclude (use with -d)
-q / --quiet hide debate transcript, show only verdict
--json output raw JSON
--judge MODEL override judge model
--debaters MODEL ... override debater models
--budget USD max cost in USD — abort if estimate exceeds budget
--timeout SECONDS override timeout per model call
--version show version and exit

MCP — Claude Code / Cursor

verd setup # select "MCP" and your provider

This prints the exact JSON to paste into ~/.claude/settings.json (Claude Code) or ~/.cursor/mcp.json (Cursor), with the correct absolute path to verd-mcp and model overrides for your provider. Then use verd, verdl, or verdh as tools directly in chat.

Slack

Install with Slack dependencies:

pip install "verd[slack]"

Create a Slack app with Socket Mode enabled, add bot scopes (app_mentions:read, channels:history, groups:history, chat:write, reactions:write, im:history, im:write, users:read), then:

export SLACK_BOT_TOKEN=xoxb-...
export SLACK_APP_TOKEN=xapp-...
export SLACK_SIGNING_SECRET=...
verd-slack

Optional: restrict access via environment variables:

export VERD_ALLOWED_CHANNELS=C123,C456 # empty = all channelsexport VERD_ALLOWED_USERS=U123,U456 # empty = all users

About

Multi-LLM debate engine for higher-confidence decisions

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages