Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - henrywang/herdr-workflow: Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery. · GitHub
Skip to content

Repository files navigation

wq — agent workflows on top of herdr

wq runs opinionated multi-agent loops in real terminal panes: a planner drafts, a different model reviews adversarially, they iterate under a hard round cap, and the result ships as a merged pull request.

Status: all 13 commands work, and every one has been driven end to end against a real herdr daemon and real agents. The exception is the push-to-merge half of wq go, which is covered by tests against a fake gh rather than by a real merge.

Two ideas worth stealing, even if you never run this

1. Cross-model adversarial review with bounded rounds. The reviewer is deliberately not the model that wrote. Findings are classified BLOCKING or NON-BLOCKING, and the reviewer ends its file with exactly one machine-readable line:

VERDICT: APPROVED | VERDICT: CHANGES

Convergence becomes a grep, not an interpretation. Rounds are capped so a disagreeing pair cannot loop forever. Content moves between panes by path, never pasted — the single largest token sink in a loop like this.

2. docs/behaviors.md — what actually breaks when you drive coding-agent TUIs unattended. Thirteen failure modes discovered the hard way, each with the test that pins it. A pane reports "ready" while it will silently swallow your prompt. agent prompt returning OK does not mean the agent took the text. done is not a state that persists. gh pr checks says "no checks reported" before CI starts, which reads exactly like a failure and means the opposite. If you are scripting Claude Code, Codex, or Amp, you will hit these whether or not you use herdr.

Install

uv tool install herdr-workflow
wq doctor

You also need herdr running, git, and at least two agent CLIs — claude and pi are what the defaults assume. gh is needed only by wq ship / wq go. wq doctor checks all of it and explains anything missing.

Architecture

wq orchestrates; it does not run models itself. Herdr owns the terminal workspaces and panes, while the agent CLIs do their work against shared files and an isolated Git worktree.

flowchart LR
actor[User or agent router] --> wq[wq process]
wq -->|NDJSON socket API| herdr[herdr daemon]
subgraph terminal[Terminal workspaces managed by herdr]
panes[Planner, coder, and reviewer panes]
agents[Claude Code and Pi agents]
panes --> agents
end
herdr -->|Create panes, start agents, deliver prompts| panes
agents <--> artifacts[(Plans, reviews, and patches)]
agents <--> worktree[(Isolated Git worktree)]
wq -->|Check status and artifacts| artifacts
wq -->|Create branches, diff, push, and merge| worktree
Loading

Workflow

The same slug carries the plan, build, and review artifacts through the whole workflow. Both agent loops are bounded; when a cap is reached, you decide whether to continue.

flowchart TD
request["wq plan &lt;slug&gt; &lt;request&gt;"] --> draft[Planner drafts plan.md]
draft --> planReview[Different model reviews the plan]
planReview -->|Changes and rounds remain| planFix[Planner revises]
planFix --> planReview
planReview -->|Approved or round cap| inspect[You inspect plan.md]
inspect -->|Proceed| build["wq build &lt;slug&gt; &lt;repo&gt;"]
inspect -->|Do not proceed| stop[Stop or plan again]
build --> implement[Code agent implements in a worktree]
implement --> codeReview[Different model reviews the diff]
codeReview -->|Changes and rounds remain| codeFix[Code agent fixes and commits]
codeFix --> codeReview
codeReview -->|Approved| ready[Build ready]
codeReview -->|Round cap; exit 2| revise["wq revise &lt;slug&gt; &lt;comment&gt;"]
ready -->|Request another change| revise
revise --> oneRound[One code turn and one review turn]
oneRound --> ready
ready -->|Accept| ship["wq ship &lt;slug&gt;"]
ship --> go[wq go in a plain shell tab]
go --> pr[Push branch and open PR]
pr --> ci{CI passes?}
ci -->|No| repair[Fix the build, then run wq go again]
repair --> go
ci -->|Yes| merge[Squash-merge PR]
merge --> clean[Remove worktree, branches, and workspaces]
Loading

A worked example

One feature, start to finish. Each command blocks until its agents are done, then tells you the next one.

wq plan auth "add token refresh to the API client"

A planner and a reviewer open side by side in their own workspace. The planner writes plan.md; the reviewer attacks it and writes review.md ending in a verdict line. They iterate until it is approved or the round cap is hit. Read the plan before continuing — this is the cheap place to disagree.

wq build auth ~/code/my-api

A worktree on wq/auth, cut from whatever your repo actually branches from. A code agent implements and commits; a different model reviews the diff. Exits 2 if the round cap is reached with findings outstanding, so a script can tell "unreviewed code on a branch" from "wq broke".

wq revise auth "use the existing retry helper instead of a new one"

One more code turn and one review turn, driven by you rather than by the reviewer. Writes revise.patch — just what this turn changed — alongside the full diff.patch.

wq ship auth

Opens a plain shell tab and runs wq go there: push, PR, wait for CI, squash-merge, then remove the worktree, delete both branches and close the workspaces. ship returns immediately; you watch it happen in the tab.

At any point:

wq list # what is running; `*` marks the most recently worked build
wq clean auth # drop it all and start over

Commands

wq up # bring up inbox + router (idempotent)
wq chat "<message>"# reuse the inbox chat tab
wq ask "<question>"# new inbox tab, scoped to $PWD
wq tidy # close finished ask tabs
wq brainstorm <slug>"<idea>"# interactive; note lands in your notes sink
wq plan <slug>"<request>"# plan <-> review loop -> plan.md
wq build <slug> [repo] # worktree, code <-> review loop, commit
wq revise <slug>"<comment>"# one more code + review round on a build
wq ship <slug># run `wq go` in an inbox shell tab
wq go <slug># push, PR, wait for CI, merge, clean up
wq list # show active wq workspaces
wq clean <slug># drop the workspace and scratch dir
wq doctor # check the environment

Global: --json, --verbose, --debug, --config, --version.

Configuration

~/.config/wq/config.toml, overridden by .wq.toml in the project, overridden by WQ_* environment variables.

[agents]
plan = "claude:opus:high"code = "claude:sonnet:high"review = "pi:openai-codex/gpt-5.6-sol:high"# not the model that wrote
[loops]
plan_rounds = 3code_rounds = 3
[paths]
root = "~/Workspace/.wq"notes = ""# note sink for `wq brainstorm`
[herdr]
socket = "auto"

Roles are kind:model:level. Either side of a loop can be either agent — the rule that matters is that the reviewer is not the model that wrote.

Development

git clone https://github.com/henrywang/herdr-workflow
cd herdr-workflow
uv sync
uv run pytest # no herdr installation required

Tests run against a fake herdr daemon (tests/fake_herdr.py) that speaks the real newline-delimited JSON protocol over a real unix socket, so framing, id correlation, and event interleaving are all under test.

uv run ruff format .&& uv run ruff check .&& uv run pyright && uv run pytest

Integration tests need a running daemon: uv run pytest -m integration.

See CONTRIBUTING.md for how the pieces fit together.

  • docs/behaviors.md — the failure-mode catalogue, and the most useful file here if you are scripting agent TUIs at all
  • docs/protocol-framing.md — the herdr socket contract, as established by probing a real daemon
  • CHANGELOG.md — what changed, and the design decisions behind it

License

MIT

About

Opinionated multi-agent workflows for herdr, with cross-model planning, adversarial code review, bounded iteration, Git worktrees, and automated pull-request delivery.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages