Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - ConnorBritain/agent-time: Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history. · GitHub
Skip to content

Latest commit

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

agent-time-estimator

An agent-native estimation layer for Claude Code. Estimates coding tasks in the unit agents actually work in — tool-call rounds (discovery, implementation, verification, rework) — then converts to a low/expected/high wall-clock range using local calibration history that self-corrects as outcomes are logged.

Why: coding agents inherit human developer heuristics from training data, so a 30-minute agent task gets called "2–3 days". This skill never starts from human hours.

What's here

skills/agent-time-estimator/
SKILL.md # the Claude Code skill (3 modes: estimate / re-estimate / calibrate)
estimator.py # stdlib-only engine — all the math, transparent constants up top
references/model.md # full model: priors tables, risk math, calibration cascade
hooks/ # optional PostToolUse/Stop hooks (auto round-counting)
examples/
sample-history.jsonl # seed calibration data (7 realistic outcomes)
example-session.md # pre → during → post walkthrough
.claude/settings.json # live hook wiring (+ .example to port elsewhere)
.claude/skills/agent-time-estimator # symlink so the skill is live in this repo
tests/test_estimator.py # stdlib unittest suite

Calibration history lives at <repo>/.claude/agent-time/history.jsonl (JSONL, one record per line, human-editable). Override with --data or $AGENT_TIME_DATA. A user-level ~/.claude/agent-time/history.jsonl is auto-blended when the project's own history is sparse (see cross-repo calibration below).

Beyond the basics

Three layers make estimates sharper than a static prior:

  • Automatic actuals (hooks). An optional PostToolUse hook counts rounds and records edited files + commands per session, so log/reestimate fill in actuals with no self-report. Wired live here via .claude/settings.json; copy .example to port it. See hooks/README.md.
  • Per-model reliability (METR-style). Each model has a single-run horizon (opus-4.8 = 90 min, sonnet-4.6 = 30 min, …). When expected wall-clock exceeds its model's horizon, the estimate recommends a checkpoint/split. A pace factor also adjusts minutes-per-round once enough per-model data exists. Pass --model.
  • Cross-repo calibration. User-level history blends into sparse local buckets (local weighs 3× user), so a fresh repo is useful on day one. Disable with --no-user-history.

Pairs with roadmap

roadmap — a tool for scoping and orchestrating AI coding sessions across a repo — consumes these estimates at the macro scale. It tags each planned slice with a shape (+ risks), calls estimate --json per slice, and rolls the durations up its dependency/concurrency schedule into projected per-project target dates on a Linear timeline (roadmap estimate / roadmap estimate timeline), feeding actuals back to log as slices complete.

This skill stays standalone and unaware of it — roadmap is just another caller of the CLI, so nothing here changes. The payoff of the pairing: the same calibrated numbers that size one task also plan the whole program, and every completed slice sharpens the next estimate. See roadmap's Estimation & timeline section.

Install in another repo

Copy (or submodule) the skill directory and expose it to Claude Code:

cp -r skills/agent-time-estimator /path/to/repo/.claude/skills/
# or user-wide:
cp -r skills/agent-time-estimator ~/.claude/skills/
# optional: enable auto round-counting hooks
cp .claude/settings.json.example /path/to/repo/.claude/settings.json

No dependencies. Python 3.8+.

The three modes

1. Before a task — estimate

Claude runs a bounded discovery pass (≤10 tool calls), classifies the task shape, selects evidenced risk factors, then calls:

python3 skills/agent-time-estimator/estimator.py estimate \
--summary "add rate limiting to the webhook endpoint" \
--shape localized-feature \
--risks external-api,no-tests \
--findings "handler in src/hooks.py; no middleware; no tests for hooks" \
--assumptions "in-memory limiter acceptable"

Output (abridged):

## Agent Time Estimate
Task shape: localized-feature
Round estimate (mode per category, before risk): Discovery 3, Implementation 5, Verification 3, Rework 2 ...
Risk factors: external-api (×1.25), no-tests (×1.20) → combined ×1.45
Calibration basis: calibrated (global median, n=7) · 2.82 min/round · sparse-data widening ×1.27
Agent rounds (low/expected/high): 10.2 / 18.9 / 30.9
Wall-clock estimate: Low: 19 min · Expected: 45 min · High: 93 min
Confidence: low-medium
Checkpoint: re-estimate after the first verification round if anything surprised you

The estimate is logged as pending with a task id; the wall clock starts there.

2. During a task — re-estimate

When a hidden dependency, failing suite, or environment issue appears (or rounds spent exceed ~1.5× expected):

python3 skills/agent-time-estimator/estimator.py reestimate \
--task-id t-20260708-a1b2 --rounds-spent 12 \
--reason "test suite fails on main; must fix fixtures first" \
--add-risks slow-flaky-tests \
--rounds-json '{"implementation": 3, "verification": 5, "rework": 4}'

The engine blends the task's own observed pace (60% weight after 5+ rounds), shows the original vs. revised numbers, and appends a revision record — the original estimate is never rewritten.

3. After a task — calibrate

python3 skills/agent-time-estimator/estimator.py log \
--task-id t-20260708-a1b2 --status pass \
--variance "fixtures were broken on main, cost ~5 extra rounds"

With the hook enabled, actual rounds, files changed, and commands are auto-filled from session activity — no --actual-rounds needed (pass it to override). Elapsed wall time is computed from the estimate timestamp. Each logged outcome tightens future ranges: median minutes-per-round by similar-task match → task-shape → global, plus a recency-weighted bias correction. estimator.py stats shows calibration health, including a per-model pace/horizon table.

How estimates improve over time

Logged outcomesBehavior
0 local, but user-level existsblends in cross-repo history (labeled user-level/blended)
0–4 (no user data)UNCALIBRATED static prior (3 min/round), ranges widened ×~1.75
5+global median pace, bias correction kicks in, ranges narrow as 1 + 0.75/√(n+1)
3+ of same shapeper-shape pace
3+ similar tasks (keyword match)similar-task pace — the strongest signal
5+ per modelper-model pace factor + reliability-horizon checkpointing
15+ predicted outcomes in bucketfully empirical p20/p80 spread

Full math: skills/agent-time-estimator/references/model.md.

Seeding

Start from the sample data to see calibrated behavior immediately:

mkdir -p .claude/agent-time
cp skills/agent-time-estimator/examples/sample-history.jsonl .claude/agent-time/history.jsonl

Or backfill real past tasks: estimator.py log without --task-id (requires --summary --shape --actual-minutes --actual-rounds).

Tests

python3 -m unittest discover tests

Design principles

Practical over academic · local-first (no services, no network) · transparent math (every constant in one CONFIG block) · useful from day one, better with history · never a single precise ETA.

About

Agent-native time estimation for coding agents: predicts work in discovery, implementation, verification, and rework rounds, then calibrates wall-clock ranges from real task history.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages