feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: add context efficiency metrics (#505) - #515

Open
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric
Open

feat: add context efficiency metrics (#505)#515
willwashburn wants to merge 7 commits into
mainfrom
issue-505-context-efficiency-metric

Conversation

@willwashburn

@willwashburnwillwashburn commented Aug 3, 2026

Copy link
Copy Markdown
Member

Summary

  • Defines context efficiency in relayburn-sdk as input, cache-read, and cache-creation tokens per generated token, with consistent reasoning normalization across harnesses.
  • Adds aggregate ratios and bounded per-session p50/p95/max context-size distributions to Rust SDK, CLI, JSON, and Node summary surfaces.
  • Adds a cost-independent context-output-ratio finding with configurable ratio and minimum-context thresholds; defaults flag the motivating 382:1 incident at a 1M-context floor.
  • Handles zero-output turns without non-finite JSON and documents the metric definition in public API comments and changelogs.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace --quiet

Fixes#505

🤖 Generated with Claude Code.

Review in cubic

@coderabbitai

coderabbitaiBot commented Aug 3, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@willwashburn, you've reached your PR review limit, so we couldn't start this review.

Next review available in:37 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 093697b6-b172-46d4-a0b1-947030f506a9

📥 Commits

Reviewing files that changed from the base of the PR and between cca7987 and 5ffdfce.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/mcp_server.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/compute.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts
  • packages/sdk-node/src/index.js
  • packages/sdk-node/test/conformance.test.js
  • tests/fixtures/cli-golden/snapshots/summary-json.stdout.txt
  • tests/fixtures/cli-golden/snapshots/summary.stdout.txt
📝 Walkthrough

Walkthrough

The SDK now computes context-efficiency metrics for turns and sessions. Summary commands and Node bindings expose these metrics. Hotspots support configurable context-to-output ratio findings with minimum context-token thresholds.

Changes

Context efficiency

Layer / File(s)Summary
Context-efficiency metrics
crates/relayburn-sdk/src/query_verbs/context_efficiency.rs, crates/relayburn-sdk/src/query_verbs/mod.rs, crates/relayburn-sdk/src/analyze/*
The SDK calculates context tokens, normalized output tokens, ratios, percentiles, session rankings, projections, and validation results.
Summary integration and presentation
crates/relayburn-sdk/src/query_verbs/summary/mod.rs, crates/relayburn-cli/src/commands/summary/*, crates/relayburn-sdk/src/query_verbs/tests.rs
Summary reports include context-efficiency data. CLI output renders aggregate metrics and up to ten ranked sessions. Tests cover JSON, grouped reports, and session values.
Context-output hotspot findings
crates/relayburn-sdk/src/query_verbs/hotspots.rs, crates/relayburn-sdk/src/analyze/findings.rs, crates/relayburn-cli/src/commands/hotspots/mod.rs, crates/relayburn-sdk/src/query_verbs/tests.rs
Hotspots accept ratio and minimum-token thresholds and generate cost-independent context-output-ratio findings.
Node API and documentation
crates/relayburn-sdk-node/src/lib.rs, packages/sdk-node/src/index.d.ts, CHANGELOG.md, packages/sdk-node/CHANGELOG.md
Node bindings and TypeScript declarations expose context-efficiency results and BigInt counters. Changelogs document the new metrics and options.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
participant LedgerHandle
participant ContextEfficiency
participant Summary
participant CLI_or_Node
LedgerHandle->>ContextEfficiency: compute turn and session metrics
ContextEfficiency->>Summary: populate contextEfficiency
Summary->>CLI_or_Node: serialize or render metrics
LedgerHandle->>ContextEfficiency: evaluate hotspot thresholds
ContextEfficiency-->>LedgerHandle: return ratio findings
Loading

Poem

I’m a rabbit with tokens to spare,
Counting context through sessions with care.
Ratios now rise,
Findings surprise,
And BigInts hop everywhere.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check nameStatusExplanation
Title check✅ PassedThe title clearly identifies the primary change: adding context efficiency metrics.
Description check✅ PassedThe description directly explains the context efficiency metrics, findings, API surfaces, validation, and linked issue.
Linked Issues check✅ PassedThe changes implement the linked issue objectives for summary metrics, per-session distributions, and configurable cost-independent ratio findings [#505].
Out of Scope Changes check✅ PassedThe changes remain within the linked issue scope and support context efficiency metrics, findings, integrations, tests, and documentation.
Docstring Coverage✅ PassedDocstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch issue-505-context-efficiency-metric

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
packages/sdk-node/CHANGELOG.md (1)

5-5: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Split the bullet into one entry per API change.

This bullet covers two separate user-visible changes: the summary() result shape and the new hotspots() options. The content is accurate against packages/sdk-node/src/index.d.ts Lines 305-320.

📝 Proposed changelog edit
-- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions, where context is input + cache-read + cache-creation tokens; `hotspots()` accepts ratio and minimum-context options for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).+- `summary()` now returns context tokens per normalized generated-output token (including reasoning) plus p50/p95/max context sizes for the ten highest-ratio sessions. Context is input + cache-read + cache-creation tokens.+- `hotspots()` accepts `contextOutputRatioThreshold` and `contextOutputMinTokens` for the new cost-independent finding (defaults: 382:1 inclusive and 1M context tokens).

Based on the guideline "Prefer one short bullet per user-visible change: name the command/API/schema touched and the practical effect".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/sdk-node/CHANGELOG.md` at line 5, Split the changelog bullet into
two separate entries: one describing the `summary()` result changes, including
context-token normalization and context-size statistics, and another describing
the `hotspots()` ratio and minimum-context options with their defaults. Keep
each entry concise and focused on the affected API and practical user-visible
effect.

Source: Coding guidelines

crates/relayburn-sdk/src/query_verbs/tests.rs (1)

1130-1137: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the vacuous negative assertion with a positive detail check.

context_output_ratio_finding in crates/relayburn-sdk/src/analyze/findings.rs builds detail from a fixed format string that lists context tokens, output tokens, the threshold, and the minimum context tokens. The substring "169/6,348" belongs to the changelog calibration note, not to that format string. The assertion at Line 1134 therefore always passes and guards nothing. Assert the real contract instead.

♻️ Proposed test assertion change
 match result {
HotspotsResult::Findings { findings, .. } => {
assert_eq!(findings.len(), 1);
assert_eq!(findings[0].session_id, "incident-382");
- assert!(!findings[0].detail.contains("169/6,348"));+ assert_eq!(findings[0].kind, "context-output-ratio");+ assert!(findings[0].detail.contains("1146000 context tokens"));+ assert!(findings[0].detail.contains("3000 generated output tokens"));
}
other => panic!("expected findings, got {other:?}"),
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/relayburn-sdk/src/query_verbs/tests.rs` around lines 1130 - 1137,
Update the assertion in the HotspotsResult::Findings branch of the test to
positively verify the detail format produced by context_output_ratio_finding,
including the expected context tokens, output tokens, threshold, and minimum
context tokens. Remove the vacuous "169/6,348" negative check and assert the
actual contract represented by the fixed detail format string.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@CHANGELOG.md`:
- Line 8: Update the changelog bullet for `hotspots --findings` by removing the
private calibration statistics and related corpus wording, including the counts
and percentages. Retain the configurable ratio and context-token floor, the
default ratio, the shipped selection impact if applicable, and the caveat that
this is an inspection signal rather than a length-normalized anomaly score.
In `@packages/sdk-node/src/index.d.ts`:
- Around line 315-318: Update the contextOutputMinTokens declaration in the SDK
options type to accept only bigint, matching the native
context_output_min_tokens: Option<BigInt> interface and ensuring values reach
bigint_to_u64 without number conversion failures.
---
Nitpick comments:
In `@crates/relayburn-sdk/src/query_verbs/tests.rs`:
- Around line 1130-1137: Update the assertion in the HotspotsResult::Findings
branch of the test to positively verify the detail format produced by
context_output_ratio_finding, including the expected context tokens, output
tokens, threshold, and minimum context tokens. Remove the vacuous "169/6,348"
negative check and assert the actual contract represented by the fixed detail
format string.
In `@packages/sdk-node/CHANGELOG.md`:
- Line 5: Split the changelog bullet into two separate entries: one describing
the `summary()` result changes, including context-token normalization and
context-size statistics, and another describing the `hotspots()` ratio and
minimum-context options with their defaults. Keep each entry concise and focused
on the affected API and practical user-visible effect.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77842559-7eb6-43d4-9941-168fb16c03a4

📥 Commits

Reviewing files that changed from the base of the PR and between 962b2b7 and cca7987.

📒 Files selected for processing (16)
  • CHANGELOG.md
  • crates/relayburn-cli/src/commands/hotspots/mod.rs
  • crates/relayburn-cli/src/commands/summary/human.rs
  • crates/relayburn-cli/src/commands/summary/json.rs
  • crates/relayburn-cli/src/commands/summary/mod.rs
  • crates/relayburn-sdk-node/src/lib.rs
  • crates/relayburn-sdk/src/analyze.rs
  • crates/relayburn-sdk/src/analyze/cost.rs
  • crates/relayburn-sdk/src/analyze/findings.rs
  • crates/relayburn-sdk/src/query_verbs/context_efficiency.rs
  • crates/relayburn-sdk/src/query_verbs/hotspots.rs
  • crates/relayburn-sdk/src/query_verbs/mod.rs
  • crates/relayburn-sdk/src/query_verbs/summary/mod.rs
  • crates/relayburn-sdk/src/query_verbs/tests.rs
  • packages/sdk-node/CHANGELOG.md
  • packages/sdk-node/src/index.d.ts

Comment threadCHANGELOG.md Outdated
Comment threadpackages/sdk-node/src/index.d.ts

@devin-ai-integrationdevin-ai-integrationBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@willwashburn

Copy link
Copy Markdown
MemberAuthor

Addressed both review-summary nitpicks in 9bab8ae: the Node changelog now has one bullet per API change, and the incident fixture now positively asserts finding kind plus context/output/threshold/floor detail instead of a vacuous negative calibration-string check. Full workspace tests and strict clippy pass.

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:cca7987a14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment threadpackages/sdk-node/src/index.d.ts Outdated

@cubic-dev-aicubic-dev-aiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment threadcrates/relayburn-sdk-node/src/lib.rs
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-sdk/src/analyze/findings.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/summary/mod.rs
Comment threadcrates/relayburn-cli/src/commands/summary/human.rs Outdated
Comment threadcrates/relayburn-sdk/src/query_verbs/context_efficiency.rs
willwashburnand others added 2 commits August 3, 2026 04:49
Resolves the merged-main overlap: unioned findings exports and hotspots
tests, pricing_status on the ratio finding, context-output options in the
MCP hotspots wrapper, and golden snapshots regenerated now that the golden
suite is enforced in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Context-size / cache-efficiency is not a first-class metric

1 participant

@willwashburn