Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(pydantic_ai): strip unnecessary wrapper tokens by colinbennettbrain · Pull Request #321 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic_ai): strip unnecessary wrapper tokens - #321

Merged
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens
Apr 20, 2026
Merged

fix(pydantic_ai): strip unnecessary wrapper tokens#321
Abhijeet Prasad (AbhiPrasad) merged 5 commits into
mainfrom
colin/pydantic-strip-wrapper-tokens

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

retyped wrapper spans (agent_run, model_request, streaming wrappers) from LLM to TASK, but they kept logging the same prompt_tokens/completion_tokens/tokens as their nested leaf chat <model> span. The server derives estimated_cost per-span from tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace totals with coalesce_add over every non-scorer span regardless of type (summary.rs accumulate_metrics), and sums experiment-level token/cost over all non-scorer spans without filtering on span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to TASK therefore did not stop double-counting on any of those three axes.

Route every wrapper log site through a new _wrapper_span_metrics helper that emits only {start, end, duration, optional time_to_first_token}. Leaf chat <model> spans (from _wrap_concrete_model_class and _DirectStreamWrapper when span_type=LLM) keep full _extract_response_metrics. _DirectStreamWrapper now branches on span_type since it serves as both leaf and wrapper. Delete now-dead _extract_usage_metrics and _extract_stream_usage_metrics.

Flip existing cassette-backed assertions (test_agent_run_async, test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools, test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens / tokens / prompt_cached_tokens are absent from wrapper spans and present only on the leaf. No cassette re-recording needed -- the change is purely in post-processing.

colinbennettbrainand others added 5 commits April 16, 2026 22:20
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Stripping wrapper-span metrics in 6a6cb76 deleted `_extract_usage_metrics`,
which had been the only call site that read `usage.input_audio_tokens`,
`usage.output_audio_tokens`, and dict-style `usage.details`. Leaf chat spans
went through `_extract_response_metrics`, which (a) never knew about the audio
fields and (b) reached for `usage.details.reasoning_tokens` as an attribute.
pydantic_ai's `RequestUsage.details` is `dict[str, int]` (see
pydantic_ai/usage.py and `models/openai.py:_map_usage` populating
`details['reasoning_tokens']` from `output_tokens_details` plus spreading
`completion_tokens_details.model_dump()`). `hasattr(dict, "reasoning_tokens")`
is `False`, so the existing branch silently dropped reasoning tokens on every
leaf span -- including for o1/o3.
`_extract_response_metrics` now reads:
- prompt_audio_tokens <- usage.input_audio_tokens
- completion_audio_tokens <- usage.output_audio_tokens
- completion_reasoning_tokens <- details["reasoning_tokens"]
- prompt_cached_tokens <- details["cached_tokens"] (overrides cache_read_tokens
when both present, matching the pre-strip helper)
Adds a synthesized-RequestUsage unit test covering all of the above. Rewrites
`test_reasoning_tokens_extraction` to use `RequestUsage` instead of `MagicMock`
-- the mock satisfied attribute-style `hasattr` lookups and was the reason the
broken `details.reasoning_tokens` branch passed CI while never firing in
production. Adds wrapper-no-tokens regressions to test_agent_run_stream_events,
test_direct_model_request_stream, and test_direct_model_request_stream_sync,
mirroring the assertions added to non-stream agent_run tests in 6a6cb76.
pylint's astroid doesn't narrow Optional via 'assert x is not None',
so subscripting tripped unsubscriptable-object. Assign through a
typed local to make the type explicit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title pydantic strip wrapper tokensfix(pydantic_ai): unnecessary strip wrapper tokensApr 20, 2026
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title fix(pydantic_ai): unnecessary strip wrapper tokensfix(pydantic_ai): strip unnecessary wrapper tokensApr 20, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 2345ce9 into mainApr 20, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin/pydantic-strip-wrapper-tokens branch April 20, 2026 22:53
Abhijeet Prasad (AbhiPrasad) added a commit that referenced this pull request Apr 27, 2026
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad