Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(pydantic): retype wrapper spans as task to stop cost double-counting by colinbennettbrain · Pull Request #312 · braintrustdata/braintrust-sdk-python · GitHub
Skip to content

fix(pydantic): retype wrapper spans as task to stop cost double-counting - #312

Merged
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai
Apr 16, 2026
Merged

fix(pydantic): retype wrapper spans as task to stop cost double-counting#312
Abhijeet Prasad (AbhiPrasad) merged 1 commit into
mainfrom
colin-pydantic_ai

Conversation

@colinbennettbrain

Copy link
Copy Markdown
Contributor

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*, agent_to_cli_sync, model_request*) were tagged type=llm and logged the same usage metrics as their nested leaf chat <model> span. Experiment aggregations that sum metrics across type=llm spans therefore counted a single provider call twice (wrapper + leaf), inflating reported tokens and cost ~2x for a single-turn agent and more for multi-turn runs.

Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf chat <model> emitted by _wrap_concrete_model_class stays LLM. _DirectStreamWrapper is used both as wrapper (direct.model_request_stream) and as leaf (Model.request_stream), so it gains a span_type parameter defaulting to LLM; wrapper callers pass TASK explicitly.

Test coverage: flip existing wrapper-type assertions to TASK (leaf chat_span assertions stay LLM) and extend the cassette-backed test_agent_run_async with a regression check that exactly one type=llm span exists and that prompt/completion tokens summed across llm spans equal the leaf's values.

Agent wrapper spans (agent_run, agent_run_sync, agent_run_stream*,
agent_to_cli_sync, model_request*) were tagged type=llm and logged the
same usage metrics as their nested leaf `chat <model>` span. Experiment
aggregations that sum metrics across type=llm spans therefore counted a
single provider call twice (wrapper + leaf), inflating reported tokens
and cost ~2x for a single-turn agent and more for multi-turn runs.
Retag every wrapper span as SpanTypeAttribute.TASK; only the leaf
`chat <model>` emitted by _wrap_concrete_model_class stays LLM.
_DirectStreamWrapper is used both as wrapper (direct.model_request_stream)
and as leaf (Model.request_stream), so it gains a span_type parameter
defaulting to LLM; wrapper callers pass TASK explicitly.
Test coverage: flip existing wrapper-type assertions to TASK (leaf
chat_span assertions stay LLM) and extend the cassette-backed
test_agent_run_async with a regression check that exactly one type=llm
span exists and that prompt/completion tokens summed across llm spans
equal the leaf's values.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@AbhiPrasadAbhijeet Prasad (AbhiPrasad) changed the title retype wrapper spans as task to stop cost double-countingfix(pydantic): retype wrapper spans as task to stop cost double-countingApr 16, 2026
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit 31d7c07 into mainApr 16, 2026
84 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the colin-pydantic_ai branch April 16, 2026 19:10
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
colinbennettbrain added a commit that referenced this pull request Apr 16, 2026
#312 retyped wrapper spans (agent_run, model_request, streaming
wrappers) from LLM to TASK, but they kept logging the same
prompt_tokens/completion_tokens/tokens as their nested leaf
`chat <model>` span. The server derives estimated_cost per-span from
tokens+metadata.model (brainstore estimated_cost.rs), rolls up trace
totals with `coalesce_add` over every non-scorer span regardless of
type (summary.rs accumulate_metrics), and sums experiment-level
token/cost over all non-scorer spans without filtering on
span_type='llm' (summary.ts experimentScanSpanSummary). Retyping to
TASK therefore did not stop double-counting on any of those three
axes.
Route every wrapper log site through a new `_wrapper_span_metrics`
helper that emits only {start, end, duration, optional
time_to_first_token}. Leaf `chat <model>` spans (from
_wrap_concrete_model_class and _DirectStreamWrapper when
span_type=LLM) keep full _extract_response_metrics.
`_DirectStreamWrapper` now branches on span_type since it serves as
both leaf and wrapper. Delete now-dead `_extract_usage_metrics` and
`_extract_stream_usage_metrics`.
Flip existing cassette-backed assertions (test_agent_run_async,
test_agent_run_sync, test_agent_run_stream_async, test_agent_with_tools,
test_agent_run_stream_sync) to assert prompt_tokens / completion_tokens
/ tokens / prompt_cached_tokens are absent from wrapper spans and
present only on the leaf. No cassette re-recording needed -- the
change is purely in post-processing.
@AbhiPrasad

Copy link
Copy Markdown
Member

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@colinbennettbrain@AbhiPrasad