perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

perf: optimize tracing hot paths (5x faster user thread) - #101

Closed
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing
Closed

perf: optimize tracing hot paths (5x faster user thread)#101
Matt Perpick (clutchski) wants to merge 6 commits into
mainfrom
perf/optimize-tracing

Conversation

@clutchski

Copy link
Copy Markdown
Contributor

Summary

  • User-thread tracing overhead reduced 5.1x (884us -> 174us per request)
  • Total CPU cost reduced 4.2x (957us -> 226us per request)
  • All 377 core tests pass, pylint clean

End-to-end benchmark (5000 requests, root + child span + logs each)

 MAIN OPTIMIZED SPEEDUP
User thread: 4421.9 ms 870.1 ms 5.1x
Flush: 362.7 ms 257.2 ms 1.4x
Total CPU: 4784.6 ms 1127.3 ms 4.2x

Micro-benchmarks

 MAIN OPTIMIZED SPEEDUP
start_span (medium) 282.5 us 51.5 us 5.5x
start_span (large) 1267.1 us 103.2 us 12.3x
root + child + end 740.0 us 221.4 us 3.3x
span.log (medium) 186.3 us 13.6 us 13.7x
bt_safe_deep_copy (medium) 114.1 us 3.7 us 30.8x
bt_safe_deep_copy (large) 1110.9 us 29.2 us 38.0x

Key optimizations

  1. _to_bt_safe primitive fast-path: Check int/str/float/bool/None via type(v) is identity before expensive isinstance checks against abstract classes and Pydantic model_dump with warnings suppression
  2. _deep_copy_object rewrite: Inline primitive checks, use type(v) is dict/list instead of isinstance(v, Mapping), separate fast branches
  3. bt_safe_deep_copy orjson fast-path: JSON roundtrip via orjson for common all-JSON data (30-38x faster), falls back to tree walk for rich objects
  4. Skip deep copy for internal-only data: end() and set_attributes() don't reference user objects -- skip the copy
  5. split_logging_data: Avoid unnecessary merge_dicts when only event or internal_data present
  6. _strip_nones: Fast-path when no None values, use type(d) is dict
  7. merge_dicts: Inline simple path, avoid tuple path-tracking overhead
  8. Caching: Cache _get_exporter() result and object_id_fields per span
  9. itertools.count() instead of lock+global counter for exec_counter
  10. Lazy get_caller_location(): Skip stack walk when span name is provided

Test plan

  • make test-core -- 377 passed
  • make pylint -- clean
  • Benchmark on main vs branch confirms improvements

Generated with Claude Code

Matt Perpickand others added 4 commits March 19, 2026 20:44
Add end-to-end benchmark (bench_e2e.py) and detailed profiling analysis
(PERF_IDEAS.md) identifying 12 optimization opportunities in the
tracing hot paths.
Baseline: 967 us/req user thread, 39 us/item flush.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_to_bt_safe: check primitives (int/str/float/bool/None) first via
type identity before expensive isinstance checks against abstract
classes and Pydantic model_dump with warnings suppression. Guard
Pydantic v2/v1 attempts with hasattr() so only actual models pay cost.
_deep_copy_object: inline primitive fast path to avoid calling
_to_bt_safe for every leaf value. Use type(v) is dict/list instead
of isinstance(v, Mapping) for the common container types.
E2e benchmark (5000 reqs): 904 -> 264 us/req (3.4x faster)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Skip bt_safe_deep_copy when log_internal has no user event data
(e.g. end() and set_attributes()). Internal data only contains
primitives that don't reference user objects.
- Only call get_caller_location() when span name is not provided,
avoiding the stack walk in the common case.
E2e benchmark (5000 reqs): 264 -> 232 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- _strip_nones: fast-path when no None values, skip unnecessary copies.
- split_logging_data: avoid merge_dicts when only one side has data.
- _validate_and_sanitize: early return for empty event.
- merge_dicts: inline fast path avoiding tuple path tracking.
E2e benchmark (5000 reqs): 232 -> 217 us/req
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matt Perpickand others added 2 commits March 19, 2026 20:53
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type identity checks (type(v) is str) don't match str subclasses
like SpanTypeAttribute(str, Enum). Add isinstance fallback so these
are preserved rather than being converted to plain strings.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@AbhiPrasad

Copy link
Copy Markdown
Member

The py/src/braintrust/bt_json.py and utils changes seem correct to me, but I'm less sure about the logger changes from the pathological case. Like for example I think caching _get_exporter actually doesn't do anything. Shall we split this PR up?

We can add a proper benchmarking workflow as well, it would be interesting to hook up something like https://codspeed.io/

@AbhiPrasad

Copy link
Copy Markdown
Member

I merged some of these improvements in #139. Will keep working on it. Closing for now

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@clutchski@AbhiPrasad