Add a repeatable performance benchmark harness for tracing hot paths #134

Description

Add a repeatable performance benchmark harness for tracing hot paths

Summary

We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

  • validate that a performance-oriented refactor helps the intended hot path
  • separate large wins from noise
  • compare branch results against main
  • catch accidental regressions in shared helpers like merge_dicts()
  • evaluate behavior with and without optional perf dependencies like orjson

This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

Goals

  • Add repeatable microbenchmarks for tracing hot-path functions
  • Add one realistic end-to-end tracing benchmark
  • Make it easy to compare main vs a feature branch locally
  • Cover both minimal installs and the performance extra
  • Keep performance checks separate from normal functional test runs

Non-goals

  • Do not block CI on strict timing thresholds initially
  • Do not turn benchmark timings into flaky pytest assertions
  • Do not try to benchmark every SDK area at once

Proposed tooling

Use pyperf as the primary benchmark runner.

Rationale:

  • much better statistical discipline than ad hoc time.perf_counter() loops
  • supports warmups, calibration, repetition, and JSON output
  • gives us a clean compare_to workflow for branch vs branch
  • better fit for stable microbenchmarks than plain pytest timing

Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

Proposed layout

Under py/:

benchmarks/
README.md
conftest.py
cases/
bench_bt_json.py
bench_logger.py
bench_tracing_e2e.py
fixtures.py

Notes:

  • bench_bt_json.py should focus on serialization and deep-copy hot paths
  • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
  • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
  • fixtures.py should centralize representative payload builders so cases stay consistent

What to benchmark

1. bt_json microbenchmarks

Cover:

  • _to_bt_safe() on primitives
  • _to_bt_safe() on str/int subclasses and enum-backed strings
  • _to_bt_safe() on dataclasses
  • _to_bt_safe() on pydantic-like objects
  • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
  • bt_safe_deep_copy() on payloads with circular references
  • bt_safe_deep_copy() on payloads with non-string dict keys

Why:

  • these are the core hot paths in the PR under discussion
  • they are easy to measure deterministically in isolation

2. logger microbenchmarks

Cover:

  • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
  • _strip_nones() on shallow and nested dicts, with and without None
  • split_logging_data() for:
    • event only
    • internal_data only
    • both present
    • BraintrustStream present
  • SpanImpl creation with explicit name
  • SpanImpl creation without name
  • log_internal() with user event payloads
  • log_internal() with internal-only payloads like end() / set_attributes()

Why:

  • this isolates the logger-local optimizations from broader utility changes
  • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

3. End-to-end tracing benchmark

Cover one realistic scenario:

  • create root span
  • add representative input / metadata
  • create child span
  • log output / metadata / metrics
  • end child and root spans
  • exercise background logger record construction without external I/O

Parameters:

  • run enough iterations to reduce noise
  • keep payload shapes fixed and version-controlled
  • report separate scenarios for:
    • medium payload
    • large payload
    • internal-only updates

Why:

  • microbenchmarks show where a win comes from
  • the e2e benchmark shows whether the win is real in the path users care about

Benchmark environments

Run each benchmark suite in at least two environments:

A. Minimal environment

  • install the package plus benchmark deps
  • do not install optional perf extras

Purpose:

  • captures baseline behavior for default installs

B. Performance environment

  • install .[performance]

Purpose:

  • measures the effect of optional orjson
  • ensures future performance work does not accidentally optimize only one environment

If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

Nox integration

Add dedicated nox sessions instead of folding this into test_core.

Suggested sessions:

  • perf_bt_json
  • perf_logger
  • perf_e2e
  • optional umbrella session: perf

Behavior:

  • install pyperf
  • install the package from source
  • optionally install .[performance] in dedicated sessions or via a flag
  • run benchmark modules directly
  • emit JSON results to a temp or ignored output path

Example local workflows:

cd py
nox -s perf_bt_json
nox -s perf_logger
nox -s perf_e2e

Comparison workflow:

cd py
nox -s perf_bt_json -- --output /tmp/main-bt-json.json
nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

Local-first workflow

The initial implementation should optimize for local developer use only.

That means:

  • no CI integration yet
  • no nightly jobs
  • no hard performance thresholds

Instead, the benchmark harness should make it easy to:

  • establish a baseline on main
  • run the same benchmark on a feature branch
  • compare pyperf JSON outputs locally

The command surface should stay very small and obvious:

cd py
nox -s perf_bt_json
nox -s perf_logger
nox -s perf_e2e

And branch comparison should be a first-class local workflow:

cd py
nox -s perf_bt_json -- --output /tmp/main.json
nox -s perf_bt_json -- --output /tmp/branch.json
pyperf compare_to /tmp/main.json /tmp/branch.json

Reporting expectations for performance PRs

Any PR that claims performance improvement should include:

  • benchmark command(s) used
  • environment details:
    • Python version
    • whether .[performance] was installed
  • before/after results for the affected cases
  • note on variance if results are noisy

Prefer reporting per-benchmark improvements rather than a single blended headline number.

Rollout plan

Phase 1

  • add benchmark directory structure
  • add shared benchmark fixtures
  • add bt_json microbenchmarks

Phase 2

  • add logger microbenchmarks
  • add end-to-end tracing benchmark

Phase 3

  • add dedicated nox sessions
  • document local benchmark workflow in py/README.md or benchmark README

Acceptance criteria

  • There is a documented benchmark workflow on main
  • We can benchmark bt_json hot paths in isolation
  • We can benchmark logger hot paths in isolation
  • We can run one realistic tracing end-to-end benchmark
  • We can compare JSON results across branches with pyperf compare_to
  • Normal functional CI remains separate from benchmark execution

Why this should happen before PR #101-style perf work

PR #101 combines several categories of optimizations:

  • serialization dispatch changes
  • deep-copy specialization
  • logger fast paths
  • shared helper changes

Without a benchmark harness, it is too easy to:

  • merge broad refactors based on one aggregate number
  • overestimate wins from noisy runs
  • miss regressions in shared helpers
  • lose the ability to review each incremental optimization on its own merit

This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      Add a repeatable performance benchmark harness for tracing hot paths #134

      Description

      Add a repeatable performance benchmark harness for tracing hot paths

      Summary

      We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

      Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

      • validate that a performance-oriented refactor helps the intended hot path
      • separate large wins from noise
      • compare branch results against main
      • catch accidental regressions in shared helpers like merge_dicts()
      • evaluate behavior with and without optional perf dependencies like orjson

      This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

      Goals

      • Add repeatable microbenchmarks for tracing hot-path functions
      • Add one realistic end-to-end tracing benchmark
      • Make it easy to compare main vs a feature branch locally
      • Cover both minimal installs and the performance extra
      • Keep performance checks separate from normal functional test runs

      Non-goals

      • Do not block CI on strict timing thresholds initially
      • Do not turn benchmark timings into flaky pytest assertions
      • Do not try to benchmark every SDK area at once

      Proposed tooling

      Use pyperf as the primary benchmark runner.

      Rationale:

      • much better statistical discipline than ad hoc time.perf_counter() loops
      • supports warmups, calibration, repetition, and JSON output
      • gives us a clean compare_to workflow for branch vs branch
      • better fit for stable microbenchmarks than plain pytest timing

      Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

      Proposed layout

      Under py/:

      benchmarks/
      README.md
      conftest.py
      cases/
      bench_bt_json.py
      bench_logger.py
      bench_tracing_e2e.py
      fixtures.py
      

      Notes:

      • bench_bt_json.py should focus on serialization and deep-copy hot paths
      • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
      • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
      • fixtures.py should centralize representative payload builders so cases stay consistent

      What to benchmark

      1. bt_json microbenchmarks

      Cover:

      • _to_bt_safe() on primitives
      • _to_bt_safe() on str/int subclasses and enum-backed strings
      • _to_bt_safe() on dataclasses
      • _to_bt_safe() on pydantic-like objects
      • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
      • bt_safe_deep_copy() on payloads with circular references
      • bt_safe_deep_copy() on payloads with non-string dict keys

      Why:

      • these are the core hot paths in the PR under discussion
      • they are easy to measure deterministically in isolation

      2. logger microbenchmarks

      Cover:

      • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
      • _strip_nones() on shallow and nested dicts, with and without None
      • split_logging_data() for:
        • event only
        • internal_data only
        • both present
        • BraintrustStream present
      • SpanImpl creation with explicit name
      • SpanImpl creation without name
      • log_internal() with user event payloads
      • log_internal() with internal-only payloads like end() / set_attributes()

      Why:

      • this isolates the logger-local optimizations from broader utility changes
      • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

      3. End-to-end tracing benchmark

      Cover one realistic scenario:

      • create root span
      • add representative input / metadata
      • create child span
      • log output / metadata / metrics
      • end child and root spans
      • exercise background logger record construction without external I/O

      Parameters:

      • run enough iterations to reduce noise
      • keep payload shapes fixed and version-controlled
      • report separate scenarios for:
        • medium payload
        • large payload
        • internal-only updates

      Why:

      • microbenchmarks show where a win comes from
      • the e2e benchmark shows whether the win is real in the path users care about

      Benchmark environments

      Run each benchmark suite in at least two environments:

      A. Minimal environment

      • install the package plus benchmark deps
      • do not install optional perf extras

      Purpose:

      • captures baseline behavior for default installs

      B. Performance environment

      • install .[performance]

      Purpose:

      • measures the effect of optional orjson
      • ensures future performance work does not accidentally optimize only one environment

      If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

      Nox integration

      Add dedicated nox sessions instead of folding this into test_core.

      Suggested sessions:

      • perf_bt_json
      • perf_logger
      • perf_e2e
      • optional umbrella session: perf

      Behavior:

      • install pyperf
      • install the package from source
      • optionally install .[performance] in dedicated sessions or via a flag
      • run benchmark modules directly
      • emit JSON results to a temp or ignored output path

      Example local workflows:

      cd py
      nox -s perf_bt_json
      nox -s perf_logger
      nox -s perf_e2e

      Comparison workflow:

      cd py
      nox -s perf_bt_json -- --output /tmp/main-bt-json.json
      nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
      pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

      We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

      Local-first workflow

      The initial implementation should optimize for local developer use only.

      That means:

      • no CI integration yet
      • no nightly jobs
      • no hard performance thresholds

      Instead, the benchmark harness should make it easy to:

      • establish a baseline on main
      • run the same benchmark on a feature branch
      • compare pyperf JSON outputs locally

      The command surface should stay very small and obvious:

      cd py
      nox -s perf_bt_json
      nox -s perf_logger
      nox -s perf_e2e

      And branch comparison should be a first-class local workflow:

      cd py
      nox -s perf_bt_json -- --output /tmp/main.json
      nox -s perf_bt_json -- --output /tmp/branch.json
      pyperf compare_to /tmp/main.json /tmp/branch.json

      Reporting expectations for performance PRs

      Any PR that claims performance improvement should include:

      • benchmark command(s) used
      • environment details:
        • Python version
        • whether .[performance] was installed
      • before/after results for the affected cases
      • note on variance if results are noisy

      Prefer reporting per-benchmark improvements rather than a single blended headline number.

      Rollout plan

      Phase 1

      • add benchmark directory structure
      • add shared benchmark fixtures
      • add bt_json microbenchmarks

      Phase 2

      • add logger microbenchmarks
      • add end-to-end tracing benchmark

      Phase 3

      • add dedicated nox sessions
      • document local benchmark workflow in py/README.md or benchmark README

      Acceptance criteria

      • There is a documented benchmark workflow on main
      • We can benchmark bt_json hot paths in isolation
      • We can benchmark logger hot paths in isolation
      • We can run one realistic tracing end-to-end benchmark
      • We can compare JSON results across branches with pyperf compare_to
      • Normal functional CI remains separate from benchmark execution

      Why this should happen before PR #101-style perf work

      PR #101 combines several categories of optimizations:

      • serialization dispatch changes
      • deep-copy specialization
      • logger fast paths
      • shared helper changes

      Without a benchmark harness, it is too easy to:

      • merge broad refactors based on one aggregate number
      • overestimate wins from noisy runs
      • miss regressions in shared helpers
      • lose the ability to review each incremental optimization on its own merit

      This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          Add a repeatable performance benchmark harness for tracing hot paths #134

          Description

          Add a repeatable performance benchmark harness for tracing hot paths

          Summary

          We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

          Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

          • validate that a performance-oriented refactor helps the intended hot path
          • separate large wins from noise
          • compare branch results against main
          • catch accidental regressions in shared helpers like merge_dicts()
          • evaluate behavior with and without optional perf dependencies like orjson

          This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

          Goals

          • Add repeatable microbenchmarks for tracing hot-path functions
          • Add one realistic end-to-end tracing benchmark
          • Make it easy to compare main vs a feature branch locally
          • Cover both minimal installs and the performance extra
          • Keep performance checks separate from normal functional test runs

          Non-goals

          • Do not block CI on strict timing thresholds initially
          • Do not turn benchmark timings into flaky pytest assertions
          • Do not try to benchmark every SDK area at once

          Proposed tooling

          Use pyperf as the primary benchmark runner.

          Rationale:

          • much better statistical discipline than ad hoc time.perf_counter() loops
          • supports warmups, calibration, repetition, and JSON output
          • gives us a clean compare_to workflow for branch vs branch
          • better fit for stable microbenchmarks than plain pytest timing

          Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

          Proposed layout

          Under py/:

          benchmarks/
          README.md
          conftest.py
          cases/
          bench_bt_json.py
          bench_logger.py
          bench_tracing_e2e.py
          fixtures.py
          

          Notes:

          • bench_bt_json.py should focus on serialization and deep-copy hot paths
          • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
          • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
          • fixtures.py should centralize representative payload builders so cases stay consistent

          What to benchmark

          1. bt_json microbenchmarks

          Cover:

          • _to_bt_safe() on primitives
          • _to_bt_safe() on str/int subclasses and enum-backed strings
          • _to_bt_safe() on dataclasses
          • _to_bt_safe() on pydantic-like objects
          • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
          • bt_safe_deep_copy() on payloads with circular references
          • bt_safe_deep_copy() on payloads with non-string dict keys

          Why:

          • these are the core hot paths in the PR under discussion
          • they are easy to measure deterministically in isolation

          2. logger microbenchmarks

          Cover:

          • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
          • _strip_nones() on shallow and nested dicts, with and without None
          • split_logging_data() for:
            • event only
            • internal_data only
            • both present
            • BraintrustStream present
          • SpanImpl creation with explicit name
          • SpanImpl creation without name
          • log_internal() with user event payloads
          • log_internal() with internal-only payloads like end() / set_attributes()

          Why:

          • this isolates the logger-local optimizations from broader utility changes
          • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

          3. End-to-end tracing benchmark

          Cover one realistic scenario:

          • create root span
          • add representative input / metadata
          • create child span
          • log output / metadata / metrics
          • end child and root spans
          • exercise background logger record construction without external I/O

          Parameters:

          • run enough iterations to reduce noise
          • keep payload shapes fixed and version-controlled
          • report separate scenarios for:
            • medium payload
            • large payload
            • internal-only updates

          Why:

          • microbenchmarks show where a win comes from
          • the e2e benchmark shows whether the win is real in the path users care about

          Benchmark environments

          Run each benchmark suite in at least two environments:

          A. Minimal environment

          • install the package plus benchmark deps
          • do not install optional perf extras

          Purpose:

          • captures baseline behavior for default installs

          B. Performance environment

          • install .[performance]

          Purpose:

          • measures the effect of optional orjson
          • ensures future performance work does not accidentally optimize only one environment

          If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

          Nox integration

          Add dedicated nox sessions instead of folding this into test_core.

          Suggested sessions:

          • perf_bt_json
          • perf_logger
          • perf_e2e
          • optional umbrella session: perf

          Behavior:

          • install pyperf
          • install the package from source
          • optionally install .[performance] in dedicated sessions or via a flag
          • run benchmark modules directly
          • emit JSON results to a temp or ignored output path

          Example local workflows:

          cd py
          nox -s perf_bt_json
          nox -s perf_logger
          nox -s perf_e2e

          Comparison workflow:

          cd py
          nox -s perf_bt_json -- --output /tmp/main-bt-json.json
          nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
          pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

          We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

          Local-first workflow

          The initial implementation should optimize for local developer use only.

          That means:

          • no CI integration yet
          • no nightly jobs
          • no hard performance thresholds

          Instead, the benchmark harness should make it easy to:

          • establish a baseline on main
          • run the same benchmark on a feature branch
          • compare pyperf JSON outputs locally

          The command surface should stay very small and obvious:

          cd py
          nox -s perf_bt_json
          nox -s perf_logger
          nox -s perf_e2e

          And branch comparison should be a first-class local workflow:

          cd py
          nox -s perf_bt_json -- --output /tmp/main.json
          nox -s perf_bt_json -- --output /tmp/branch.json
          pyperf compare_to /tmp/main.json /tmp/branch.json

          Reporting expectations for performance PRs

          Any PR that claims performance improvement should include:

          • benchmark command(s) used
          • environment details:
            • Python version
            • whether .[performance] was installed
          • before/after results for the affected cases
          • note on variance if results are noisy

          Prefer reporting per-benchmark improvements rather than a single blended headline number.

          Rollout plan

          Phase 1

          • add benchmark directory structure
          • add shared benchmark fixtures
          • add bt_json microbenchmarks

          Phase 2

          • add logger microbenchmarks
          • add end-to-end tracing benchmark

          Phase 3

          • add dedicated nox sessions
          • document local benchmark workflow in py/README.md or benchmark README

          Acceptance criteria

          • There is a documented benchmark workflow on main
          • We can benchmark bt_json hot paths in isolation
          • We can benchmark logger hot paths in isolation
          • We can run one realistic tracing end-to-end benchmark
          • We can compare JSON results across branches with pyperf compare_to
          • Normal functional CI remains separate from benchmark execution

          Why this should happen before PR #101-style perf work

          PR #101 combines several categories of optimizations:

          • serialization dispatch changes
          • deep-copy specialization
          • logger fast paths
          • shared helper changes

          Without a benchmark harness, it is too easy to:

          • merge broad refactors based on one aggregate number
          • overestimate wins from noisy runs
          • miss regressions in shared helpers
          • lose the ability to review each incremental optimization on its own merit

          This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Add a repeatable performance benchmark harness for tracing hot paths #134

              Description

              Add a repeatable performance benchmark harness for tracing hot paths

              Summary

              We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

              Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

              • validate that a performance-oriented refactor helps the intended hot path
              • separate large wins from noise
              • compare branch results against main
              • catch accidental regressions in shared helpers like merge_dicts()
              • evaluate behavior with and without optional perf dependencies like orjson

              This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

              Goals

              • Add repeatable microbenchmarks for tracing hot-path functions
              • Add one realistic end-to-end tracing benchmark
              • Make it easy to compare main vs a feature branch locally
              • Cover both minimal installs and the performance extra
              • Keep performance checks separate from normal functional test runs

              Non-goals

              • Do not block CI on strict timing thresholds initially
              • Do not turn benchmark timings into flaky pytest assertions
              • Do not try to benchmark every SDK area at once

              Proposed tooling

              Use pyperf as the primary benchmark runner.

              Rationale:

              • much better statistical discipline than ad hoc time.perf_counter() loops
              • supports warmups, calibration, repetition, and JSON output
              • gives us a clean compare_to workflow for branch vs branch
              • better fit for stable microbenchmarks than plain pytest timing

              Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

              Proposed layout

              Under py/:

              benchmarks/
              README.md
              conftest.py
              cases/
              bench_bt_json.py
              bench_logger.py
              bench_tracing_e2e.py
              fixtures.py
              

              Notes:

              • bench_bt_json.py should focus on serialization and deep-copy hot paths
              • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
              • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
              • fixtures.py should centralize representative payload builders so cases stay consistent

              What to benchmark

              1. bt_json microbenchmarks

              Cover:

              • _to_bt_safe() on primitives
              • _to_bt_safe() on str/int subclasses and enum-backed strings
              • _to_bt_safe() on dataclasses
              • _to_bt_safe() on pydantic-like objects
              • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
              • bt_safe_deep_copy() on payloads with circular references
              • bt_safe_deep_copy() on payloads with non-string dict keys

              Why:

              • these are the core hot paths in the PR under discussion
              • they are easy to measure deterministically in isolation

              2. logger microbenchmarks

              Cover:

              • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
              • _strip_nones() on shallow and nested dicts, with and without None
              • split_logging_data() for:
                • event only
                • internal_data only
                • both present
                • BraintrustStream present
              • SpanImpl creation with explicit name
              • SpanImpl creation without name
              • log_internal() with user event payloads
              • log_internal() with internal-only payloads like end() / set_attributes()

              Why:

              • this isolates the logger-local optimizations from broader utility changes
              • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

              3. End-to-end tracing benchmark

              Cover one realistic scenario:

              • create root span
              • add representative input / metadata
              • create child span
              • log output / metadata / metrics
              • end child and root spans
              • exercise background logger record construction without external I/O

              Parameters:

              • run enough iterations to reduce noise
              • keep payload shapes fixed and version-controlled
              • report separate scenarios for:
                • medium payload
                • large payload
                • internal-only updates

              Why:

              • microbenchmarks show where a win comes from
              • the e2e benchmark shows whether the win is real in the path users care about

              Benchmark environments

              Run each benchmark suite in at least two environments:

              A. Minimal environment

              • install the package plus benchmark deps
              • do not install optional perf extras

              Purpose:

              • captures baseline behavior for default installs

              B. Performance environment

              • install .[performance]

              Purpose:

              • measures the effect of optional orjson
              • ensures future performance work does not accidentally optimize only one environment

              If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

              Nox integration

              Add dedicated nox sessions instead of folding this into test_core.

              Suggested sessions:

              • perf_bt_json
              • perf_logger
              • perf_e2e
              • optional umbrella session: perf

              Behavior:

              • install pyperf
              • install the package from source
              • optionally install .[performance] in dedicated sessions or via a flag
              • run benchmark modules directly
              • emit JSON results to a temp or ignored output path

              Example local workflows:

              cd py
              nox -s perf_bt_json
              nox -s perf_logger
              nox -s perf_e2e

              Comparison workflow:

              cd py
              nox -s perf_bt_json -- --output /tmp/main-bt-json.json
              nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
              pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

              We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

              Local-first workflow

              The initial implementation should optimize for local developer use only.

              That means:

              • no CI integration yet
              • no nightly jobs
              • no hard performance thresholds

              Instead, the benchmark harness should make it easy to:

              • establish a baseline on main
              • run the same benchmark on a feature branch
              • compare pyperf JSON outputs locally

              The command surface should stay very small and obvious:

              cd py
              nox -s perf_bt_json
              nox -s perf_logger
              nox -s perf_e2e

              And branch comparison should be a first-class local workflow:

              cd py
              nox -s perf_bt_json -- --output /tmp/main.json
              nox -s perf_bt_json -- --output /tmp/branch.json
              pyperf compare_to /tmp/main.json /tmp/branch.json

              Reporting expectations for performance PRs

              Any PR that claims performance improvement should include:

              • benchmark command(s) used
              • environment details:
                • Python version
                • whether .[performance] was installed
              • before/after results for the affected cases
              • note on variance if results are noisy

              Prefer reporting per-benchmark improvements rather than a single blended headline number.

              Rollout plan

              Phase 1

              • add benchmark directory structure
              • add shared benchmark fixtures
              • add bt_json microbenchmarks

              Phase 2

              • add logger microbenchmarks
              • add end-to-end tracing benchmark

              Phase 3

              • add dedicated nox sessions
              • document local benchmark workflow in py/README.md or benchmark README

              Acceptance criteria

              • There is a documented benchmark workflow on main
              • We can benchmark bt_json hot paths in isolation
              • We can benchmark logger hot paths in isolation
              • We can run one realistic tracing end-to-end benchmark
              • We can compare JSON results across branches with pyperf compare_to
              • Normal functional CI remains separate from benchmark execution

              Why this should happen before PR #101-style perf work

              PR #101 combines several categories of optimizations:

              • serialization dispatch changes
              • deep-copy specialization
              • logger fast paths
              • shared helper changes

              Without a benchmark harness, it is too easy to:

              • merge broad refactors based on one aggregate number
              • overestimate wins from noisy runs
              • miss regressions in shared helpers
              • lose the ability to review each incremental optimization on its own merit

              This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  Add a repeatable performance benchmark harness for tracing hot paths #134

                  Description

                  Add a repeatable performance benchmark harness for tracing hot paths

                  Summary

                  We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

                  Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

                  • validate that a performance-oriented refactor helps the intended hot path
                  • separate large wins from noise
                  • compare branch results against main
                  • catch accidental regressions in shared helpers like merge_dicts()
                  • evaluate behavior with and without optional perf dependencies like orjson

                  This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

                  Goals

                  • Add repeatable microbenchmarks for tracing hot-path functions
                  • Add one realistic end-to-end tracing benchmark
                  • Make it easy to compare main vs a feature branch locally
                  • Cover both minimal installs and the performance extra
                  • Keep performance checks separate from normal functional test runs

                  Non-goals

                  • Do not block CI on strict timing thresholds initially
                  • Do not turn benchmark timings into flaky pytest assertions
                  • Do not try to benchmark every SDK area at once

                  Proposed tooling

                  Use pyperf as the primary benchmark runner.

                  Rationale:

                  • much better statistical discipline than ad hoc time.perf_counter() loops
                  • supports warmups, calibration, repetition, and JSON output
                  • gives us a clean compare_to workflow for branch vs branch
                  • better fit for stable microbenchmarks than plain pytest timing

                  Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

                  Proposed layout

                  Under py/:

                  benchmarks/
                  README.md
                  conftest.py
                  cases/
                  bench_bt_json.py
                  bench_logger.py
                  bench_tracing_e2e.py
                  fixtures.py
                  

                  Notes:

                  • bench_bt_json.py should focus on serialization and deep-copy hot paths
                  • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
                  • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
                  • fixtures.py should centralize representative payload builders so cases stay consistent

                  What to benchmark

                  1. bt_json microbenchmarks

                  Cover:

                  • _to_bt_safe() on primitives
                  • _to_bt_safe() on str/int subclasses and enum-backed strings
                  • _to_bt_safe() on dataclasses
                  • _to_bt_safe() on pydantic-like objects
                  • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
                  • bt_safe_deep_copy() on payloads with circular references
                  • bt_safe_deep_copy() on payloads with non-string dict keys

                  Why:

                  • these are the core hot paths in the PR under discussion
                  • they are easy to measure deterministically in isolation

                  2. logger microbenchmarks

                  Cover:

                  • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
                  • _strip_nones() on shallow and nested dicts, with and without None
                  • split_logging_data() for:
                    • event only
                    • internal_data only
                    • both present
                    • BraintrustStream present
                  • SpanImpl creation with explicit name
                  • SpanImpl creation without name
                  • log_internal() with user event payloads
                  • log_internal() with internal-only payloads like end() / set_attributes()

                  Why:

                  • this isolates the logger-local optimizations from broader utility changes
                  • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

                  3. End-to-end tracing benchmark

                  Cover one realistic scenario:

                  • create root span
                  • add representative input / metadata
                  • create child span
                  • log output / metadata / metrics
                  • end child and root spans
                  • exercise background logger record construction without external I/O

                  Parameters:

                  • run enough iterations to reduce noise
                  • keep payload shapes fixed and version-controlled
                  • report separate scenarios for:
                    • medium payload
                    • large payload
                    • internal-only updates

                  Why:

                  • microbenchmarks show where a win comes from
                  • the e2e benchmark shows whether the win is real in the path users care about

                  Benchmark environments

                  Run each benchmark suite in at least two environments:

                  A. Minimal environment

                  • install the package plus benchmark deps
                  • do not install optional perf extras

                  Purpose:

                  • captures baseline behavior for default installs

                  B. Performance environment

                  • install .[performance]

                  Purpose:

                  • measures the effect of optional orjson
                  • ensures future performance work does not accidentally optimize only one environment

                  If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

                  Nox integration

                  Add dedicated nox sessions instead of folding this into test_core.

                  Suggested sessions:

                  • perf_bt_json
                  • perf_logger
                  • perf_e2e
                  • optional umbrella session: perf

                  Behavior:

                  • install pyperf
                  • install the package from source
                  • optionally install .[performance] in dedicated sessions or via a flag
                  • run benchmark modules directly
                  • emit JSON results to a temp or ignored output path

                  Example local workflows:

                  cd py
                  nox -s perf_bt_json
                  nox -s perf_logger
                  nox -s perf_e2e

                  Comparison workflow:

                  cd py
                  nox -s perf_bt_json -- --output /tmp/main-bt-json.json
                  nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
                  pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

                  We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

                  Local-first workflow

                  The initial implementation should optimize for local developer use only.

                  That means:

                  • no CI integration yet
                  • no nightly jobs
                  • no hard performance thresholds

                  Instead, the benchmark harness should make it easy to:

                  • establish a baseline on main
                  • run the same benchmark on a feature branch
                  • compare pyperf JSON outputs locally

                  The command surface should stay very small and obvious:

                  cd py
                  nox -s perf_bt_json
                  nox -s perf_logger
                  nox -s perf_e2e

                  And branch comparison should be a first-class local workflow:

                  cd py
                  nox -s perf_bt_json -- --output /tmp/main.json
                  nox -s perf_bt_json -- --output /tmp/branch.json
                  pyperf compare_to /tmp/main.json /tmp/branch.json

                  Reporting expectations for performance PRs

                  Any PR that claims performance improvement should include:

                  • benchmark command(s) used
                  • environment details:
                    • Python version
                    • whether .[performance] was installed
                  • before/after results for the affected cases
                  • note on variance if results are noisy

                  Prefer reporting per-benchmark improvements rather than a single blended headline number.

                  Rollout plan

                  Phase 1

                  • add benchmark directory structure
                  • add shared benchmark fixtures
                  • add bt_json microbenchmarks

                  Phase 2

                  • add logger microbenchmarks
                  • add end-to-end tracing benchmark

                  Phase 3

                  • add dedicated nox sessions
                  • document local benchmark workflow in py/README.md or benchmark README

                  Acceptance criteria

                  • There is a documented benchmark workflow on main
                  • We can benchmark bt_json hot paths in isolation
                  • We can benchmark logger hot paths in isolation
                  • We can run one realistic tracing end-to-end benchmark
                  • We can compare JSON results across branches with pyperf compare_to
                  • Normal functional CI remains separate from benchmark execution

                  Why this should happen before PR #101-style perf work

                  PR #101 combines several categories of optimizations:

                  • serialization dispatch changes
                  • deep-copy specialization
                  • logger fast paths
                  • shared helper changes

                  Without a benchmark harness, it is too easy to:

                  • merge broad refactors based on one aggregate number
                  • overestimate wins from noisy runs
                  • miss regressions in shared helpers
                  • lose the ability to review each incremental optimization on its own merit

                  This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      Add a repeatable performance benchmark harness for tracing hot paths #134

                      Description

                      Add a repeatable performance benchmark harness for tracing hot paths

                      Summary

                      We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

                      Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

                      • validate that a performance-oriented refactor helps the intended hot path
                      • separate large wins from noise
                      • compare branch results against main
                      • catch accidental regressions in shared helpers like merge_dicts()
                      • evaluate behavior with and without optional perf dependencies like orjson

                      This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

                      Goals

                      • Add repeatable microbenchmarks for tracing hot-path functions
                      • Add one realistic end-to-end tracing benchmark
                      • Make it easy to compare main vs a feature branch locally
                      • Cover both minimal installs and the performance extra
                      • Keep performance checks separate from normal functional test runs

                      Non-goals

                      • Do not block CI on strict timing thresholds initially
                      • Do not turn benchmark timings into flaky pytest assertions
                      • Do not try to benchmark every SDK area at once

                      Proposed tooling

                      Use pyperf as the primary benchmark runner.

                      Rationale:

                      • much better statistical discipline than ad hoc time.perf_counter() loops
                      • supports warmups, calibration, repetition, and JSON output
                      • gives us a clean compare_to workflow for branch vs branch
                      • better fit for stable microbenchmarks than plain pytest timing

                      Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

                      Proposed layout

                      Under py/:

                      benchmarks/
                      README.md
                      conftest.py
                      cases/
                      bench_bt_json.py
                      bench_logger.py
                      bench_tracing_e2e.py
                      fixtures.py
                      

                      Notes:

                      • bench_bt_json.py should focus on serialization and deep-copy hot paths
                      • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
                      • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
                      • fixtures.py should centralize representative payload builders so cases stay consistent

                      What to benchmark

                      1. bt_json microbenchmarks

                      Cover:

                      • _to_bt_safe() on primitives
                      • _to_bt_safe() on str/int subclasses and enum-backed strings
                      • _to_bt_safe() on dataclasses
                      • _to_bt_safe() on pydantic-like objects
                      • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
                      • bt_safe_deep_copy() on payloads with circular references
                      • bt_safe_deep_copy() on payloads with non-string dict keys

                      Why:

                      • these are the core hot paths in the PR under discussion
                      • they are easy to measure deterministically in isolation

                      2. logger microbenchmarks

                      Cover:

                      • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
                      • _strip_nones() on shallow and nested dicts, with and without None
                      • split_logging_data() for:
                        • event only
                        • internal_data only
                        • both present
                        • BraintrustStream present
                      • SpanImpl creation with explicit name
                      • SpanImpl creation without name
                      • log_internal() with user event payloads
                      • log_internal() with internal-only payloads like end() / set_attributes()

                      Why:

                      • this isolates the logger-local optimizations from broader utility changes
                      • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

                      3. End-to-end tracing benchmark

                      Cover one realistic scenario:

                      • create root span
                      • add representative input / metadata
                      • create child span
                      • log output / metadata / metrics
                      • end child and root spans
                      • exercise background logger record construction without external I/O

                      Parameters:

                      • run enough iterations to reduce noise
                      • keep payload shapes fixed and version-controlled
                      • report separate scenarios for:
                        • medium payload
                        • large payload
                        • internal-only updates

                      Why:

                      • microbenchmarks show where a win comes from
                      • the e2e benchmark shows whether the win is real in the path users care about

                      Benchmark environments

                      Run each benchmark suite in at least two environments:

                      A. Minimal environment

                      • install the package plus benchmark deps
                      • do not install optional perf extras

                      Purpose:

                      • captures baseline behavior for default installs

                      B. Performance environment

                      • install .[performance]

                      Purpose:

                      • measures the effect of optional orjson
                      • ensures future performance work does not accidentally optimize only one environment

                      If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

                      Nox integration

                      Add dedicated nox sessions instead of folding this into test_core.

                      Suggested sessions:

                      • perf_bt_json
                      • perf_logger
                      • perf_e2e
                      • optional umbrella session: perf

                      Behavior:

                      • install pyperf
                      • install the package from source
                      • optionally install .[performance] in dedicated sessions or via a flag
                      • run benchmark modules directly
                      • emit JSON results to a temp or ignored output path

                      Example local workflows:

                      cd py
                      nox -s perf_bt_json
                      nox -s perf_logger
                      nox -s perf_e2e

                      Comparison workflow:

                      cd py
                      nox -s perf_bt_json -- --output /tmp/main-bt-json.json
                      nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
                      pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

                      We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

                      Local-first workflow

                      The initial implementation should optimize for local developer use only.

                      That means:

                      • no CI integration yet
                      • no nightly jobs
                      • no hard performance thresholds

                      Instead, the benchmark harness should make it easy to:

                      • establish a baseline on main
                      • run the same benchmark on a feature branch
                      • compare pyperf JSON outputs locally

                      The command surface should stay very small and obvious:

                      cd py
                      nox -s perf_bt_json
                      nox -s perf_logger
                      nox -s perf_e2e

                      And branch comparison should be a first-class local workflow:

                      cd py
                      nox -s perf_bt_json -- --output /tmp/main.json
                      nox -s perf_bt_json -- --output /tmp/branch.json
                      pyperf compare_to /tmp/main.json /tmp/branch.json

                      Reporting expectations for performance PRs

                      Any PR that claims performance improvement should include:

                      • benchmark command(s) used
                      • environment details:
                        • Python version
                        • whether .[performance] was installed
                      • before/after results for the affected cases
                      • note on variance if results are noisy

                      Prefer reporting per-benchmark improvements rather than a single blended headline number.

                      Rollout plan

                      Phase 1

                      • add benchmark directory structure
                      • add shared benchmark fixtures
                      • add bt_json microbenchmarks

                      Phase 2

                      • add logger microbenchmarks
                      • add end-to-end tracing benchmark

                      Phase 3

                      • add dedicated nox sessions
                      • document local benchmark workflow in py/README.md or benchmark README

                      Acceptance criteria

                      • There is a documented benchmark workflow on main
                      • We can benchmark bt_json hot paths in isolation
                      • We can benchmark logger hot paths in isolation
                      • We can run one realistic tracing end-to-end benchmark
                      • We can compare JSON results across branches with pyperf compare_to
                      • Normal functional CI remains separate from benchmark execution

                      Why this should happen before PR #101-style perf work

                      PR #101 combines several categories of optimizations:

                      • serialization dispatch changes
                      • deep-copy specialization
                      • logger fast paths
                      • shared helper changes

                      Without a benchmark harness, it is too easy to:

                      • merge broad refactors based on one aggregate number
                      • overestimate wins from noisy runs
                      • miss regressions in shared helpers
                      • lose the ability to review each incremental optimization on its own merit

                      This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          Add a repeatable performance benchmark harness for tracing hot paths #134

                          Description

                          Add a repeatable performance benchmark harness for tracing hot paths

                          Summary

                          We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

                          Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

                          • validate that a performance-oriented refactor helps the intended hot path
                          • separate large wins from noise
                          • compare branch results against main
                          • catch accidental regressions in shared helpers like merge_dicts()
                          • evaluate behavior with and without optional perf dependencies like orjson

                          This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

                          Goals

                          • Add repeatable microbenchmarks for tracing hot-path functions
                          • Add one realistic end-to-end tracing benchmark
                          • Make it easy to compare main vs a feature branch locally
                          • Cover both minimal installs and the performance extra
                          • Keep performance checks separate from normal functional test runs

                          Non-goals

                          • Do not block CI on strict timing thresholds initially
                          • Do not turn benchmark timings into flaky pytest assertions
                          • Do not try to benchmark every SDK area at once

                          Proposed tooling

                          Use pyperf as the primary benchmark runner.

                          Rationale:

                          • much better statistical discipline than ad hoc time.perf_counter() loops
                          • supports warmups, calibration, repetition, and JSON output
                          • gives us a clean compare_to workflow for branch vs branch
                          • better fit for stable microbenchmarks than plain pytest timing

                          Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

                          Proposed layout

                          Under py/:

                          benchmarks/
                          README.md
                          conftest.py
                          cases/
                          bench_bt_json.py
                          bench_logger.py
                          bench_tracing_e2e.py
                          fixtures.py
                          

                          Notes:

                          • bench_bt_json.py should focus on serialization and deep-copy hot paths
                          • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
                          • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
                          • fixtures.py should centralize representative payload builders so cases stay consistent

                          What to benchmark

                          1. bt_json microbenchmarks

                          Cover:

                          • _to_bt_safe() on primitives
                          • _to_bt_safe() on str/int subclasses and enum-backed strings
                          • _to_bt_safe() on dataclasses
                          • _to_bt_safe() on pydantic-like objects
                          • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
                          • bt_safe_deep_copy() on payloads with circular references
                          • bt_safe_deep_copy() on payloads with non-string dict keys

                          Why:

                          • these are the core hot paths in the PR under discussion
                          • they are easy to measure deterministically in isolation

                          2. logger microbenchmarks

                          Cover:

                          • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
                          • _strip_nones() on shallow and nested dicts, with and without None
                          • split_logging_data() for:
                            • event only
                            • internal_data only
                            • both present
                            • BraintrustStream present
                          • SpanImpl creation with explicit name
                          • SpanImpl creation without name
                          • log_internal() with user event payloads
                          • log_internal() with internal-only payloads like end() / set_attributes()

                          Why:

                          • this isolates the logger-local optimizations from broader utility changes
                          • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

                          3. End-to-end tracing benchmark

                          Cover one realistic scenario:

                          • create root span
                          • add representative input / metadata
                          • create child span
                          • log output / metadata / metrics
                          • end child and root spans
                          • exercise background logger record construction without external I/O

                          Parameters:

                          • run enough iterations to reduce noise
                          • keep payload shapes fixed and version-controlled
                          • report separate scenarios for:
                            • medium payload
                            • large payload
                            • internal-only updates

                          Why:

                          • microbenchmarks show where a win comes from
                          • the e2e benchmark shows whether the win is real in the path users care about

                          Benchmark environments

                          Run each benchmark suite in at least two environments:

                          A. Minimal environment

                          • install the package plus benchmark deps
                          • do not install optional perf extras

                          Purpose:

                          • captures baseline behavior for default installs

                          B. Performance environment

                          • install .[performance]

                          Purpose:

                          • measures the effect of optional orjson
                          • ensures future performance work does not accidentally optimize only one environment

                          If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

                          Nox integration

                          Add dedicated nox sessions instead of folding this into test_core.

                          Suggested sessions:

                          • perf_bt_json
                          • perf_logger
                          • perf_e2e
                          • optional umbrella session: perf

                          Behavior:

                          • install pyperf
                          • install the package from source
                          • optionally install .[performance] in dedicated sessions or via a flag
                          • run benchmark modules directly
                          • emit JSON results to a temp or ignored output path

                          Example local workflows:

                          cd py
                          nox -s perf_bt_json
                          nox -s perf_logger
                          nox -s perf_e2e

                          Comparison workflow:

                          cd py
                          nox -s perf_bt_json -- --output /tmp/main-bt-json.json
                          nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
                          pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

                          We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

                          Local-first workflow

                          The initial implementation should optimize for local developer use only.

                          That means:

                          • no CI integration yet
                          • no nightly jobs
                          • no hard performance thresholds

                          Instead, the benchmark harness should make it easy to:

                          • establish a baseline on main
                          • run the same benchmark on a feature branch
                          • compare pyperf JSON outputs locally

                          The command surface should stay very small and obvious:

                          cd py
                          nox -s perf_bt_json
                          nox -s perf_logger
                          nox -s perf_e2e

                          And branch comparison should be a first-class local workflow:

                          cd py
                          nox -s perf_bt_json -- --output /tmp/main.json
                          nox -s perf_bt_json -- --output /tmp/branch.json
                          pyperf compare_to /tmp/main.json /tmp/branch.json

                          Reporting expectations for performance PRs

                          Any PR that claims performance improvement should include:

                          • benchmark command(s) used
                          • environment details:
                            • Python version
                            • whether .[performance] was installed
                          • before/after results for the affected cases
                          • note on variance if results are noisy

                          Prefer reporting per-benchmark improvements rather than a single blended headline number.

                          Rollout plan

                          Phase 1

                          • add benchmark directory structure
                          • add shared benchmark fixtures
                          • add bt_json microbenchmarks

                          Phase 2

                          • add logger microbenchmarks
                          • add end-to-end tracing benchmark

                          Phase 3

                          • add dedicated nox sessions
                          • document local benchmark workflow in py/README.md or benchmark README

                          Acceptance criteria

                          • There is a documented benchmark workflow on main
                          • We can benchmark bt_json hot paths in isolation
                          • We can benchmark logger hot paths in isolation
                          • We can run one realistic tracing end-to-end benchmark
                          • We can compare JSON results across branches with pyperf compare_to
                          • Normal functional CI remains separate from benchmark execution

                          Why this should happen before PR #101-style perf work

                          PR #101 combines several categories of optimizations:

                          • serialization dispatch changes
                          • deep-copy specialization
                          • logger fast paths
                          • shared helper changes

                          Without a benchmark harness, it is too easy to:

                          • merge broad refactors based on one aggregate number
                          • overestimate wins from noisy runs
                          • miss regressions in shared helpers
                          • lose the ability to review each incremental optimization on its own merit

                          This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              Add a repeatable performance benchmark harness for tracing hot paths #134

                              Description

                              Add a repeatable performance benchmark harness for tracing hot paths

                              Summary

                              We should add a durable benchmarking setup for the Python SDK before landing tracing performance changes like the ones proposed in PR #101.

                              Today we have strong functional coverage, but no formal performance harness on main. That makes it hard to:

                              • validate that a performance-oriented refactor helps the intended hot path
                              • separate large wins from noise
                              • compare branch results against main
                              • catch accidental regressions in shared helpers like merge_dicts()
                              • evaluate behavior with and without optional perf dependencies like orjson

                              This issue proposes a benchmarking design that fits the repo's current nox and packaging setup.

                              Goals

                              • Add repeatable microbenchmarks for tracing hot-path functions
                              • Add one realistic end-to-end tracing benchmark
                              • Make it easy to compare main vs a feature branch locally
                              • Cover both minimal installs and the performance extra
                              • Keep performance checks separate from normal functional test runs

                              Non-goals

                              • Do not block CI on strict timing thresholds initially
                              • Do not turn benchmark timings into flaky pytest assertions
                              • Do not try to benchmark every SDK area at once

                              Proposed tooling

                              Use pyperf as the primary benchmark runner.

                              Rationale:

                              • much better statistical discipline than ad hoc time.perf_counter() loops
                              • supports warmups, calibration, repetition, and JSON output
                              • gives us a clean compare_to workflow for branch vs branch
                              • better fit for stable microbenchmarks than plain pytest timing

                              Keep a separate scenario-style benchmark for end-to-end tracing flows. That can still be implemented in Python, but it should be structured as a benchmark module rather than a one-off exploratory script.

                              Proposed layout

                              Under py/:

                              benchmarks/
                              README.md
                              conftest.py
                              cases/
                              bench_bt_json.py
                              bench_logger.py
                              bench_tracing_e2e.py
                              fixtures.py
                              

                              Notes:

                              • bench_bt_json.py should focus on serialization and deep-copy hot paths
                              • bench_logger.py should focus on span creation, split/sanitize paths, and internal-only logging
                              • bench_tracing_e2e.py should model a realistic root-span + child-span + logging workload
                              • fixtures.py should centralize representative payload builders so cases stay consistent

                              What to benchmark

                              1. bt_json microbenchmarks

                              Cover:

                              • _to_bt_safe() on primitives
                              • _to_bt_safe() on str/int subclasses and enum-backed strings
                              • _to_bt_safe() on dataclasses
                              • _to_bt_safe() on pydantic-like objects
                              • bt_safe_deep_copy() on small, medium, and large nested dict/list payloads
                              • bt_safe_deep_copy() on payloads with circular references
                              • bt_safe_deep_copy() on payloads with non-string dict keys

                              Why:

                              • these are the core hot paths in the PR under discussion
                              • they are easy to measure deterministically in isolation

                              2. logger microbenchmarks

                              Cover:

                              • _validate_and_sanitize_experiment_log_partial_args() on empty vs populated events
                              • _strip_nones() on shallow and nested dicts, with and without None
                              • split_logging_data() for:
                                • event only
                                • internal_data only
                                • both present
                                • BraintrustStream present
                              • SpanImpl creation with explicit name
                              • SpanImpl creation without name
                              • log_internal() with user event payloads
                              • log_internal() with internal-only payloads like end() / set_attributes()

                              Why:

                              • this isolates the logger-local optimizations from broader utility changes
                              • it lets us quantify small changes like lazy caller-location lookup separately from bigger serialization wins

                              3. End-to-end tracing benchmark

                              Cover one realistic scenario:

                              • create root span
                              • add representative input / metadata
                              • create child span
                              • log output / metadata / metrics
                              • end child and root spans
                              • exercise background logger record construction without external I/O

                              Parameters:

                              • run enough iterations to reduce noise
                              • keep payload shapes fixed and version-controlled
                              • report separate scenarios for:
                                • medium payload
                                • large payload
                                • internal-only updates

                              Why:

                              • microbenchmarks show where a win comes from
                              • the e2e benchmark shows whether the win is real in the path users care about

                              Benchmark environments

                              Run each benchmark suite in at least two environments:

                              A. Minimal environment

                              • install the package plus benchmark deps
                              • do not install optional perf extras

                              Purpose:

                              • captures baseline behavior for default installs

                              B. Performance environment

                              • install .[performance]

                              Purpose:

                              • measures the effect of optional orjson
                              • ensures future performance work does not accidentally optimize only one environment

                              If needed later, add a second Python minor version to confirm interpreter-sensitive behavior, but that should not be required for the initial rollout.

                              Nox integration

                              Add dedicated nox sessions instead of folding this into test_core.

                              Suggested sessions:

                              • perf_bt_json
                              • perf_logger
                              • perf_e2e
                              • optional umbrella session: perf

                              Behavior:

                              • install pyperf
                              • install the package from source
                              • optionally install .[performance] in dedicated sessions or via a flag
                              • run benchmark modules directly
                              • emit JSON results to a temp or ignored output path

                              Example local workflows:

                              cd py
                              nox -s perf_bt_json
                              nox -s perf_logger
                              nox -s perf_e2e

                              Comparison workflow:

                              cd py
                              nox -s perf_bt_json -- --output /tmp/main-bt-json.json
                              nox -s perf_bt_json -- --output /tmp/branch-bt-json.json
                              pyperf compare_to /tmp/main-bt-json.json /tmp/branch-bt-json.json

                              We can refine exact command shape during implementation, but the key point is to make branch-vs-branch comparison first-class.

                              Local-first workflow

                              The initial implementation should optimize for local developer use only.

                              That means:

                              • no CI integration yet
                              • no nightly jobs
                              • no hard performance thresholds

                              Instead, the benchmark harness should make it easy to:

                              • establish a baseline on main
                              • run the same benchmark on a feature branch
                              • compare pyperf JSON outputs locally

                              The command surface should stay very small and obvious:

                              cd py
                              nox -s perf_bt_json
                              nox -s perf_logger
                              nox -s perf_e2e

                              And branch comparison should be a first-class local workflow:

                              cd py
                              nox -s perf_bt_json -- --output /tmp/main.json
                              nox -s perf_bt_json -- --output /tmp/branch.json
                              pyperf compare_to /tmp/main.json /tmp/branch.json

                              Reporting expectations for performance PRs

                              Any PR that claims performance improvement should include:

                              • benchmark command(s) used
                              • environment details:
                                • Python version
                                • whether .[performance] was installed
                              • before/after results for the affected cases
                              • note on variance if results are noisy

                              Prefer reporting per-benchmark improvements rather than a single blended headline number.

                              Rollout plan

                              Phase 1

                              • add benchmark directory structure
                              • add shared benchmark fixtures
                              • add bt_json microbenchmarks

                              Phase 2

                              • add logger microbenchmarks
                              • add end-to-end tracing benchmark

                              Phase 3

                              • add dedicated nox sessions
                              • document local benchmark workflow in py/README.md or benchmark README

                              Acceptance criteria

                              • There is a documented benchmark workflow on main
                              • We can benchmark bt_json hot paths in isolation
                              • We can benchmark logger hot paths in isolation
                              • We can run one realistic tracing end-to-end benchmark
                              • We can compare JSON results across branches with pyperf compare_to
                              • Normal functional CI remains separate from benchmark execution

                              Why this should happen before PR #101-style perf work

                              PR #101 combines several categories of optimizations:

                              • serialization dispatch changes
                              • deep-copy specialization
                              • logger fast paths
                              • shared helper changes

                              Without a benchmark harness, it is too easy to:

                              • merge broad refactors based on one aggregate number
                              • overestimate wins from noisy runs
                              • miss regressions in shared helpers
                              • lose the ability to review each incremental optimization on its own merit

                              This harness should make those changes measurable and allow them to land more incrementally, while staying easy to run locally before we decide whether any CI integration is worthwhile.

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions