Skip to content

[bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

Description

@braintrust-bot

Summary

The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

What is missing

In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

  • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
  • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
  • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
  • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

Braintrust docs status: not_found

Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

Upstream sources

Local repo files inspected

  • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
  • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
  • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
    Skip to content

    [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

    Description

    @braintrust-bot

    Summary

    The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

    The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

    What is missing

    In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

    • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
    • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
    • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
    • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

    Braintrust docs status: not_found

    Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

    Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

    Upstream sources

    Local repo files inspected

    • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
    • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
    • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      No labels
      No labels

      Type

      Projects

      No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
      Skip to content

      [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

      Description

      @braintrust-bot

      Summary

      The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

      The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

      What is missing

      In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

      • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
      • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
      • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
      • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

      Braintrust docs status: not_found

      Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

      Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

      Upstream sources

      Local repo files inspected

      • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
      • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
      • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        Projects

        No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
        Skip to content

        [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

        Description

        @braintrust-bot

        Summary

        The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

        The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

        What is missing

        In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

        • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
        • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
        • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
        • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

        Braintrust docs status: not_found

        Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

        Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

        Upstream sources

        Local repo files inspected

        • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
        • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
        • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

        Metadata

        Metadata

        Assignees

        No one assigned

          Labels

          No labels
          No labels

          Type

          Projects

          No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
          Skip to content

          [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

          Description

          @braintrust-bot

          Summary

          The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

          The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

          What is missing

          In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

          • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
          • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
          • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
          • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

          Braintrust docs status: not_found

          Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

          Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

          Upstream sources

          Local repo files inspected

          • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
          • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
          • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            Projects

            No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
            Skip to content

            [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

            Description

            @braintrust-bot

            Summary

            The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

            The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

            What is missing

            In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

            • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
            • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
            • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
            • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

            Braintrust docs status: not_found

            Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

            Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

            Upstream sources

            Local repo files inspected

            • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
            • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
            • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

            Metadata

            Metadata

            Assignees

            No one assigned

              Labels

              No labels
              No labels

              Type

              Projects

              No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) · Issue #164 · braintrustdata/braintrust-sdk-java · GitHub
              Skip to content

              [bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span) #164

              Description

              @braintrust-bot

              Summary

              The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

              The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

              What is missing

              In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

              • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
              • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
              • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
              • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

              Braintrust docs status: not_found

              Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

              Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

              Upstream sources

              Local repo files inspected

              • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
              • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
              • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                Projects

                No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions