fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(telemetry): emit usage event on streaming /v1/responses (#808) - #613

Merged
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808
Jun 15, 2026
Merged

fix(telemetry): emit usage event on streaming /v1/responses (#808)#613
jarvis9443 merged 2 commits into
mainfrom
fix/responses-streaming-usage-808

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Problem

A successful streaming/v1/responses request emitted no usage event. The verbatim streaming path forwarded the upstream SSE bytes without parsing them, returned usage: None, and the handler only emits a UsageEvent when usage is Some. So a streamed 200 produced no row in dpmgr_usage_events (invisible in the dashboard Logs and the budget ledger), while a 4xx/5xx on the same endpoint still produced a zero-token row via the error path.

This hit clients that always stream — the OpenAI Codex CLI talks to /v1/responses and always streams, so every successful Codex call was unlogged while its failures were logged.

Fix

Wrap the forwarded byte stream so the terminal response.completed event's usage block is parsed in-flight, and emit the UsageEvent from the stream's Drop guard at end-of-stream (or client-disconnect). response.incomplete / response.failed carry the same counts on truncation/cancellation and are handled too. Bytes still forward verbatim — the client sees the exact upstream SSE shape.

  • Verbatim streaming path → Drop-guard emit (usage_handled_by_stream guards against a double-emit).
  • Buffered output-guardrail path → parses usage from its held buffer and emits from the handler.
  • Token counts captured: input_tokens, output_tokens, plus reasoning_tokens / cached_tokens sub-counts.
  • SSE framing helpers are shared with the /v1/messages passthrough rather than duplicated.

This mirrors the end-of-stream emission the Anthropic /v1/messages streaming path already does.

Behavior change

Streaming /v1/responses 200s now emit one usage event per request (parity with non-streaming and with chat completions). No config or wire-shape change.

Tests

  • Integration test: a streamed 200 emits exactly one UsageEvent with the terminal-event token counts (reasoning + cached included). Fails before the fix (no event), passes after.
  • DP standalone E2E (responses-streaming-usage-e2e): real aisix binary + mock upstream streaming a response.completed + mock OTLP receiver, asserting the emitted gen_ai.usage.input_tokens / output_tokens / status_code. Existing responses-endpoint and per-attempt-telemetry E2E pass unchanged.

Docs

Added a Responses API usage-accounting note covering streaming vs non-streaming.

Fixes api7/AISIX-Cloud#808

Summary by CodeRabbit

  • Bug Fixes

    • Fixed duplicate usage event emission for streaming requests.
    • Streaming responses now properly emit usage events with accurate token counts at stream completion.
  • Documentation

    • Added usage accounting documentation for the /v1/responses endpoint, clarifying how token counts are tracked for streaming requests.
  • Tests

    • Added end-to-end test for streaming usage event emission.

The verbatim streaming `/v1/responses` path forwarded SSE bytes without
parsing them, returned `usage: None`, and the handler only emitted a
UsageEvent when usage was `Some` — so a successful streamed request
emitted no usage event at all, while a 4xx/5xx still produced a
zero-token row via the error path. Clients that always stream (e.g. the
OpenAI Codex CLI) were therefore invisible to the dashboard Logs and the
budget ledger on every success, but visible on failure.
Wrap the forwarded byte stream so the terminal `response.completed`
event's `usage` block (also `response.incomplete` / `response.failed`,
which carry the same counts on truncation/cancellation) is parsed
in-flight, and emit the UsageEvent from the stream's Drop guard at
end-of-stream (or client-disconnect) — the same end-of-stream emission
the Anthropic `/v1/messages` streaming path already uses. Bytes still
forward verbatim. The buffered output-guardrail path parses usage from
its held buffer and emits from the handler. The SSE framing helpers are
shared with the `/v1/messages` passthrough.
Tests: integration test asserting a streamed 200 emits one UsageEvent
with the terminal-event token counts (fails before, passes after); DP
standalone E2E (`responses-streaming-usage-e2e`) driving a real `aisix`
binary + mock upstream + OTLP receiver to assert the emitted token
counts. Docs: Responses API usage-accounting note.
@coderabbitai

coderabbitaiBot commented Jun 15, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@jarvis9443, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 41 minutes, and 9 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3a5768d8-35a9-4515-96f2-a18ce9cc93cf

📥 Commits

Reviewing files that changed from the base of the PR and between 74d40fd and 0457020.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
📝 Walkthrough

Walkthrough

Fixes missing usage log emission for streaming /v1/responses requests (#808). SSE frame parsing helpers in messages.rs are exposed as pub(crate). ResponseDispatchSuccess gains a usage_handled_by_stream flag to prevent duplicate emission. A Drop-guarded stream wrapper parses terminal response.completed/response.incomplete/response.failed SSE events in-flight and emits UsageEvent exactly once. An E2E regression test and documentation are added.

Changes

Streaming Usage Emission for /v1/responses

Layer / File(s)Summary
Expose shared SSE frame parsing helpers
crates/aisix-proxy/src/messages.rs
MAX_SSE_FRAME_BUF_BYTES, find_frame_end, and extract_sse_data_line are promoted from private to pub(crate) so the /v1/responses streaming usage parser can reuse them.
ResponseDispatchSuccess flag and signature extensions
crates/aisix-proxy/src/responses.rs
ResponseDispatchSuccess gains usage_handled_by_stream; ResponseUsage derives Default and Clone; dispatch and responses_to_target signatures are extended with started, client, model_name, api_key_id, and AttemptInfo to thread request context into the streaming path.
Conditional UsageEvent emission gate
crates/aisix-proxy/src/responses.rs
Winning-attempt UsageEvent emission is gated on !success.usage_handled_by_stream; non-streaming and guardrail-blocked paths explicitly set usage_handled_by_stream = false.
Streaming path: buffered guardrail and verbatim passthrough
crates/aisix-proxy/src/responses.rs
Buffered output-guardrail path extracts terminal SSE usage from buffered bytes with usage_handled_by_stream = false. Verbatim passthrough path wraps the upstream stream; ResponseDispatchSuccess is returned with usage = None and usage_handled_by_stream = true.
SSE parsing infrastructure and Drop-guarded stream wrapper
crates/aisix-proxy/src/responses.rs
New helpers extract terminal usage from response.completed/response.incomplete/response.failed; ResponsesUsageGuard implements Drop for exactly-once UsageEvent emission; build_responses_passthrough_stream forwards bytes verbatim while parsing in-flight.
Updated unit tests and E2E regression test
crates/aisix-proxy/src/responses.rs, tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts
Unit test for #808 is rewritten to assert exactly one UsageEvent with parsed token counts and verbatim SSE passthrough. A new E2E test uses an in-process OTLP receiver to assert emitted token counts via span attributes after a streamed request.
Usage Accounting documentation
docs/integration/responses.md
New "Usage Accounting" section describes how token usage events are emitted for both streaming and non-streaming /v1/responses calls, including terminal event sources and emission timing.

Sequence Diagram(s)

sequenceDiagram
participant Client
participant Gateway as aisix-proxy /v1/responses handler
participant Upstream as OpenAI Responses API
participant UsageGuard as ResponsesUsageGuard (Drop)
participant Telemetry as UsageEvent emitter
Client->>Gateway: POST /v1/responses (stream: true)
Gateway->>Upstream: upstream streaming request
Upstream-->>Gateway: SSE: response.created
Upstream-->>Gateway: SSE: response.output_text.delta (×N)
Upstream-->>Gateway: SSE: response.completed { usage }
Upstream-->>Gateway: SSE: [DONE]
Gateway-->>Client: verbatim SSE bytes (forwarded in-flight)
Note over Gateway,UsageGuard: Stream ends or client disconnects
UsageGuard->>Telemetry: emit UsageEvent(input_tokens, output_tokens, ...)
Telemetry-->>Gateway: usage_handled_by_stream = true (skip handler emission)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • api7/ai-gateway#363: Shares the same messages.rs SSE frame parsing and stream-path UsageEvent emission pattern for /v1/messages that this PR replicates and extends for /v1/responses.
  • api7/ai-gateway#563: Both modify responses_to_target's streaming path by wrapping the upstream byte stream; the timeout-fallback stream wrapping in #563 is structurally adjacent to the usage-emission wrapper introduced here.
  • api7/ai-gateway#495: Both touch /v1/responses usage-event extraction in responses.rs; the terminal-usage parsing introduced here is directly related to the extract_response_usage gating from #495.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly identifies the main change: fixing telemetry usage event emission for streaming /v1/responses requests, directly addressing issue #808.
Linked Issues check✅ PassedThe PR fully addresses issue #808 by implementing usage event emission for streaming /v1/responses requests, achieving consistent logging for successful requests.
Out of Scope Changes check✅ PassedAll changes are directly related to fixing streaming /v1/responses usage event emission; changes to messages.rs, responses.rs, documentation, and tests are all in-scope for issue #808.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts (1)

92-94: ⚡ Quick win

Don’t silently swallow OTLP parse failures in the test receiver.

Line 92 ignores malformed payloads, which turns exporter regressions into opaque timeouts. Persist the last parse error (with a short payload snippet) and include it in the timeout error to keep failures actionable.

Suggested patch
 interface OtlpReceiver {
url: string;
spanAttrs: Array<Record<string, string>>;
+ parseErrors: string[];
close(): Promise<void>;
}
@@
async function startOtlpReceiver(): Promise<OtlpReceiver> {
const spanAttrs: Array<Record<string, string>> = [];
+ const parseErrors: string[] = [];
@@
- } catch {- // ignore malformed bodies — assertions fail on missing spans+ } catch (err) {+ parseErrors.push(+ `parse error: ${(err as Error).message}; body=${raw.slice(0, 256)}`,+ );
}
@@
url: `http://127.0.0.1:${port}/v1/traces`,
spanAttrs,
+ parseErrors,
@@
- throw new Error(`no usage span for request_id=${requestId}`);+ const parseHint = recv.parseErrors.at(-1);+ throw new Error(+ `no usage span for request_id=${requestId}` ++ (parseHint ? `; last otlp receiver error: ${parseHint}` : ""),+ );
}

As per coding guidelines, "Handle errors gracefully with meaningful error messages."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts` around lines 92 -
94, The catch block at lines 92-94 silently swallows OTLP parse failures, which
masks exporter regressions and causes tests to fail with opaque timeouts.
Instead of ignoring the error, capture the parse failure and store it (along
with a short snippet of the malformed payload) in a variable that persists
across iterations. Then, when the test times out waiting for assertions to pass
due to missing spans, include the stored parse error details in the timeout
error message so developers can quickly diagnose what went wrong in the
exporter.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 119-123: The test uses find() to retrieve the first span matching
the request ID, which does not validate the dedup contract or ensure exactly one
usage span exists. Replace the weak first-match lookup by filtering to
usage-bearing spans for the request and asserting cardinality equals 1 before
accessing token values. This validation is needed at multiple locations in the
test file (the primary location at the recv.spanAttrs.find() call and also at a
similar pattern that occurs later in the file) to properly verify that duplicate
UsageEvent spans are not emitted for the same request.
---
Nitpick comments:
In `@tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts`:
- Around line 92-94: The catch block at lines 92-94 silently swallows OTLP parse
failures, which masks exporter regressions and causes tests to fail with opaque
timeouts. Instead of ignoring the error, capture the parse failure and store it
(along with a short snippet of the malformed payload) in a variable that
persists across iterations. Then, when the test times out waiting for assertions
to pass due to missing spans, include the stored parse error details in the
timeout error message so developers can quickly diagnose what went wrong in the
exporter.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 89cc6b80-97ca-42c3-a516-5939075b2986

📥 Commits

Reviewing files that changed from the base of the PR and between 2ba65f1 and 74d40fd.

📒 Files selected for processing (4)
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • docs/integration/responses.md
  • tests/e2e/src/cases/responses-streaming-usage-e2e.test.ts

Comment threadtests/e2e/src/cases/responses-streaming-usage-e2e.test.ts Outdated
…equest
Strengthen the #808 E2E to assert cardinality (=== 1) of usage-bearing
spans for the measured request_id instead of a first-match lookup, so a
duplicate-emit regression of the usage_handled_by_stream dedup guard is
caught at the E2E layer too (the unit test already pins single emission).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443