feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP) - #425

Merged
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit
May 27, 2026
Merged

feat(responses): emit UsageEvent on /v1/responses 200 non-streaming (#404 MVP)#425
moonming merged 2 commits into
mainfrom
feat/issue-404-responses-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 27, 2026

Copy link
Copy Markdown
Member

Summary

Fixes#404 (non-streaming MVP). Sibling of #402 (embeddings).

Pre-fix, `/v1/responses` dropped the `UsageEvent` entirely. Every o1/o3/GPT-5 request via OpenAI's Responses API (the modern default entry for reasoning models) was invisible to cp-api's budget ledger and customer-facing /logs analytics — meaning customer spend on the most expensive model class wasn't being counted at all.

This PR emits a UsageEvent on every successful non-streaming 200 with:

UsageEvent fieldMirrors
`prompt_tokens``usage.input_tokens`
`completion_tokens``usage.output_tokens`
`reasoning_tokens``usage.output_tokens_details.reasoning_tokens` (o1/o3/GPT-5 only)
`cached_prompt_tokens``usage.input_tokens_details.cached_tokens`
`inbound_protocol``"openai"`

Scope: non-streaming only

Streaming path is deferred. Responses-API streaming surfaces usage in the final `response.completed` SSE event, but the current handler passes the byte stream through verbatim and the gateway doesn't parse SSE chunks here. SSE byte-stream interception is its own design surface; tracked as a #404 streaming follow-up so the MVP can land and start capturing the majority of traffic (non-streaming).

Same MVP pattern PR #402 used for embeddings.

Emit semantics

PathEmit?Reason
Non-stream 200 with `usage`yesupstream-reported counters
Non-stream 200 without `usage`noedge/error shape, nothing to attribute
Streaming 200no (today)SSE parse not wired; follow-up
4xx/5xxnono usage data

Test plan

  • `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4 token counters, status code, model_id, api_key_id, protocol against a canonical Responses-API response body
  • `skips_usage_event_when_upstream_omits_usage_block` — pins the no-emit edge for 200 without `usage`
  • `streaming_path_does_not_emit_today_but_passes_through` — pins the streaming MVP contract (no regression, no emission yet)
  • All 5 pre-existing `responses::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Summary by CodeRabbit

Release Notes

  • New Features
    • Enhanced usage tracking: Non-streaming API responses now capture and emit usage metrics, including detailed token consumption data, providing improved visibility into API resource utilization and enabling better usage monitoring.

Review Change Stack

…404)
Pre-#404, /v1/responses dropped the UsageEvent entirely. Every
o1/o3/GPT-5 traffic through OpenAI's Responses API (the modern
default entry point for reasoning models) was invisible to
cp-api's budget ledger and customer-facing /logs analytics.
This PR ships the non-streaming MVP. The handler now extracts the
upstream `usage` block and emits a UsageEvent with:
- `prompt_tokens` = usage.input_tokens
- `completion_tokens` = usage.output_tokens
- `reasoning_tokens` = usage.output_tokens_details.reasoning_tokens
(uniquely surfaced for o1/o3/GPT-5 class models)
- `cached_prompt_tokens` = usage.input_tokens_details.cached_tokens
(OpenAI prompt-cache hit subset of input_tokens)
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
Streaming path scope-deferred: SSE byte-stream interception is
needed to extract usage from the final `response.completed` event,
and the current responses.rs handler passes the byte stream
through verbatim. Tracked as a #404 streaming follow-up — same
MVP pattern #402 used for embeddings.
Emit-on-success-only:
- non-stream 200 with `usage` present → emit
- non-stream 200 without `usage` block (edge / error shapes) → no emit
- streaming 200 → no emit today (follow-up)
- 4xx/5xx → no emit (no usage data to attribute)
Tests (3 new):
- `emits_usage_event_on_200_non_streaming_issue_404` — pins all 4
token counters, status_code, model_id, api_key_id, protocol
- `skips_usage_event_when_upstream_omits_usage_block` — pins the
no-emit edge for 200 without usage
- `streaming_path_does_not_emit_today_but_passes_through` — pins
the streaming MVP contract (no regression, no emission until
the follow-up lands)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVP: #402 (embeddings)
- Spec: <https://platform.openai.com/docs/api-reference/responses>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 9 minutes and 54 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 88779164-c549-4ae1-b04d-05b16dfda032

📥 Commits

Reviewing files that changed from the base of the PR and between e8a9888 and 981c2c2.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/responses.rs
📝 Walkthrough

Walkthrough

The Responses API handler now emits usage telemetry from upstream successful responses. Internal structs (ResponseDispatchSuccess and ResponseUsage) capture usage metadata and token counters. The dispatch function's return type changes to flow provider, model, and optional usage back to the handler, which extracts usage from non-streaming responses and publishes UsageEvent to the telemetry pipeline while streaming responses return None for usage.

Changes

Usage Telemetry Integration

Layer / File(s)Summary
Type contracts and imports
crates/aisix-proxy/src/responses.rs
Adds UsageEvent import and defines ResponseDispatchSuccess with usage: Option<ResponseUsage> and ResponseUsage with token fields to carry metadata through dispatch and handler.
Dispatch function return type
crates/aisix-proxy/src/responses.rs
Changes dispatch return from (Response, String) tuple to Result<ResponseDispatchSuccess, ProxyError> to flow provider, model, and optional usage metadata with the response.
Dispatch branches with usage extraction
crates/aisix-proxy/src/responses.rs
Streaming branch returns usage: None; non-streaming branch parses upstream JSON, extracts usage counters via extract_response_usage, and returns populated ResponseDispatchSuccess. Adds emit_usage_event helper to construct and publish UsageEvent to the telemetry sink.
Handler success path and event emission
crates/aisix-proxy/src/responses.rs
Unpacks ResponseDispatchSuccess in handler success path, emits AccessLog with resolved model id, and additionally emits UsageEvent to OTLP exporters when dispatch supplies non-None usage.
Usage telemetry tests
crates/aisix-proxy/src/responses.rs
Tests validate UsageEvent emission for non-streaming responses with correct token mapping, absence when upstream omits usage, and successful streaming passthrough without usage events.

Sequence Diagram

sequenceDiagram
participant Client
participant ResponsesHandler
participant Dispatch
participant UpstreamAPI
participant UsageSink
participant OTLPExporter
Client->>ResponsesHandler: HTTP request
ResponsesHandler->>Dispatch: invoke dispatch
Dispatch->>UpstreamAPI: forward request
UpstreamAPI-->>Dispatch: response + usage block (non-streaming)
Dispatch->>Dispatch: extract usage counters
Dispatch-->>ResponsesHandler: ResponseDispatchSuccess{response, usage, model_id}
ResponsesHandler->>ResponsesHandler: emit AccessLog
ResponsesHandler->>ResponsesHandler: construct UsageEvent
ResponsesHandler->>UsageSink: publish UsageEvent
UsageSink->>OTLPExporter: fanout usage telemetry
ResponsesHandler-->>Client: HTTP response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…OW-1, LOW-2)
PR #425 audit raised 2 MEDIUM + 2 LOW findings. All addressed:
MEDIUM-1 — `extract_response_usage` returned `Some(zeros)` for
`usage: {}` (malformed but present). Per the OpenAI Responses-API
spec, `input_tokens` is required on every non-streaming 200, so
its absence is upstream-malformed rather than a legitimate
zero-spend reply. Tightened: gate emit on `input_tokens` being
present and parseable; missing → no emit. New test
`skips_usage_event_when_usage_block_is_empty_audit_m1` pins this.
MEDIUM-2 — `upstream_error_returns_502` asserted only the status
code, not the absence of a UsageEvent on the sink. A future
regression that moved `emit_usage_event` into the error branch
would silently ship. New test `upstream_5xx_does_not_emit_usage_event`
adds the negative pinning.
LOW-1 — Doc comment on `emit_usage_event` claimed completion /
cached / reasoning were "left at Default::default()" but those
are precisely the fields *populated*. Copy-paste bug from
embeddings.rs (where they really were defaulted). Corrected to
list cache_creation / cache_read / provider_request_id / etc.
LOW-2 — Streaming pass-through test drained the body but didn't
assert its shape. A regression that broke the byte-stream wiring
in the new ResponseDispatchSuccess path could surface as 200 +
empty body and still pass. Added `body_bytes.starts_with("data:")`
to pin the SSE shape survives the refactor.
@moonming
moonming merged commit 4727dfc into mainMay 27, 2026
8 checks passed
@moonming
moonming deleted the feat/issue-404-responses-usage-emit branch May 27, 2026 04:13
moonming added a commit that referenced this pull request May 27, 2026
PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/responses (#226 follow-up)

1 participant

@moonming