feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(completions): emit UsageEvent on /v1/completions 200 (#403) - #426

Merged
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit
May 27, 2026
Merged

feat(completions): emit UsageEvent on /v1/completions 200 (#403)#426
moonming merged 2 commits into
mainfrom
feat/issue-403-completions-usage-emit

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Fixes#403. Sibling of PR #402 (embeddings) and PR #425 (responses).

Pre-fix, `/v1/completions` dropped the `UsageEvent` entirely. Despite OpenAI deprecating the legacy completions endpoint, customer traffic on it is still non-trivial in production (every `gpt-3.5-turbo-instruct` integration, every old OpenAI Python SDK path) — and all of it was invisible to cp-api's budget ledger.

This PR mirrors the #402 pattern: extract the upstream `usage` block, emit a UsageEvent on success.

Emit semantics

PathEmit?Reason
200 with `usage.prompt_tokens` presentyesupstream-reported
200 with `usage: {}` or missing `prompt_tokens`nomalformed upstream; avoids noise rows (audit MEDIUM-1 precedent from #425)
200 with no `usage` blocknoedge / error shape
501 NotImplementednono upstream call happened
4xx / 5xxnono usage data to attribute

Per-PR audit-precedent fixes applied preemptively

The audit on the parallel PR #425 (responses) raised two MEDIUMs that apply structurally to this handler too:

  • MEDIUM-1 — tighten the gate so `usage: {}` doesn't emit zero-everything rows. Applied here via `usage.prompt_tokens` presence requirement in `extract_completion_usage`.
  • MEDIUM-2 — add negative pinning so 4xx/5xx code paths can't silently start emitting. New test `upstream_5xx_does_not_emit_usage_event` covers this.

Both lessons baked in from the start rather than addressed post-merge.

Test plan

  • `emits_usage_event_on_200_with_tokens_issue_403` — pins prompt + completion tokens, status code, model_id, api_key_id, protocol against a canonical legacy completions response
  • `skips_usage_event_when_upstream_usage_block_is_empty` — pins the malformed `usage: {}` edge → no emit
  • `upstream_5xx_does_not_emit_usage_event` — negative pinning for the error path
  • All 5 pre-existing `completions::tests` still pass
  • `cargo clippy -p aisix-proxy -- -D warnings` clean

References

Pre-#403, /v1/completions dropped the UsageEvent entirely.
Customers using the legacy completions endpoint (still in
production at significant volume despite deprecation) had spend
invisible to cp-api's budget ledger and customer-facing /logs
analytics.
This PR mirrors PR #402 (embeddings) and PR #425 (responses).
The handler now extracts the upstream `usage` block and emits a
UsageEvent with:
- `prompt_tokens` = usage.prompt_tokens
- `completion_tokens` = usage.completion_tokens
- `status_code`, `model_id`, `api_key_id`, `latency_ms`
- `inbound_protocol` = "openai"
- `handler` = "completions" (matches #408 enumeration)
Emit semantics (mirrors #402 / #425, with PR #425 audit MEDIUM-1
applied preemptively):
- 200 with `usage.prompt_tokens` present → emit
- 200 with `usage: {}` or missing `prompt_tokens` → no emit
(upstream-malformed; avoids zero-everything noise rows)
- 200 with no `usage` block → no emit
- 501 NotImplemented (provider lacks completions) → no emit
- 4xx/5xx error path → no emit
The `usage.prompt_tokens` gate is tightened per audit MEDIUM-1 on
PR #425: per the OpenAI spec the field is required on every
legitimate completion response, so its absence is upstream-
malformed rather than a real zero-spend reply.
Tests (3 new):
- `emits_usage_event_on_200_with_tokens_issue_403` — pins all
fields against a canonical legacy completions response
- `skips_usage_event_when_upstream_usage_block_is_empty` — pins
the `usage: {}` edge (audit-precedent from #425 MEDIUM-1)
- `upstream_5xx_does_not_emit_usage_event` — negative pinning,
ensures 4xx/5xx paths never reach the emit site (audit MEDIUM-2
from #425 applied preemptively)
References:
- Parent: #226 (non-chat handlers don't emit)
- Sibling MVPs: #402 (embeddings), #425 (responses)
- Spec: <https://platform.openai.com/docs/api-reference/completions/object>
@coderabbitai

coderabbitaiBot commented May 27, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 4 minutes and 28 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: d0a69a80-5bfa-482f-9947-5729562383f4

📥 Commits

Reviewing files that changed from the base of the PR and between ba47e37 and be0fa96.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/completions.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

PR #426 audit raised 3 MEDIUM + 3 LOW findings. Addressed
inline; LOW-2 filed as #429 (cross-handler tightening that
touches #425 too).
MEDIUM-1 — Doc comment on `CompletionDispatchSuccess.model_id`
referenced a non-existent `upstream_called` field (stale copy
from embeddings.rs where it does exist). Corrected to describe
the actual gating channel (`usage.is_some()`).
MEDIUM-2 — 501 NotImplemented path was recording `status=200` in
the access log and prometheus metrics (hardcoded `200u16`),
making it impossible for operators to distinguish real successes
from "provider does not support completions". Same systemic bug
PR #404 and PR #405 fix in their own handlers. Switched to
`success.response.status().as_u16()` + `RequestOutcome::from_status`
to mirror the convention. UsageEvent emission was already
correctly skipped via `usage: None`, so billing wasn't affected
— only observability.
MEDIUM-3 — No test exercised the 501 path. Added
`provider_lacking_complete_returns_501_without_emit`: registers
an `AnthropicBridge` (which doesn't override `Bridge::complete()`
so the trait default returns `BridgeError::Config` → 501),
routes a `/v1/completions` request at an Anthropic model, and
pins both the 501 response status AND the absence of any
UsageEvent on the sink.
LOW-1 — Added `skips_usage_event_when_upstream_omits_usage_block_entirely`
test that exercises the outer `body.get("usage")?` short-circuit
in `extract_completion_usage` (the existing
`*_when_upstream_usage_block_is_empty` test only covered the
inner `prompt_tokens` missing case).
LOW-3 — Comment referenced `#404` PR as if merged; on this branch
it's now true (#425 landed) so the reference is correct.
LOW-2 (gate `completion_tokens` on presence) deliberately deferred
to #429 for cross-handler symmetry with #425's responses.rs.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

UsageEvent emission missing on /v1/completions (#226 follow-up)

1 participant

@moonming