feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonmingmoonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

PathEmit?Reason
200 OKyesupstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplementednono upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx errornono usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.
This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).
Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
call happened, so prompt_tokens=0 signals "skip emit" to the
handler (avoids attributing zero-token spend to the api_key and
bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)
Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.
Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.
Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitaiBot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s)Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)
PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:
- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
for empty input or provider-specific billing) — emit MUST happen
The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.
Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant

@moonming