Uh oh!
There was an error while loading. Please reload this page.
feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402
Conversation
Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every embeddings call was invisible to cp-api's budget ledger and the customer-facing /logs analytics. Chat completions has emitted UsageEvents since #302 M17, so the gap was specifically on the non-chat OpenAI-shape handlers. This PR ships the MVP for #226: /v1/embeddings now emits a UsageEvent on 200 with the upstream-reported prompt_tokens, the resolved model_id, the authenticated api_key_id, the status code, and inbound_protocol = "openai" (matching chat.rs convention). Emit-on-success-only mirrors chat.rs: - 200 path → emit with real prompt_tokens from upstream usage block - 501 Not Implemented (provider lacks embed support) → no upstream call happened, so prompt_tokens=0 signals "skip emit" to the handler (avoids attributing zero-token spend to the api_key and bloating /logs with noise) - Upstream error path → no UsageEvent (no usage data to attribute) Per-PK telemetry attribution (provider_kind / featured / branded_provider / pk_label / byo_label) is wired for chat only; filed as a follow-up so the non-chat handlers gain the same dashboard-slicing surface. Follow-ups (separate PRs) for /v1/completions, /v1/responses, /v1/rerank, /v1/audio/*, /v1/images/* — same shape, same emit-on-success-only convention. Test: unit test asserts a successful /v1/embeddings call enqueues exactly one UsageEvent on the sink with the expected prompt_tokens, model_id, api_key_id, status_code, inbound_protocol, and non-empty request_id + occurred_at. Mirrors chat.rs's existing UsageSink- based unit-test pattern.
Warning Review limit reached
More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Free Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughThis PR enhances the embeddings endpoint to emit ChangesEmbeddings usage event emission
🎯 3 (Moderate) | ⏱️ ~25 minutes Note 🎁 Summarized by CodeRabbit FreeYour organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above. Comment |
…mpt_tokens (#226 audit M1) PR #402 audit found that gating emission on `prompt_tokens > 0` conflates two distinct cases: - 501 NotImplemented (no upstream call) — emit MUST be skipped - 200 with upstream-reported `prompt_tokens=0` (rare but possible for empty input or provider-specific billing) — emit MUST happen The audit-suggested fix: replace the numeric sentinel with an explicit `upstream_called: bool` flag on EmbedDispatchSuccess. Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the handler gates on the flag directly. Regression test pins the post-fix contract: a 200 with `usage.prompt_tokens=0` still produces exactly one UsageEvent on the sink, attributed to the authenticated api_key for compliance / audit even when the billable count is zero.
Uh oh!
There was an error while loading. Please reload this page.
Summary
Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.
Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.
This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.
Scope
MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.
Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.
Emission convention
Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:
Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.
Audit response
Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:
Test plan
References