feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a) - #321

Merged
moonming merged 2 commits into
mainfrom
feat/vertex-wire
May 17, 2026
Merged

feat(provider-vertex): wire Gemini chat via pre-minted OAuth2 bearer (D5.2.a)#321
moonming merged 2 commits into
mainfrom
feat/vertex-wire

Conversation

@moonming

@moonmingmoonming commented May 17, 2026

Copy link
Copy Markdown
Member

Summary

Replaces the skeleton's "not yet implemented" stubs in aisix-provider-vertex with real Vertex AI dispatch for the google publisher (Gemini chat). Other Vertex publishers + streaming surface clear publisher-named follow-up errors.

Scope: deliberately minimalchat() for google only, with credentials supplied as a pre-minted OAuth2 access token (operator manages refresh). Mirrors the D6 #319 / D7.2.a #320 pattern.

Wire shape (Gemini on Vertex)

Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:

  • URL:POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent
  • Body: Gemini's generateContent JSON:
    • contents[].role is \"user\" or \"model\" (NOT \"assistant\")
    • contents[].parts[].text (not content)
    • systemInstruction is top-level — system messages do NOT appear in contents[]
    • generationConfig.temperature / topP / maxOutputTokens (camelCase, emitted only when set)
  • Auth:Authorization: Bearer <access_token> (pre-minted OAuth2 token from operator)

Credentials convention

ProviderKey.secret is JSON-encoded {access_token, project, region}. The operator manages token refresh (GCP TTL ~1 hour). D5.1 follow-up adds in-process token mint via service_account_json (e.g. yup-oauth2 / gcp_auth).

project, region, and upstream_id are validated for URL-control chars (/, ?, #, \\n, ..) before being interpolated into the path — defense in depth against malicious model_name redirecting dispatch or corrupting metrics labels.

Audit lessons applied proactively

Written after the D6 audit (#319) and D7.2.a audit (#320) results came back. The lessons baked in from the start:

  • bridge.chat() exercised end-to-end via wiremock + a #[cfg(test)] api_base_override seam
  • 4xx error body redacted to canned status-keyed phrases (does NOT echo project id from Vertex's error envelopes)
  • VertexSecret::parse error uses generic JSON-shape hint; does NOT echo raw secret bytes
  • BridgeContext.deadline threaded into chat() via with_deadline
  • Header-injection guard on access_token value (HeaderValue::from_str rejects newlines)

References

Test plan

38 unit tests, all passing:

Publisher resolution (8 tests, preserved from skeleton)

gemini-, claude-, meta/llama-, llama, mistral-, codestral-, jamba-*; case-insensitive; URL segment for each; Meta::url_segment() == None.

VertexSecret parsing (4 tests)

Full form, empty rejected, non-JSON rejected, raw secret bytes do NOT leak into error messages (pinned by vertex_secret_error_does_not_leak_secret_content).

URL token validation (3 tests)

Canonical IDs pass; URL-injection chars (/, ?, #, \\n, ..) rejected; empty rejected.

Gemini request body translation (5 tests)

User turn → role=user; assistant → role=model (NOT "assistant"); system → top-level systemInstruction (NOT in contents[]); multiple system messages concatenated; generationConfig only emitted when set.

Gemini response translation (4 tests)

STOP → Stop; MAX_TOKENS → Length; safety reasons (SAFETY / RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter; missing usageMetadata.

Pre-dispatch validation (6 tests)

Unknown publisher, non-Google publisher named in error, invalid secret, missing model_name, chat_ignores_req_model_and_uses_ctx_model_name, chat_stream references D5.2.b.

Bridge dispatch via bridge.chat() end-to-end against wiremock (8 tests)

  • URL path exactly /v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent
  • Body uses Gemini wire shape (camelCase, no model/stream)
  • System role lifted to top-level systemInstruction
  • Assistant role uses model, NOT assistant
  • Authorization header carries Bearer <access_token> verbatim
  • 4xx upstream error body redacted (does NOT echo project id)
  • Path injection in project field rejected before URL stitching
  • MAX_TOKENS finishReason → FinishReason::Length

Test plan TODO

  • cargo test -p aisix-provider-vertex → 38/38 pass
  • cargo clippy --workspace --all-targets -- -D warnings clean
  • cargo fmt --check clean
  • CI: cargo test --workspace
  • CI: cargo clippy
  • CI: cargo fmt --check

Summary by CodeRabbit

  • New Features

    • Vertex AI provider now supports functional Gemini chat operations with proper request/response translation, authentication handling, and error management.
  • Tests

    • Added comprehensive test coverage for publisher resolution, credential validation, request routing, and error handling scenarios.
  • Documentation

    • Updated crate-level documentation with current implementation status and feature notes.

Review Change Stack

…(D5.2.a, #302 Phase E)
Replaces the skeleton's `BridgeError::Config("not yet implemented")`
stubs in `aisix-provider-vertex` with real Vertex AI dispatch for the
`google` publisher (Gemini chat). Other Vertex publishers
(`anthropic.*`, `meta.*`, `mistral.*`, `ai21.*`) surface a clear
publisher-named "not yet implemented — D5.3/D5.4" error.
`chat_stream()` returns "streaming not yet implemented — D5.2.b"
for all publishers.
## Wire shape pinned (Gemini on Vertex)
Per https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini:
- URL: `POST https://<region>-aiplatform.googleapis.com/v1/projects/<project>/locations/<region>/publishers/google/models/<model>:generateContent`
- Body: Gemini's `generateContent` JSON shape:
1. `contents[].role` is `"user"` or `"model"` (NOT `"assistant"`)
2. `contents[].parts[].text` (not `content`)
3. `systemInstruction` is top-level — system messages do NOT
appear in `contents[]` (Gemini 400s otherwise)
4. `generationConfig.temperature` / `topP` / `maxOutputTokens`
(camelCase, emitted only when set)
- Auth: `Authorization: Bearer <access_token>` where `access_token`
comes from the operator-supplied pre-minted GCP OAuth2 token
## Credentials convention
`ProviderKey.secret` is a JSON-encoded `{access_token, project,
region}` blob. The `access_token` is operator-managed — they refresh
it before the GCP TTL (~1 hour) and re-PUT the ProviderKey. D5.1
follow-up will add in-process token mint via `yup-oauth2` /
`gcp_auth` from a `service_account_json` field.
`project`, `region`, and `upstream_id` (model name from
`ctx.model.model_name`) are validated for URL-control chars before
being interpolated into the path: `/`, `?`, `#`, whitespace, `..`
all rejected at registration time so a malicious model_name can't
redirect dispatch or corrupt metrics labels.
## Audit lessons applied proactively
This PR was written after the D6 audit (#319) and D7.2.a audit (#320)
returned their reports. The lessons baked in from the start:
- `bridge.chat()` is exercised end-to-end via wiremock + a
`#[cfg(test)] api_base_override` seam — credentials, URL stitching,
body shaping all run normally; only the destination host is
different.
- 4xx error body redacted to canned status-keyed phrases (does NOT
echo Vertex's `Permission denied on project my-proj-prod-123`
envelope into customer-visible errors).
- `VertexSecret::parse` error message uses generic JSON-shape hint;
does NOT echo raw secret bytes (serde error messages can leak
partial content). Pinned by `vertex_secret_error_does_not_leak_secret_content`.
- `BridgeContext.deadline` threaded into `chat()` via `with_deadline`.
- Header-injection guard on `access_token` value (`HeaderValue::from_str`
rejects newlines / NULs); same for `request_id`.
## Test coverage
38 unit tests, all passing:
### Publisher resolution (8 tests, preserved from skeleton)
gemini-*, claude-*, meta/llama-*, llama*, mistral-*, codestral-*,
jamba-*; case-insensitive on model name; URL segment for each
publisher; `Meta::url_segment() == None` (Llama uses OpenAPI shim).
### `VertexSecret` parsing (4 tests)
Full form, empty rejected, non-JSON rejected with generic shape
error, raw secret bytes do NOT leak into error messages.
### URL token validation (3 tests)
Canonical IDs pass (`my-proj-prod-123`, `us-central1`,
`europe-west4`, `gemini-1.5-pro`); URL-injection chars
(`/`, `?`, `#`, `\n`, `..`) rejected; empty rejected.
### Gemini request body translation (5 tests)
User turn → role=user; assistant turn → role=model (NOT
"assistant"); system messages → top-level `systemInstruction`
(NOT in contents[]); multiple system messages concatenated with
`\n\n`; `generationConfig` only emitted when at least one field
is set.
### Gemini response translation (4 tests)
Text + STOP finishReason → ChatResponse + FinishReason::Stop;
MAX_TOKENS → FinishReason::Length; safety reasons (SAFETY /
RECITATION / BLOCKLIST / PROHIBITED_CONTENT / SPII) → ContentFilter;
missing usageMetadata → 0 tokens.
### Pre-dispatch validation (6 tests)
Unknown publisher, non-Google publisher named in error, invalid
secret, missing model_name, `chat_ignores_req_model_and_uses_ctx_model_name`
(D6 audit HIGH-1 regression carried over), `chat_stream` not-implemented
error references D5.2.b.
### Bridge dispatch via `bridge.chat()` end-to-end against wiremock (8 tests)
- URL path exactly `/v1/projects/my-proj/locations/us-central1/publishers/google/models/gemini-1.5-pro:generateContent`
- Body uses Gemini wire shape (camelCase, no `model` field, no `stream`)
- System role lifted to top-level `systemInstruction`
- Assistant role uses `model`, NOT `assistant`
- Authorization header carries `Bearer <access_token>` verbatim
- 4xx upstream error body redacted (does NOT echo project id)
- Path injection in project field rejected before URL stitching
- MAX_TOKENS finishReason maps to FinishReason::Length
## References
- Vertex AI REST API — https://cloud.google.com/vertex-ai/docs/reference/rest
- Gemini generateContent — https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini
- Gemini finishReason enum — https://ai.google.dev/api/generate-content#FinishReason
- Google Gemini Python SDK — https://github.com/google-gemini/generative-ai-python
CopilotAI review requested due to automatic review settings May 17, 2026 13:18
@coderabbitai

coderabbitaiBot commented May 17, 2026

Copy link
Copy Markdown
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: c73ab2bb-4b20-4257-a8de-7b3071294d97

📥 Commits

Reviewing files that changed from the base of the PR and between fdc5f52 and 67bad57.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (3)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs
  • crates/aisix-provider-vertex/src/lib.rs

📝 Walkthrough

Walkthrough

This PR transforms the Vertex provider bridge from a skeleton into a fully functional Gemini chat implementation, including HTTP client integration, publisher dispatch, Gemini request/response translation, comprehensive security validation, and integration tests with mock HTTP endpoints.

Changes

Vertex Gemini Bridge

Layer / File(s)Summary
Dependencies and Bridge Structure
crates/aisix-provider-vertex/Cargo.toml, crates/aisix-provider-vertex/src/bridge.rs
Add reqwest, serde, tokio, http, and wiremock dependencies; update VertexBridge struct to own a reqwest::Client; introduce VertexPublisher enum with Google and Meta variants; add default_client construction with fixed user-agent.
Publisher Resolution and Naming
crates/aisix-provider-vertex/src/bridge.rs
Implement VertexPublisher::from_upstream_id to resolve publishers from upstream model_name prefixes; add publisher-to-name mapping for error messages; include unit tests for publisher resolution and URL-segment pinning behavior.
Dispatch Support Functions
crates/aisix-provider-vertex/src/bridge.rs
Parse ProviderKey.secret JSON credentials without echoing in errors; validate URL path tokens to reject control/traversal characters; map HTTP status responses to customer-visible BridgeError variants with retry-after headers; add optional deadline/timeout support. Include unit tests for secret parsing, redaction guarantees, and URL-token validation.
Chat Dispatch Entry Point
crates/aisix-provider-vertex/src/bridge.rs
Implement Bridge::chat to resolve publisher from context.model.model_name, dispatch Google to chat_gemini, and return "not yet implemented" for other publishers. Include validation tests for unknown publishers, invalid secrets, missing model_name, model resolution precedence, and chat_stream not-yet-implemented messaging.
Gemini Request/Response Implementation
crates/aisix-provider-vertex/src/bridge.rs
Translate gateway ChatFormat to Gemini generateContent wire shape (system → systemInstruction, assistant → "model", conditional generation config); parse Gemini responses with token usage and finish-reason mapping; validate bearer tokens, build safe headers, fail fast for empty contents, support request deadlines. Comprehensive wiremock integration tests validate URL construction, request authorization, body shape, role mapping, security (path-injection rejection, error-redaction), and edge cases (system-only messages, MAX_TOKENS finish reason).
Crate Documentation
crates/aisix-provider-vertex/src/lib.rs
Replace skeleton status with Phase E feature checklist; document pre-minted access_token expectation; clarify publisher resolution flow; add Vertex AI REST and Gemini API references.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

… + LOW-4)
PR #321 audit surfaced 1 MEDIUM + 4 LOW; this commit addresses the
MEDIUM + 2 of the 4 LOW items that have concrete code fixes. LOW-1
(generationConfig missing topK/stopSequences) and LOW-2 (usageMetadata
missing cachedContentTokenCount/thoughtsTokenCount) deferred as
scope-extension follow-ups.
## MEDIUM (fixed)
MEDIUM-1 — `BridgeError::Config` no longer formats the underlying
`InvalidHeaderValue` Display into the customer-visible message.
The `http` crate's current Display impl is opaque, but it's an
implementation detail; a future change including the offending
byte position would leak partial bearer-token content. Since the
bytes being validated ARE the customer's bearer token, defense in
depth wins. Pinned by two new tests
(`header_invalid_access_token_error_does_not_leak_bytes` +
`header_invalid_request_id_error_does_not_leak_bytes`).
## LOW (fixed)
LOW-3 — `finishReason` IMAGE_SAFETY and LANGUAGE now map to
`FinishReason::ContentFilter` (previously fell through to Stop,
misleading tracing — a customer's dashboard would show a
successful "stop" when Google in fact filtered the response).
Added to the existing safety-finish-reason loop test.
LOW-4 — `chat_gemini` now fails fast with a clear Config error
when `contents[]` ends up empty (system-only messages, since
system lifts to top-level `systemInstruction`). Gemini's schema
requires `contents` non-empty; without this guard, Vertex 400s
with a generic envelope and the customer sees the canned
"upstream returned 400" phrase. Pinned by
`chat_gemini_with_system_only_messages_fails_fast` using
`Mock::expect(0)` to catch a regression that leaks the request
to dispatch.
## Deferred (justified)
LOW-1 — `generationConfig` missing `topK` / `stopSequences` /
`responseMimeType` / `candidateCount` / `seed` /
`thinkingConfig`. Out of scope for the minimal D5.2.a; will
land alongside the `ChatFormat.extra` pass-through follow-up
that lifts these fields uniformly across all bridges.
LOW-2 — `usageMetadata` missing `cachedContentTokenCount` and
`thoughtsTokenCount`. Needs `UsageStats` schema additions
(cross-bridge change) — will land alongside D5.5 metrics work.
## Verification
cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit
regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean
@moonming

Copy link
Copy Markdown
MemberAuthor

Audit follow-up pushed (67bad57): Audit returned 0 HIGH + 1 MEDIUM + 4 LOW. This commit addresses MEDIUM-1 + 2 of the 4 LOW items.

MEDIUM (fixed)

  • MEDIUM-1BridgeError::Config no longer formats the underlying InvalidHeaderValue Display into the customer-visible message. Defense in depth against a future http-crate Display impl change that might include the offending byte position (which would leak partial bearer-token content). Pinned by two new tests (header_invalid_access_token_error_does_not_leak_bytes + header_invalid_request_id_error_does_not_leak_bytes).

LOW (fixed)

  • LOW-3finishReasonIMAGE_SAFETY and LANGUAGE now map to FinishReason::ContentFilter (previously fell through to Stop, misleading tracing). Added to the existing safety-finish-reason test loop.
  • LOW-4chat_gemini fails fast with a clear Config error when contents[] ends up empty (system-only messages). Gemini's schema requires contents non-empty; without this guard, Vertex 400s and the customer sees the canned "upstream returned 400". Pinned by chat_gemini_with_system_only_messages_fails_fast using Mock::expect(0).

Deferred (justified)

  • LOW-1generationConfig missing topK / stopSequences / responseMimeType / candidateCount / seed / thinkingConfig. Out of scope for the minimal D5.2.a; will land alongside the ChatFormat.extra pass-through follow-up.
  • LOW-2usageMetadata missing cachedContentTokenCount and thoughtsTokenCount. Needs UsageStats schema additions (cross-bridge change); will land alongside D5.5 metrics work.

Verification

cargo test -p aisix-provider-vertex → 41 passed (was 38; +3 audit regression tests)
cargo clippy --workspace --all-targets -- -D warnings → clean
cargo fmt --check → clean

@moonming
moonming merged commit 83fad3d into mainMay 17, 2026
8 checks passed
@moonming
moonming deleted the feat/vertex-wire branch May 17, 2026 13:35
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming