feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E) - #446

Merged
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex
May 29, 2026
Merged

feat(vertex): Llama + OpenAI-compatible MaaS family via the openapi shim (#302 Phase E)#446
moonming merged 3 commits into
mainfrom
feat/llama-on-vertex

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the OpenAI chat-completions shim rail on Vertex. Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss, MiniMax, Moonshot, and Z.ai through a single endpoint endpoints/openapi/chat/completions (model id in the body, not a publishers/<vendor>/...:rawPredict URL). One handler ⇒ the whole MaaS family.

  • VertexPublisher: old unimplemented MetaOpenAiCompat; from_upstream_id resolves meta/|llama|deepseek|qwen|openai/gpt-oss|minimaxai/|moonshotai/|zai-org/ to it. Mistral/AI21 stay separate (:rawPredict, deferred).
  • chat_openai_shim (non-stream) + chat_openai_shim_stream (streaming): POST the OpenAI body to the shim with the GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the OpenAI request serializer + response/stream decoders verbatim (new dep on aisix-provider-openai) so the wire matches direct OpenAI. Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
  • Model id KEPT in the body (the shim keys off it) — opposite of the Gemini/Anthropic publisher paths.

Reference-impl + upstream

Tests

  • cargo test -p aisix-provider-vertex — 77 pass (resolver maps the full MaaS family to OpenAiCompat; wiremock dispatch test pins openapi URL + Bearer + model-in-body; stream-not-implemented test repointed off Llama to the still-deferred rawPredict publishers).
  • cargo clippy --all-targets -- -D warnings + cargo check --workspace clean.

Scope / follow-ups

  • Live e2e (mock-vertex openapi shim + spec, Llama representative) lands in the AISIX-Cloud companion PR.
  • Claude-on-Vertex streaming (:streamRawPredict, D5.3) + Mistral/AI21 (D5.4) remain.

🤖 Generated with Claude Code

…him (#302 Phase E, D5.4)
Vertex Model Garden serves Meta Llama, DeepSeek, Qwen, gpt-oss,
MiniMax, Moonshot, and Z.ai through a single OpenAI chat-completions
shim at `endpoints/openapi/chat/completions` (the model id rides in the
request body, NOT a `publishers/<vendor>/...:rawPredict` URL). This
wires that one rail, so the whole family dispatches through it.
- `VertexPublisher`: the old (unimplemented) `Meta` variant becomes
`OpenAiCompat`, and `from_upstream_id` resolves the documented MaaS
prefixes (`meta/` | `llama` | `deepseek` | `qwen` | `openai/gpt-oss` |
`minimaxai/` | `moonshotai/` | `zai-org/`) to it. Mistral / AI21 stay
separate (they use `:rawPredict`, still deferred).
- `chat_openai_shim` (non-stream) + `chat_openai_shim_stream`
(streaming): POST the OpenAI chat body to the openapi shim with the
GCP OAuth2 Bearer (same token path as Gemini/Claude). Reuses the
OpenAI request serializer + response/stream-chunk decoders verbatim
(new dep on aisix-provider-openai) so the wire matches direct OpenAI.
Streaming reuses the shared SseDecoder ([DONE]-terminated OpenAI SSE).
- The model id is KEPT in the body (the shim keys off it) — opposite of
the Gemini/Anthropic publisher paths which strip it from the body and
carry it in the URL.
Reference impl confirms the shim endpoint + the OpenAI-handler family
grouping + the OpenAI body/response wire; Google's Vertex Llama docs
confirm the `endpoints/openapi/chat/completions` path.
Tests: resolver maps the whole MaaS family to OpenAiCompat (+ Mistral/
AI21 stay on their own arms); wiremock dispatch test pins the openapi
URL + Bearer + model-in-body; the stream-not-implemented test repointed
off Llama (now wired) to the still-deferred rawPredict publishers.
77 unit tests pass; clippy + workspace check clean.
Live e2e (mock-vertex openapi shim + spec) lands in the AISIX-Cloud
companion PR. Streaming for Claude-on-Vertex (`:streamRawPredict`, D5.3)
+ Mistral/AI21 (D5.4) remain.
@coderabbitai

coderabbitaiBot commented May 29, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 14 minutes and 43 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: 5ca3fa45-8c3c-4fb9-8b35-71af4fde44c0

📥 Commits

Reviewing files that changed from the base of the PR and between 3100fc5 and a443801.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (2)
  • crates/aisix-provider-vertex/Cargo.toml
  • crates/aisix-provider-vertex/src/bridge.rs

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…stream)
The non-stream openapi-shim dispatch had a wire test pinning the URL,
Bearer, and model-in-body; the streaming handler had none. Add a
streaming wire test (capturing responder + OpenAI chat.completion.chunk
SSE + [DONE]) that pins the openapi URL, Bearer auth, `stream:true` in
the body, the model id staying in the body (never the URL), and that the
SSE frames decode into ChatChunks whose aggregated content and terminal
finish_reason match the upstream stream.
@moonming
moonming merged commit b3a6992 into mainMay 29, 2026
8 checks passed
@jarvis9443
jarvis9443 deleted the feat/llama-on-vertex branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@moonming