feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsBuilds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.tomlAdds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rsIntroduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rsRenders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rsMounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rsDefines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rsImplements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rsAdds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.tomlAdds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lockLocks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
ifletSome(rest) = s.strip_prefix("Bearer "){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
ifletSome((scheme, rest)) = s.split_once(' '){
if scheme.eq_ignore_ascii_case("bearer"){
let rest = rest.trim();
if rest.is_empty(){
returnErr(ProxyError::MissingAuth);
}
returnOk(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming