feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(server): startup sequence + health listeners - #5

Merged
moonming merged 1 commit into
mainfrom
feat/server-bootstrap
Apr 17, 2026
Merged

feat(server): startup sequence + health listeners#5
moonming merged 1 commit into
mainfrom
feat/server-bootstrap

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Wires the 10-step startup sequence from spec §1 into the `aisix`
binary:

  1. Parse CLI args (`--config`, with `AISIX_CONFIG` env fallback)
  2. Load + validate bootstrap config
  3. `init_tracing` — `aisix-obs` now installs an `EnvFilter` subscriber
  4. `EtcdConfigProvider::connect` (5s × 5 retry, reused from PR feat(core): Model, ApiKey, RateLimit entities + JSON Schema validators #3)
  5. Initial snapshot load via `Supervisor::load_once`
  6. Spawn `Supervisor::run` on a dedicated task (cancel via `watch`)
  7. Build proxy router (`aisix-proxy` now exposes `build_router`)
  8. Build admin router (`aisix-admin` now exposes `build_router`)
  9. Bind + serve both listeners with `axum::serve` + graceful shutdown
  10. SIGINT/SIGTERM → flip cancel → drain serves → join supervisor

Both routers only mount `/health` for now so the bootstrap is
observable without dragging in feature code that hasn't landed yet.
The health handler reports current snapshot table sizes so the full
path is end-to-end verifiable.

Test plan

  • `cargo test --workspace` — 75 tests pass (new + existing)
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • `cargo build --bin aisix` succeeds
  • CI green on all 6 jobs

Wires the 10-step startup sequence from spec §1 into the aisix binary:
1. Parse CLI args (--config, with AISIX_CONFIG env fallback)
2. Load + validate bootstrap config
3. init_tracing — aisix-obs now installs an EnvFilter-backed subscriber
4. EtcdConfigProvider::connect (5s × 5 retry, reused from PR #3)
5. Initial snapshot load via Supervisor::load_once
6. Spawn Supervisor::run on a dedicated task (cancel via watch channel)
7. Build proxy router (aisix-proxy now exposes build_router + ProxyState)
8. Build admin router (aisix-admin now exposes build_router + AdminState)
9. Bind + serve both listeners with axum::serve + graceful shutdown
10. SIGINT/SIGTERM → flip cancel channel → drain both serves → join supervisor
For now both routers only mount /health so the startup wiring is
observable without reaching into feature code that hasn't landed yet.
The health handler reports the current snapshot table sizes so the
bootstrap path is end-to-end verifiable.
Ancillary changes:
- ObsError with Filter + AlreadyInitialised variants (thiserror)
- aisix-obs drops #![forbid(unsafe_code)] so test code can use the
now-unsafe std::env::remove_var without a lint override
- ProxyState / AdminState are Clone and cheap (Arc-backed handles and
Arc<[String]> admin keys)
6 new unit tests (CLI parsing, ObsError display, proxy health JSON
shape, admin health JSON shape, admin-keys Arc sharing) land on top of
the existing 68 — 75 passing workspace-wide.
CopilotAI review requested due to automatic review settings April 17, 2026 05:53
@moonming
moonming merged commit 5692500 into mainApr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/server-bootstrap branch April 17, 2026 05:57

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the spec §1 startup sequence wiring in the aisix server binary, introducing minimal /health endpoints for both proxy and admin listeners so bootstrap + snapshot plumbing can be exercised end-to-end.

Changes:

  • Adds an async aisix-server startup flow: CLI config path, config load, tracing init, etcd supervisor spawn, and dual Axum listeners with graceful shutdown.
  • Exposes build_router + state types in aisix-proxy and aisix-admin, each mounting a basic /health handler reporting snapshot table sizes.
  • Introduces aisix-obs::init_tracing to install a tracing_subscriber with an EnvFilter derived from config and/or RUST_LOG.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/main.rsImplements the server startup orchestration, supervisor wiring, listener binding/serving, and shutdown coordination.
crates/aisix-proxy/src/lib.rsAdds ProxyState and build_router() with a /health endpoint and tests.
crates/aisix-proxy/Cargo.tomlAdds tokio dev-dependency features needed for async tests.
crates/aisix-obs/src/lib.rsAdds init_tracing() and error types for observability bootstrap.
crates/aisix-admin/src/lib.rsAdds AdminState and build_router() with a /health endpoint and tests.
crates/aisix-admin/Cargo.tomlAdds tokio dev-dependency features needed for async tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +38 to +43
let filter = EnvFilter::try_from_default_env()
.or_else(|_| EnvFilter::try_new(&cfg.log_level))
.map_err(|source| ObsError::Filter {
directive: cfg.log_level.clone(),
source,
})?;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

EnvFilter::try_from_default_env().or_else(|_| ...) also falls back when RUST_LOG is set but invalid, silently ignoring the operator override. If RUST_LOG is present but can’t be parsed, it’s usually better to return an error that references the invalid RUST_LOG value (and only fall back to cfg.log_level when the env var is absent).

Copilot uses AI. Check for mistakes.
Comment on lines +83 to +92
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));

let (proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokio::join!(proxy_serve, admin_serve) can deadlock: if either server future returns early (e.g. bind/serve error) while the other is still running, join! will wait forever and the cancellation cancel_tx.send(true) is never reached. Use tokio::select! (or spawn both serves and select on their JoinHandles) to trigger cancellation as soon as any of (signal, proxy_serve, admin_serve) completes, then gracefully shut down the remaining tasks.

Suggested change
let signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
let(proxy_res, admin_res) = tokio::join!(proxy_serve, admin_serve);
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;
let _ = watch_task.await;
letmut signal_task = tokio::spawn(wait_for_signal(cancel_tx.clone()));
letmut proxy_task = tokio::spawn(proxy_serve);
letmut admin_task = tokio::spawn(admin_serve);
let first_result: anyhow::Result<()> = tokio::select! {
res = &mut signal_task => {
res.map_err(|e| anyhow::anyhow!("signal task join error: {e}"))?;
Ok(())
}
res = &mut proxy_task => {
res.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))
}
res = &mut admin_task => {
res.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?
.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))
}
};
// Ask the supervisor and remaining servers to stop.
let _ = cancel_tx.send(true);
// The signal task may still be waiting for SIGINT/SIGTERM if a server
// completed first, so do not wait forever on it.
if !signal_task.is_finished(){
signal_task.abort();
}
let _ = signal_task.await;
if !proxy_task.is_finished(){
let proxy_res = proxy_task
.await
.map_err(|e| anyhow::anyhow!("proxy task join error: {e}"))?;
proxy_res.map_err(|e| anyhow::anyhow!("proxy serve error: {e}"))?;
}
if !admin_task.is_finished(){
let admin_res = admin_task
.await
.map_err(|e| anyhow::anyhow!("admin task join error: {e}"))?;
admin_res.map_err(|e| anyhow::anyhow!("admin serve error: {e}"))?;
}
let _ = watch_task.await;
first_result?;

Copilot uses AI. Check for mistakes.

// Ask the supervisor to stop (no-op if the signal task already did).
let _ = cancel_tx.send(true);
let _ = signal_task.await;

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

signal_task.await can hang indefinitely when shutdown is triggered by something other than SIGINT/SIGTERM (e.g. one of the serve futures errors/exits). wait_for_signal doesn’t observe the cancel channel, so it never returns in that scenario. Consider either (1) passing a watch::Receiver<bool> into wait_for_signal and select!ing between cancel and OS signals, or (2) aborting/dropping signal_task once cancellation is triggered.

Suggested change
let _ = signal_task.await;
// If shutdown was triggered by something other than SIGINT/SIGTERM,
// the signal task may still be blocked waiting on OS signals.
// Abort it so this join cannot hang indefinitely.
signal_task.abort();
match signal_task.await{
Ok(()) => {}
Err(err)if err.is_cancelled() => {}
Err(err) => returnErr(anyhow::anyhow!("signal task join error: {err}")),
}

Copilot uses AI. Check for mistakes.
Comment on lines +57 to +66
let supervisor = Arc::new(Supervisor::new(provider, cfg.etcd.prefix.clone()));
let snapshot_handle = supervisor.handle();

let (cancel_tx, cancel_rx) = watch::channel(false);
let watch_task = tokio::spawn(supervisor.clone().run(cancel_rx.clone()));

// Steps 7-8: routers.
let proxy_router =
aisix_proxy::build_router(ProxyState::new(snapshot_handle.clone(), &cfg.proxy));
let admin_router =

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup sequence comment and PR description call for bootstrapping an initial snapshot before serving, but the code starts serving immediately after spawning Supervisor::run. Since Supervisor::run performs the first load_all asynchronously, /health (and future proxy routing) can observe an empty snapshot during startup. Call supervisor.load_once().await? (or otherwise await the first successful load) before binding/serving.

Copilot uses AI. Check for mistakes.
let state = AdminState::new(handle, &cfg());
let b = state.clone();
assert_eq!(b.admin_keys.len(), 1);
assert_eq!(&*b.admin_keys[0], "k1");

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assertion does not compile: b.admin_keys[0] is an &String, so &*b.admin_keys[0] attempts to move out of a borrow. Compare using b.admin_keys[0].as_str() (or &b.admin_keys[0] with an appropriate RHS) instead.

Suggested change
assert_eq!(&*b.admin_keys[0],"k1");
assert_eq!(b.admin_keys[0].as_str(),"k1");

Copilot uses AI. Check for mistakes.
//! PRs so this crate stays focused.

#![forbid(unsafe_code)]
#![deny(rust_2018_idioms)]

CopilotAIApr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This crate drops #![forbid(unsafe_code)] while the rest of the workspace consistently forbids unsafe (e.g. aisix-core, aisix-etcd, aisix-admin, aisix-proxy). If that wasn’t intentional, re-add #![forbid(unsafe_code)] to keep the workspace policy consistent.

Suggested change
#![deny(rust_2018_idioms)]
#![deny(rust_2018_idioms)]
#![forbid(unsafe_code)]

Copilot uses AI. Check for mistakes.
moonming added a commit that referenced this pull request May 18, 2026
…ped reads, redact 5xx message, Vertex content-type guard
Five concrete fixes from the Copilot inline review on PR #323. Two
stale comments (#3, #4 — already fixed in commit 3) are skipped.
**#1+#7 — Azure OpenAI-compatible code preservation.**
Azure's envelope omits `error.type` and carries only `error.code`.
The bridge previously put the upstream code into `view.kind` and
left `view.code` as `None`. For OpenAI-compat tokens Azure inherits
unchanged (e.g. `rate_limit_exceeded`), this meant downstream OpenAI
clients received `error.type=rate_limit_exceeded` but
`error.code=null` — exactly the SDK-retry break issue #322 is about.
Fix:
- Azure parser populates BOTH `view.kind` AND `view.code` from the
upstream `error.code` field.
- `render_openai_envelope`'s AzureOpenAI branch now prefers the
translation-table-derived code (so explicit Azure tokens like
`DeploymentNotFound` → `model_not_found` still win), falling back
to `view.code` for OpenAI-compat pass-through.
**#2 — Drain the response stream after hitting the cap.**
`read_body_capped` previously broke out of the read loop the moment
`limit` bytes were buffered. With reqwest/hyper that leaves unread
bytes in the response and prevents connection reuse — during a burst
of upstream errors the gateway would churn TCP connections instead
of recycling the keep-alive pool. Fix: keep iterating the stream,
discarding chunks past the cap. Memory stays bounded by `limit`.
**#5 — Redact upstream `error.message` on 5xx.**
The 5xx branch of `render_bridge_upstream_envelope` was forwarding
`BridgeError::UpstreamStatus.message` verbatim — which for OpenAI /
Anthropic comes from the parsed upstream `error.message`. Upstream
5xx bodies routinely embed operator-internal detail (engine names,
shard ids, queue depth). Fix: on 5xx, emit a canned
`"upstream returned {status}"` message; the full upstream body
remains in operator logs via tracing.
**#6 — Stale "follow-up" comment.**
The docstring on `render_bridge_upstream_envelope` claimed cross-wire
translation would ship in a follow-up, but it already shipped in
commit 2. Rewrite the comment to describe current behaviour
(4xx → `error_translate`; 5xx → canned envelope; `Unknown` wire →
legacy generic envelope).
**#8 — Content-type guard on Vertex (and Azure, while at it).**
`capture_upstream_error_http` already gates serde parsing on
`Content-Type: application/json` so a 64 KB HTML error page from a
fronting WAF doesn't waste CPU on a doomed JSON parse. The Vertex
and Azure bridges call serde directly because they need a custom
parse path (canned message for redaction) — same guard now applies.
Promoted `content_type_is_json` and added a `response_is_json`
helper to the gateway's public surface; both bridges call it before
`parse_*_error_*`.
New tests:
- `upstream_openai_5xx_with_json_envelope_collapses_and_redacts_message`
pins the 5xx redaction (asserts `engine offline` / `shard 47` /
`engine_overloaded` don't reach the customer envelope).
- `chat_429_preserves_openai_compatible_code_for_sdk_retry` (Azure)
pins that `parsed.code` carries the OpenAI-compat upstream code.
- `chat_400_non_json_body_skips_envelope_parse` (Azure) and
`chat_gemini_non_json_body_skips_envelope_parse` (Vertex) pin the
new content-type guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@moonmingmoonming mentioned this pull request May 21, 2026
5 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming