feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(managed): one-shot /dp/register client + atomic cert persistence - #30

Merged
moonming merged 1 commit into
mainfrom
feat/dp-register
Apr 23, 2026
Merged

feat(managed): one-shot /dp/register client + atomic cert persistence#30
moonming merged 1 commit into
mainfrom
feat/dp-register

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the prd-09 §9.3.5/dp/register wire contract. Paired with the cp-api handler at api7/AISIX-Cloud#9 — both sides implement against the spec committed in api7/AISIX-Cloud d44bfc0.

Boot-time flow

managed.enabled = true
├─ bundle already on disk? → skip registration, proceed with existing cert
└─ no bundle AND token+cp_base_url set
└─ POST /dp/register
├─ write { ca.crt, client.crt, client.key } to mtls_dir (0600, atomic)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.{endpoints,tls} so the regular connect path
works as if the user had configured it by hand

New pieces

crates/aisix-server/src/register.rs (private module)

Not a new crate — the registration client is a single-use helper tied to main's boot. Two public entry points:

  • register_and_persist(&ManagedConfig) -> anyhow::Result<Registered> — full roundtrip
  • bundle_exists(mtls_dir) -> bool — idempotent boot check

Internal helpers:

  • call_register()reqwest POST with 20s timeout, User-Agent: aisix-dp/<semver>, non-2xx surfaces upstream body text (truncated to 300 chars) so operators see WHICH error code the CP returned.
  • persist_mtls() / persist_dp_id()tokio::fs writes via write_atomic().
  • write_atomic() — tmp-file + fsync + rename. Permissions applied on .mode() at open time so a concurrent reader never sees the wider mode window. cfg(unix) only; a stub panics on non-Unix targets (managed mode is Unix-only per prd-09).
  • hostname::get() — inline libc::gethostname wrapper under cfg(unix) so we don't pull in the hostname crate just for this.

aisix-core::ManagedConfig additions

FieldPurpose
registration_token: Option<String>Single-use token from the create-gateway response on cp-api. Empty = skip registration.
cp_base_url: Option<String>e.g. https://api.us.aisix.cloud. Required alongside the token.
mtls_dir: StringWhere to persist the cert bundle. Default /var/lib/aisix/mtls.
dp_id_file: StringWhere to persist dp_id. Default /var/lib/aisix/dp_id.
registration_enabled() helperTrue iff both token + URL set.

main.rs wiring

New block at the top of run() that invokes the register flow conditionally and then mutates cfg.etcd.{endpoints, tls} so the rest of the startup is oblivious to whether certs came from an out-of-band install or just-now registration.

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
register_and_persist_happy_pathFull wiremock server + tempfile dir; asserts file contents + 0600 permissions + dp_id persistence + interval decode
register_propagates_4xx_body401 from CP surfaces the exact upstream error code (INVALID_TOKEN) in the error string so operators don't stare at status-only messages
register_requires_token_and_urlTwo missing-field cases cleanly rejected
bundle_exists_detects_complete_setReturns false until ALL three PEM files are present — partial installs don't mask a failed earlier register

Relationship

SidePRSpec ref
cp-api serverapi7/AISIX-Cloud#9§9.3.5
DP client (this)#30§9.3.5

Both implementations land against the same PRD commit, so review independence is preserved — they'll merge in either order.

Explicitly out of scope (tracked for the heartbeat PR)

  • POST /dp/heartbeat worker. The Registered struct this PR returns already captures heartbeat_url + heartbeat_interval so the next PR spawns a tokio task over those values without another register roundtrip.
  • Local config snapshot so the DP serves cached config while etcd is unreachable (prd-09 §9.7.2).
  • mTLS cert rotation as the 10-year validity approaches — currently manual "revoke + re-register".

DP side of the prd-09 §9.3.5 wire contract. Paired with the cp-api
handler in api7/AISIX-Cloud#9.
## Boot-time behaviour
```
managed.enabled = true AND
bundle_exists(managed.mtls_dir) is false AND
managed.{registration_token, cp_base_url} both set
──► POST /dp/register
├─ write ca.crt / client.crt / client.key to mtls_dir (0600)
├─ write dp_id to dp_id_file
└─ override cfg.etcd.endpoints + cfg.etcd.tls so the regular
connect path sees the fresh cert
```
Any of the three conditions false → skip registration entirely.
Subsequent boots hit the "bundle already on disk" branch and go
straight to etcd connect.
## Design decisions
- **`register.rs` as a private module** (not a new crate): the
registration client is a single-use helper tied to `main`'s boot
sequence. A separate crate would need its own test harness for
little gain.
- **No `hostname` crate dependency** — inline `libc::gethostname` wrapper
under `cfg(unix)` inside a private submodule. Managed-mode is
Unix-only anyway (prd-09 assumes the DP runs on the user's server).
- **`write_atomic()`** uses the standard tmp-file-then-rename dance
plus explicit `fsync` before rename. Crashes between write and
rename leave either the old file or nothing — never a truncated
file. Permissions are set on open (0600) rather than chmod'd after
the rename so a concurrent reader never sees the wider mode.
- **Registration response schema** mirrors §9.3.5 exactly. The
`Registered` struct captures `heartbeat_*` + `telemetry_*` fields
even though this PR doesn't use them; the follow-up heartbeat PR
will wire them into the supervisor without another round-trip.
## Config additions (`ManagedConfig`)
- `registration_token: Option<String>` — single-use token from the
create-gateway response on cp-api.
- `cp_base_url: Option<String>` — e.g. `https://api.us.aisix.cloud`.
- `mtls_dir: String` — default `/var/lib/aisix/mtls`.
- `dp_id_file: String` — default `/var/lib/aisix/dp_id`.
- `registration_enabled()` helper — true iff both token + URL set.
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `register_and_persist_happy_path` — full wiremock server +
tempfile dir; asserts file contents AND `0600` permissions on each
of the three cert files AND dp_id persistence.
- `register_propagates_4xx_body` — 401 from CP surfaces the exact
upstream error code (INVALID_TOKEN) in the error string so operators
don't have to decode status-only messages.
- `register_requires_token_and_url` — table-driven for the two missing-
field cases.
- `bundle_exists_detects_complete_set` — returns false until all three
PEM files are present (doesn't accept partial installs).
## Explicitly out of scope (next PR)
- `POST /dp/heartbeat` worker (the `Registered.heartbeat_*` fields
this PR captures will drive it).
- Local config snapshot for etcd-disconnect resilience (prd-09 §9.7.2).
- Rotating mTLS bundle when it nears expiry (10y today; re-register
is a manual operator action for now).
CopilotAI review requested due to automatic review settings April 23, 2026 09:16

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode bootstrap flow to register against the control plane (POST /dp/register) at startup, persist the returned mTLS bundle + dp_id, and wire the returned etcd endpoint/certs into the existing etcd connect path.

Changes:

  • Introduces a one-shot registration client (register.rs) that calls /dp/register and persists {ca.crt, client.crt, client.key} + dp_id.
  • Wires managed boot logic into aisix-server startup to conditionally register and override cfg.etcd.{endpoints,tls}.
  • Extends ManagedConfig with registration token/CP URL and persistence path fields.

Reviewed changes

Copilot reviewed 4 out of 5 changed files in this pull request and generated 9 comments.

Show a summary per file
FileDescription
crates/aisix-server/src/register.rsNew registration client + persistence helpers + tests.
crates/aisix-server/src/main.rsStartup wiring to run registration and override etcd config when needed.
crates/aisix-server/Cargo.tomlAdds reqwest and test dependencies for the new module.
crates/aisix-core/src/config.rsAdds managed-mode registration/persistence settings to config schema.
Cargo.lockLocks new dependency additions.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +199 to +203
// Implement inline via libc.
mod hostname {
use std::ffi::{CStr, OsString};
use std::os::unix::ffi::OsStringExt;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hostname helper unconditionally uses std::os::unix::*, so this module won’t compile on non-Unix targets even though write_atomic has a cfg(not(unix)) stub. Consider gating the hostname module + gather_host_info() behind cfg(unix) and providing a small non-Unix fallback (e.g., return an explicit error) so cross-builds remain possible.

Copilot uses AI. Check for mistakes.
Comment on lines +268 to +273
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

write_atomic builds the tmp path via with_extension("{ext}.tmp"). For paths with no extension (e.g., dp_id_file default /var/lib/aisix/dp_id), this produces a double-dot filename like dp_id..tmp. Consider generating the tmp name by appending .tmp to the full filename (or using a random tmp name in the same dir) to avoid odd filenames and potential collisions.

Suggested change
let tmp = path.with_extension(format!(
"{}.tmp",
path.extension()
.map(|e| e.to_string_lossy().into_owned())
.unwrap_or_default()
));
let file_name = path
.file_name()
.ok_or_else(|| anyhow!("cannot create temporary path for {}", path.display()))?;
let tmp = path.with_file_name(format!("{}.tmp", file_name.to_string_lossy()));

Copilot uses AI. Check for mistakes.
Comment on lines +275 to +279
let mut f = tokio::fs::OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.mode(mode)

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security: the tmp file is opened with create(true) + truncate(true). If a previous crash leaves *.tmp behind with permissive mode (e.g., 0644), reopening it will not apply .mode(0o600) (mode is only used on creation) and the renamed private key can end up world-readable. Use create_new(true) with a unique tmp name (or delete any existing tmp before open) and/or explicitly set_permissions after open to guarantee 0600.

Copilot uses AI. Check for mistakes.
}
tokio::fs::rename(&tmp, path)
.await
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational durability: after rename(tmp, path), the parent directory isn’t fsynced. On crash/power loss, the rename can be lost even though the tmp file was fsynced, which undermines the “persisted atomically” guarantee for certs/keys. Consider fsyncing the containing directory after the rename to make the update durable.

Suggested change
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))?;
let parent_dir = path.parent().unwrap_or_else(|| Path::new(".")).to_path_buf();
tokio::task::spawn_blocking(move || -> anyhow::Result<()>{
let dir = std::fs::File::open(&parent_dir)
.with_context(|| format!("open directory {} for fsync", parent_dir.display()))?;
dir.sync_all()
.with_context(|| format!("fsync directory {}", parent_dir.display()))?;
Ok(())
})
.await
.context("join directory fsync task")??;

Copilot uses AI. Check for mistakes.
Comment on lines +228 to +244
async fn persist_mtls(
dir: &str,
bundle: &MTLSBundle,
) -> anyhow::Result<(PathBuf, PathBuf, PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;

let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

write_atomic(&ca, bundle.ca_certificate.as_bytes(), 0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(), 0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(), 0o600).await?;

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The mTLS bundle isn’t persisted atomically as a set: these three write_atomic calls commit each file independently. If the process crashes between them, the DP can be left with a partial bundle while the registration token has already been consumed, and next boot may be unable to recover. Consider writing all bundle files to staging (or *.tmp names) and only “commit” them together (e.g., via a marker file or directory rename).

Suggested change
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
tokio::fs::create_dir_all(&dir)
.await
.with_context(|| format!("create {}", dir.display()))?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");
write_atomic(&ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&key, bundle.private_key.as_bytes(),0o600).await?;
asyncfn commit_staged_bundle_dir(staging:&Path,target:&Path) -> anyhow::Result<()>{
match tokio::fs::rename(staging, target).await{
Ok(()) => Ok(()),
Err(err)if err.kind() == std::io::ErrorKind::AlreadyExists => {
let backup = target.with_extension(format!("bak-{}", std::process::id()));
match tokio::fs::metadata(&backup).await{
Ok(_) => {
tokio::fs::remove_dir_all(&backup)
.await
.with_context(|| format!("remove {}", backup.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", backup.display()));
}
}
tokio::fs::rename(target,&backup)
.await
.with_context(|| format!("rename {} -> {}", target.display(), backup.display()))?;
ifletErr(rename_err) = tokio::fs::rename(staging, target).await{
let _ = tokio::fs::rename(&backup, target).await;
returnErr(rename_err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display()));
}
let _ = tokio::fs::remove_dir_all(&backup).await;
Ok(())
}
Err(err) => Err(err)
.with_context(|| format!("rename {} -> {}", staging.display(), target.display())),
}
}
asyncfn persist_mtls(
dir:&str,
bundle:&MTLSBundle,
) -> anyhow::Result<(PathBuf,PathBuf,PathBuf)> {
let dir = PathBuf::from(dir);
let parent = dir
.parent()
.map(Path::to_path_buf)
.unwrap_or_else(|| PathBuf::from("."));
tokio::fs::create_dir_all(&parent)
.await
.with_context(|| format!("create {}", parent.display()))?;
let stage_name = format!(
".{}.tmp-{}",
dir.file_name().and_then(|name| name.to_str()).unwrap_or("mtls"),
std::process::id()
);
let staging = parent.join(stage_name);
match tokio::fs::metadata(&staging).await{
Ok(_) => {
tokio::fs::remove_dir_all(&staging)
.await
.with_context(|| format!("remove {}", staging.display()))?;
}
Err(e)if e.kind() == std::io::ErrorKind::NotFound => {}
Err(e) => {
returnErr(e).with_context(|| format!("stat {}", staging.display()));
}
}
tokio::fs::create_dir_all(&staging)
.await
.with_context(|| format!("create {}", staging.display()))?;
let staged_ca = staging.join("ca.crt");
let staged_cert = staging.join("client.crt");
let staged_key = staging.join("client.key");
write_atomic(&staged_ca, bundle.ca_certificate.as_bytes(),0o600).await?;
write_atomic(&staged_cert, bundle.certificate.as_bytes(),0o600).await?;
write_atomic(&staged_key, bundle.private_key.as_bytes(),0o600).await?;
commit_staged_bundle_dir(&staging,&dir).await?;
let ca = dir.join("ca.crt");
let cert = dir.join("client.crt");
let key = dir.join("client.key");

Copilot uses AI. Check for mistakes.
Comment on lines 111 to 113
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields, default)]
pub struct ManagedConfig {

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ManagedConfig derives Default, but the new mtls_dir / dp_id_file defaults are only applied during serde deserialization (via #[serde(default = ...)]). Calling ManagedConfig::default() will still produce empty strings for those paths, which is surprising for a public config type. Consider implementing Default manually so the struct’s Default matches the documented defaults.

Copilot uses AI. Check for mistakes.
Comment on lines +49 to +50
tempfile = "3"
wiremock = "0.6"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repo convention: other crates use workspace-pinned test deps (e.g., wiremock.workspace = true, tempfile.workspace = true). Consider switching these dev-dependencies to workspace to avoid version drift across crates.

Suggested change
tempfile = "3"
wiremock = "0.6"
tempfile.workspace = true
wiremock.workspace = true

Copilot uses AI. Check for mistakes.
Comment on lines +133 to +134
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs vs behavior: this comment says the mtls_dir parent directory must already exist, but persist_mtls() currently calls create_dir_all. Either update the documentation to match (directory will be created) or change the code to enforce the documented requirement.

Suggested change
/// written `0600`. Parent directory must already exist and be
/// writable by the aisix process user.
/// written `0600`. The directory will be created if it does not
/// already exist and must be writable by the aisix process user.

Copilot uses AI. Check for mistakes.
Comment on lines +378 to +382
use std::os::unix::fs::PermissionsExt;
assert_eq!(
m.permissions().mode() & 0o777,
0o600,
"file {p:?} perms wrong"

CopilotAIApr 23, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests are Unix-specific: this assertion uses std::os::unix::fs::PermissionsExt unconditionally, so the test module won’t compile on non-Unix targets. If cross-platform test builds matter, gate the permission checks (or the whole test) behind cfg(unix) and provide a non-Unix alternative.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bd4a4fe into mainApr 23, 2026
13 of 17 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-register branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming