Uh oh!
There was an error while loading. Please reload this page.
feat(managed): spawn /dp/heartbeat worker at boot - #31
Merged
Conversation
There was a problem hiding this comment.
Pull request overview
Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.
Changes:
- Introduces a new
heartbeatmodule implementing periodicPOST /dp/heartbeatwith interval clamping and wiremock-based tests. - Extends
mainmanaged-mode bootstrap to deriveHeartbeatConfigfrom either registration response (first boot) or persisteddp_id+cp_base_url(subsequent boots). - Spawns the heartbeat task during startup and awaits it during shutdown.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.
| File | Description |
|---|---|
crates/aisix-server/src/main.rs | Computes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk. |
crates/aisix-server/src/heartbeat.rs | New worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+270
to
+271
| if let Some(task) = heartbeat_task { | ||
| let _ = task.await; |
Comment on lines
+95
to
+98
| _ = cancel.changed() => { | ||
| if *cancel.borrow() { | ||
| tracing::info!("heartbeat shutting down"); | ||
| return; |
Comment on lines
+90
to
+104
| match send(&client, &cfg, uptime).await { | ||
| Ok(()) => tracing::debug!("heartbeat ok"), | ||
| Err(e) => tracing::warn!(error = %e, "heartbeat failed"), | ||
| } | ||
| } | ||
| _ = cancel.changed() => { | ||
| if *cancel.borrow() { | ||
| tracing::info!("heartbeat shutting down"); | ||
| return; | ||
| } | ||
| } | ||
| } | ||
| } | ||
| } | ||
| Ok(h) => Some(h), | ||
| Err(e) => { | ||
| tracing::warn!(error = %e, | ||
| "managed mode: heartbeat worker disabled (dp_id unreadable)"); |
DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client). ## Behaviour After the managed-mode bootstrap either registers or confirms an existing bundle on disk, `main` spawns one tokio task that POSTs `/dp/heartbeat` on a fixed interval. The task: - ticks at the interval returned by the register response (clamped to [5s, 300s] as defence against a buggy CP reply) - uses MissedTickBehavior::Delay so a slow tick doesn't burst catch-up beats afterwards - logs individual failures as warnings and keeps running; a transient CP outage means "no dashboard update" but not "DP stops trying" - cancels via the shared `watch::Receiver<bool>` so graceful shutdown drains the in-flight request Request body: `{ dp_id, uptime_seconds, version }`. Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to mTLS once cp-api terminates mTLS). ## Boot paths (both covered) 1. **First boot**: register returns `Registered` with heartbeat_url + dp_id + interval; these flow straight into `HeartbeatConfig`. 2. **Subsequent boot**: bundle already on disk; `dp_id` is read from `managed.dp_id_file` and the URL is synthesised from `managed.cp_base_url` with a default 15s interval. Failure to read dp_id → worker disabled with a warning, not a hard boot failure (the DP should still proxy traffic). ## Files - `crates/aisix-server/src/heartbeat.rs` (new) - `HeartbeatConfig` + `sanitised()` interval clamping - `spawn()` / `run()` / `send()` split so tests can drive each step - HTTP client built inside the worker (per-instance, not shared — the worker owns its lifetime) - `crates/aisix-server/src/main.rs` - `mod heartbeat;` - New pre-etcd block constructing `heartbeat_cfg: Option<_>` - `heartbeat_task: Option<JoinHandle>` awaited at the end of run() - `load_heartbeat_config_from_disk()` helper for the "bundle exists from a prior boot" branch ## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean) - `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts on the Authorization header + body fields the CP handler uses. - `heartbeat::tests::send_propagates_non_success_body` — CP error body surfaces into the anyhow chain so operators see which error code fired without decoding logs. - `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe a successful beat, flip cancel, assert the task returns inside a 2s grace window. - `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s and 86400s → 300s as sanity bounds. ## Explicitly out of scope (tracked) - `/dp/telemetry` (next PR) - Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2) - mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api mTLS listener once that lands)
moonmingforce-pushed
the
feat/dp-heartbeat
branch
from
April 23, 2026 12:31
e519423 to
ede0a69CompareUh oh!
There was an error while loading. Please reload this page.
4 tasks
moonming added a commit
that referenced
this pull request
Apr 23, 2026
The same Docker image now serves both standalone and managed (aisix.cloud tenant) deployments. Two pieces: - config.managed.yaml — bootstrap template baked at /etc/aisix/config.managed.yaml. Has placeholder etcd endpoint (overwritten by /dp/register response), managed.enabled = true, and unbindable admin (defence-in-depth if managed mode somehow flipped off). All real per-DP secrets come from env vars. - docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH (default /etc/aisix/config.yaml). Standalone users mount their config at the default path; managed users point AISIX_CONFIG_PATH at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN + AISIX_MANAGED__CP_BASE_URL. Existing main.rs bootstrap (PR #30 + #31) already does the rest: register-and-persist on first boot, reload bundle on subsequent boots, spawn heartbeat worker. Tests: parses_managed_block_with_register_fields locks the YAML shape so any new required field on ManagedConfig fails CI loudly instead of silently breaking the image. Docs: docs/managed-mode.md walks operators through first boot, restart semantics, env-var override matrix, and common errors. This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness can now `docker run` aisix with a deployment_token and have the DP register itself without prebaked certs.
8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (
/dp/registerclient); retargets to main when #30 merges.After managed-mode bootstrap completes (either register or confirm-existing-bundle),
mainspawns one tokio task that POSTs/dp/heartbeaton a fixed interval for the life of the process.Request
Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.
Worker lifecycle
[5s, 300s]as defence against a buggy CP reply (0 = burst loop, a week = useless).MissedTickBehavior::Delay— a slow beat doesn't burst catch-up beats afterwards.warn!and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".tokio::select!on the sharedwatch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.Boot paths
HeartbeatConfigRegistered { heartbeat_url, dp_id, heartbeat_interval }from the register responsedp_idread frommanaged.dp_id_file; URL =managed.cp_base_url + /dp/heartbeat; default interval 15sIf
dp_id_fileis unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".Files
crates/aisix-server/src/heartbeat.rs(new, ~220 lines) — config +spawn/run/send+ testscrates/aisix-server/src/main.rs—mod heartbeat, pre-etcdheartbeat_cfgblock,heartbeat_taskawaited alongsidewatch_taskat shutdown,load_heartbeat_config_from_diskhelperTests (
cargo test --workspacegreen,cargo clippy -D warningsclean)send_posts_dp_id_and_bearersend_propagates_non_success_bodyDP_NOT_FOUNDetc. without decoding logsrun_stops_on_cancelsanitised_interval_clamps_extremesExplicitly out of scope
/dp/telemetry(next PR)