feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
Authorization: Bearer <dp_id>
Content-Type: application/json
{ "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

ScenarioSource of HeartbeatConfig
First boot (registration just ran)Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk)dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

TestCovers
send_posts_dp_id_and_bearerwiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_bodyCP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancelspawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

CopilotAI review requested due to automatic review settings April 23, 2026 09:36

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

FileDescription
crates/aisix-server/src/main.rsComputes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rsNew worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to mainApril 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).
## Behaviour
After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:
- ticks at the interval returned by the register response (clamped
to [5s, 300s] as defence against a buggy CP reply)
- uses MissedTickBehavior::Delay so a slow tick doesn't burst
catch-up beats afterwards
- logs individual failures as warnings and keeps running; a
transient CP outage means "no dashboard update" but not "DP
stops trying"
- cancels via the shared `watch::Receiver<bool>` so graceful
shutdown drains the in-flight request
Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).
## Boot paths (both covered)
1. **First boot**: register returns `Registered` with heartbeat_url +
dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
`managed.dp_id_file` and the URL is synthesised from
`managed.cp_base_url` with a default 15s interval.
Failure to read dp_id → worker disabled with a warning, not a
hard boot failure (the DP should still proxy traffic).
## Files
- `crates/aisix-server/src/heartbeat.rs` (new)
- `HeartbeatConfig` + `sanitised()` interval clamping
- `spawn()` / `run()` / `send()` split so tests can drive each step
- HTTP client built inside the worker (per-instance, not shared —
the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
- `mod heartbeat;`
- New pre-etcd block constructing `heartbeat_cfg: Option<_>`
- `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
- `load_heartbeat_config_from_disk()` helper for the "bundle exists
from a prior boot" branch
## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)
- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
body surfaces into the anyhow chain so operators see which error
code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
a successful beat, flip cancel, assert the task returns inside a
2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
and 86400s → 300s as sanity bounds.
## Explicitly out of scope (tracked)
- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
mTLS listener once that lands)
@moonming
moonming merged commit b597129 into mainApr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:
- config.managed.yaml — bootstrap template baked at
/etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
(overwritten by /dp/register response), managed.enabled = true,
and unbindable admin (defence-in-depth if managed mode somehow
flipped off). All real per-DP secrets come from env vars.
- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
(default /etc/aisix/config.yaml). Standalone users mount their
config at the default path; managed users point AISIX_CONFIG_PATH
at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
AISIX_MANAGED__CP_BASE_URL.
Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.
Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.
Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.
This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming