Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets - #451

Merged
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class
Aug 24, 2026
Merged

Add system-node-critical priorityClass to health-monitoring-agent DaemonSets#451
zhaoqizqwang merged 1 commit into
aws:mainfrom
huaziyao:hma-priority-class

Conversation

@huaziyao

@huaziyaohuaziyao commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator

What's changing and why?

Adds a priorityClassName (defaulting to the Kubernetes built-in system-node-critical) to both health-monitoring-agent DaemonSets:

  • health-monitoring-agent (NVIDIA / GPU)
  • health-monitoring-agent-non-nvidia (Trainium / Inferentia)

The value is configurable via .Values.priorityClassName and falls back to system-node-critical when unset.

Why: Under node DiskPressure, HMA was evicted moments before a GPU failure occurred — so the failure went undetected and unremediated. system-node-critical (priority value 2000001000) makes HMA among the last pods evicted under node pressure, after training and other monitoring pods, preserving health monitoring exactly when it matters most. This mirrors what EKSNodeMonitoringAgent already does. Because system-node-critical is a built-in PriorityClass, no additional PriorityClass object needs to be created.

Before / After UX

  • Before: HMA pods run at priority 0. Under node pressure (e.g. DiskPressure), they can be evicted early — potentially before or during a hardware failure — leaving the node unmonitored.
  • After: HMA pods run at priority 2000001000 (system-node-critical) by default, so the kubelet evicts them last, keeping health monitoring alive during node pressure. Operators can override or disable via .Values.priorityClassName.

How was this change tested?

Validated through render, lint, dry-run, and a live apply/observe/revert on a real cluster:

  1. helm template — rendered both DaemonSets and confirmed priorityClassName: system-node-critical appears in each pod spec; verified --set priorityClassName=<x> propagates correctly, with fallback to system-node-critical when unset.
  2. helm lint — passes clean.
  3. Server-side dry-runkubectl apply --server-side --dry-run accepted with no schema/admission errors.
  4. Live apply → observe → revert on an e2e test cluster (~10 nodes, 4 HMA pods): after apply, all HMA pods rolled out cleanly and reported priority: 2000001000 (up from 0); reverting returned them to priority: 0 with no priorityClassName, again via a clean rollout. No DaemonSet disruption during either transition.

Are unit tests added?

No — this is a Helm chart template/values change with no application code. Coverage is provided by the helm template / helm lint render checks above.

Are integration tests added?

No new automated integration tests. Verified manually via the live apply/observe/revert described above.

Reviewer Guidelines

‼️Merge Requirements: PRs with failing integration tests cannot be merged without justification.

One of the following must be true:

  • All automated PR checks pass
  • Failed tests include local run results/screenshots proving they work
  • Changes are documentation-only

…monSets
Under node DiskPressure, HMA was evicted moments before a GPU failure went
undetected and unremediated (P438078245). Setting priorityClassName to the
built-in system-node-critical makes HMA the last pod evicted (after training
and other monitoring pods), preserving health monitoring during node pressure.
Mirrors what EKSNodeMonitoringAgent already does.
Applied to both the health-monitoring-agent (NVIDIA) and
health-monitoring-agent-non-nvidia (Trainium/Inferentia) DaemonSets.
Configurable via .Values.priorityClassName, defaulting to system-node-critical.
@huaziyao
huaziyao requested a review from a team as a code ownerAugust 21, 2026 22:43
@kethang-sm

Copy link
Copy Markdown
Contributor

LGTM

@zhaoqizqwang
zhaoqizqwang merged commit 78ce39c into aws:mainAug 24, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@huaziyao@kethang-sm@zhaoqizqwang