Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) by KeitaW · Pull Request #398 · aws/sagemaker-hyperpod-cli · GitHub
Skip to content

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra) - #398

Merged
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles
Apr 1, 2026
Merged

Add MIG profile support for ml.p6-b300.48xlarge (Blackwell Ultra)#398
nargokul merged 1 commit into
aws:mainfrom
KeitaW:feat/add-p6-b300-mig-profiles

Conversation

@KeitaW

@KeitaWKeitaW commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py with the B300 MIG profiles: mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
  • Add 17 B300-specific MIG partition profiles (7 uniform + 10 mixed) to the Helm chart default-mig-config.yaml ConfigMap

Relationship to #396

PR #396 ("Added profiles for B300") was merged on 2026-03-23 and added 2 ConfigMap profiles (all-1g.67gb and mixed-2-1g.34gb-1-2g.67gb-1-3g.135gb). However, it left two critical gaps:

1. constants.py was not updated — MIG requests on B300 are rejected before the ConfigMap is ever consulted.

_validate_accelerator_partition_parameters() in accelerator_partition_util.py checks INSTANCE_TYPE_MIG_PROFILES at line 26 as a gate. Because ml.p6-b300.48xlarge is absent from that dict, the CLI returns:

"Instance type 'ml.p6-b300.48xlarge' does not support accelerator partitions."

This blocks all MIG usage on B300 — HyperPodPyTorchJob submissions, inference endpoints with acceleratorPartitionType, and hyp list-accelerator-partition-type. The ConfigMap profiles from #396 are unreachable.

2. 15 of 17 ConfigMap profiles are missing.

Cross-referencing against the NVIDIA GPU Operator v25.3.0 upstream ConfigMap (B300 section, device-filter 0x318210DE) and the NVIDIA MIG product page (Blackwell Ultra: 7x34GB, 4x69GB, 2x139GB, 1x279GB):

ProfileUpstreamAfter #396This PR
all-1g.34gb (x7)YesMissingAdded
all-1g.67gb (x4)YesAdded
all-2g.67gb (x3)YesMissingAdded
all-3g.135gb (x2)YesMissingAdded
all-4g.135gb (x1)YesMissingAdded
all-7g.269gb (x1)YesMissingAdded
10 mixed profilesYes1 of 10+9 added

Profile Source

MIG profiles are derived from the NVIDIA GPU Operator upstream ConfigMap (v25.3.0), which defines B300 profiles under the # B300 comment section with all-balanced device-filter 0x318210DE. The NVIDIA MIG User Guide (r580) has not been updated for B300 yet.

Additional Note

The existing p6-b200.48xlarge key in INSTANCE_TYPE_MIG_PROFILES is missing the ml. prefix (unlike all other entries). This PR does not address that issue to keep scope focused, but it may warrant a separate fix.

Test plan

  • Verify INSTANCE_TYPE_MIG_PROFILES['ml.p6-b300.48xlarge'] returns the correct 6 profiles
  • Verify ALLOWED_ACCELERATOR_PARTITION_TYPES includes all B300 MIG types (mig-1g.34gb, mig-1g.67gb, mig-2g.67gb, mig-3g.135gb, mig-4g.135gb, mig-7g.269gb)
  • Verify default-mig-config.yaml parses as valid YAML
  • Verify _validate_accelerator_partition("mig-1g.34gb", ..., "ml.p6-b300.48xlarge") passes validation
  • Integration test: deploy a MIG-enabled instance group with ml.p6-b300.48xlarge and nvidia.com/mig.config: all-1g.34gb

@KeitaW
KeitaW requested a review from a team as a code ownerMarch 27, 2026 12:21
Add ml.p6-b300.48xlarge to INSTANCE_TYPE_MIG_PROFILES in constants.py
with the correct B300 MIG profiles derived from the NVIDIA GPU Operator
v25.3.0 upstream ConfigMap (device-filter 0x318210DE):
- mig-1g.34gb, mig-1g.67gb, mig-2g.67gb
- mig-3g.135gb, mig-4g.135gb, mig-7g.269gb
Also add the corresponding uniform and mixed MIG partition profiles
to the Helm chart default-mig-config.yaml ConfigMap, following the
same pattern used for existing GPU types (H100, H200, B200).
The B300 GPU (288GB HBM3e, ~269GB usable) was already registered in
INSTANCE_RESOURCES but had no MIG profile mapping, causing HyperPod
MIG validation to reject accelerator partition requests on this
instance type.
@KeitaW
KeitaWforce-pushed the feat/add-p6-b300-mig-profiles branch from 045470a to c98fd6eCompareMarch 28, 2026 00:06
KeitaW added a commit to KeitaW/sagemaker-hyperpod-cli that referenced this pull request Mar 28, 2026
Covers ml.p6-b300.48xlarge MIG profile support added in PR aws#398:
- Profile presence in INSTANCE_TYPE_MIG_PROFILES
- Complete profile list verification (6 profiles)
- All profiles in ALLOWED_ACCELERATOR_PARTITION_TYPES
- GPU slice extraction for all B300 profiles (1g→1, 2g→2, ..., 7g→7)
- CPU/memory default calculation for each profile at max instances
- Validation acceptance for valid B300 profiles
- Validation rejection for invalid profiles on B300 instance type
@nargokul
nargokul merged commit a234d67 into aws:mainApr 1, 2026
1 of 2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@KeitaW@nargokul@mx26pol