## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

## fix: resolve MLflow app discovery issues - #5924

Merged
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests
Jun 4, 2026
Merged

## fix: resolve MLflow app discovery issues#5924
lucasjia-aws merged 13 commits into
aws:masterfrom
lucasjia-aws:nova_tests

Conversation

@lucasjia-aws

@lucasjia-awslucasjia-aws commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes multiple test failures in sagemaker-train-integ-tests that emerged after the MTRL Launch PR (#5919) was merged. The root causes are: (1) an incorrect response key in MLflow app discovery, (2) deleted/hard-coded MLflow app ARNs in tests, and (3) a missing None check in feature processor lineage.

Changes

Test / ComponentProblemFix
All evaluator integ tests (benchmark, llm_as_judge, custom_scorer)_resolve_mlflow_resource_arn in finetune_utils.py reads page.get("MlflowApps", []) but the list_mlflow_apps API returns apps under "Summaries". The function always sees an empty list, tries to create a new app, and fails on quota.Change the key from "MlflowApps" to "Summaries" in _resolve_mlflow_resource_arn.
test_llm_as_judge_base_model_fixHard-coded mlflow_tracking_server_arn (app-W7FOBBXZANVX) no longer exists in the test account.Add a mlflow_resource_arn pytest fixture in conftest.py that auto-discovers an existing ready MLflow app or creates a temporary one (with cleanup). Both tests in this file use the fixture.
test_mtrl_trainer_integration
test_mtrl_evaluator
test_mtrl_evaluator_3p_agent
test_multi_turn_rl_trainer_integration
Hard-coded MLflow app ARN (app-ZG6FYITNGMMU) was deleted from the test account.Replace with existing app-O4ZGQYBYHMRH (mtrl-integ-test) which is in Created state in the same account.
_feature_processor_lineage.pyAttributeError when pipeline_version_context is None during lineage update.Add None check before accessing pipeline_version_context attributes.
Unit tests for MLflow fixTests were mocking the old "MlflowApps" response key.Fix "MlflowApps""Summaries" in MLflow unit test mocks.

Testing

  • sagemaker-train unit tests: all passing locally
  • sagemaker-train-integ-tests: MLflow-related failures resolved; remaining failures are infra flaky (AlgorithmError) or quota-limited MTRL eval jobs (pre-existing, unrelated)

…resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
…ates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
@lucasjia-aws

lucasjia-aws commented Jun 4, 2026

Copy link
Copy Markdown
CollaboratorAuthor

serve integ test succeeded in a previous commit: https://github.com/aws/sagemaker-python-sdk/actions/runs/26923043569/job/79444479419, and the following commits have nothing to do with serve module.

mujtaba1747
mujtaba1747 previously approved these changes Jun 4, 2026
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
@lucasjia-awslucasjia-aws changed the title ## fix: resolve cross-region singleton bug and MLflow app discovery issues## fix: resolve MLflow app discovery issuesJun 4, 2026
rsareddy0329
rsareddy0329 previously approved these changes Jun 4, 2026
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
@lucasjia-aws
lucasjia-aws merged commit b86c2ac into aws:masterJun 4, 2026
13 of 19 checks passed
@lucasjia-aws
lucasjia-aws deleted the nova_tests branch June 5, 2026 18:04
guanweim pushed a commit to guanweim/sagemaker-python-sdk that referenced this pull request Jun 15, 2026
* fix: bypass SageMakerClient singleton for cross-region model package resolution
The SageMakerClient singleton caches the first region it is initialized with and ignores subsequent region parameters. This causes Nova integ tests (which run in us-east-1) to fail when the singleton was already created with us-west-2 by an earlier test in the same process.
Errors observed:
- ModelPackageGroup arn:aws:sagemaker:us-west-2:784379639078:model-package-group/sdk-test-finetuned-models does not exist
- DescribeModelPackage: ARN should be scoped to correct region: us-west-2
Fix: use session.boto_session.client("sagemaker") directly instead of ModelPackageGroup.get() / ModelPackage.get() in the three call sites that resolve model package resources. This respects the session's actual region without depending on the singleton's cached state.
* test: update unit tests
* fix: handle missing pipeline version context in lineage update
_update_pipeline_lineage assumed the version context always exists.
When it's been deleted or never created (e.g. prior run failure),
DescribeContext throws ResourceNotFound. Now catches the error and
recreates the version context with proper associations.
* fix(test): add mlflow_resource_arn fixture that auto-discovers or creates app
Replace hard-coded MLflow app ARN with a conftest fixture that finds an
existing ready app or creates a temporary one (cleaned up after tests).
Prevents failures when the hard-coded app is deleted or quota is full.
X-AI-Prompt: add self-healing mlflow fixture for llm_as_judge integ tests
X-AI-Tool: kiro-cli
* fix(test): use correct response key "Summaries" for list_mlflow_apps API
* mark two slow tests as not serial
* fix: use correct response key "Summaries" in _resolve_mlflow_resource_arn
* replace not-existing mlflow app
* refactor: use session.sagemaker_client instead of boto_session.client
Per SDK coding standards, avoid calling boto3 directly. Use the
session's sagemaker_client attribute which already has the correct
region bound at session creation time.
* revert: remove SageMakerClient singleton bypass from feature code
* test: mark TestLLMAsJudgeBaseModelFix as serial
Tests share the same pipeline definition and conflict when run in
parallel (Pipeline has been modified since your last read).
X-AI-Prompt: mark llm_as_judge_base_model_fix as serial
X-AI-Tool: kiro-cli
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@lucasjia-aws@mujtaba1747@rsareddy0329