sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

Description

@akwoprosper

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

Describe the bug

sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

Two observable consequences:

  1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
  2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

To reproduce

Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

pip install 'sagemaker-core==2.15.0'
fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
role_type="training",
)

Result:

RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
ecr:GetDownloadUrlForLayer

Confirm the verdict is an org-layer artifact rather than a real permission gap:

aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
--action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
{
"EvalActionName": "cloudwatch:PutMetricData",
"EvalDecision": "implicitDeny",
"MatchedStatements": [],
"OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
// ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
}

The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

Expected behavior

The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

Screenshots or logs

.../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
570 # Permission check (definitive denial blocks; unverifiable warns)
571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
572 if verdict is False:
> 573 raise RoleValidationError(
574 _build_validation_error_message(role_arn, role_type, miss...

System information

  • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
  • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
  • Framework version: N/A
  • Python version: 3.12
  • CPU or GPU: N/A (fails before job submission)
  • Custom Docker image (Y/N): N/A

Additional context

Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

Relevant docs:

Workarounds (both verified):

  1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
  2. Pin sagemaker-core<2.15.0, which predates the pre-check.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

      Description

      @akwoprosper

      PySDK Version

      • PySDK V2 (2.x)
      • PySDK V3 (3.x)

      Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

      Describe the bug

      sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

      Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

      Two observable consequences:

      1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
      2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

      2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

      To reproduce

      Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

      pip install 'sagemaker-core==2.15.0'
      fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
      role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
      role_type="training",
      )

      Result:

      RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
      Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
      ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
      ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
      ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
      ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
      ecr:GetDownloadUrlForLayer
      

      Confirm the verdict is an org-layer artifact rather than a real permission gap:

      aws iam simulate-principal-policy \
      --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
      --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
      {
      "EvalActionName": "cloudwatch:PutMetricData",
      "EvalDecision": "implicitDeny",
      "MatchedStatements": [],
      "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
      // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
      }

      The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

      Expected behavior

      The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

      Screenshots or logs

      .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
      570 # Permission check (definitive denial blocks; unverifiable warns)
      571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
      572 if verdict is False:
      > 573 raise RoleValidationError(
      574 _build_validation_error_message(role_arn, role_type, miss...
      

      System information

      • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
      • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
      • Framework version: N/A
      • Python version: 3.12
      • CPU or GPU: N/A (fails before job submission)
      • Custom Docker image (Y/N): N/A

      Additional context

      Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

      Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

      Relevant docs:

      Workarounds (both verified):

      1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
      2. Pin sagemaker-core<2.15.0, which predates the pre-check.

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

          Description

          @akwoprosper

          PySDK Version

          • PySDK V2 (2.x)
          • PySDK V3 (3.x)

          Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

          Describe the bug

          sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

          Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

          Two observable consequences:

          1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
          2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

          2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

          To reproduce

          Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

          pip install 'sagemaker-core==2.15.0'
          fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
          role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
          role_type="training",
          )

          Result:

          RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
          Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
          ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
          ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
          ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
          ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
          ecr:GetDownloadUrlForLayer
          

          Confirm the verdict is an org-layer artifact rather than a real permission gap:

          aws iam simulate-principal-policy \
          --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
          --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
          {
          "EvalActionName": "cloudwatch:PutMetricData",
          "EvalDecision": "implicitDeny",
          "MatchedStatements": [],
          "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
          // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
          }

          The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

          Expected behavior

          The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

          Screenshots or logs

          .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
          570 # Permission check (definitive denial blocks; unverifiable warns)
          571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
          572 if verdict is False:
          > 573 raise RoleValidationError(
          574 _build_validation_error_message(role_arn, role_type, miss...
          

          System information

          • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
          • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
          • Framework version: N/A
          • Python version: 3.12
          • CPU or GPU: N/A (fails before job submission)
          • Custom Docker image (Y/N): N/A

          Additional context

          Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

          Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

          Relevant docs:

          Workarounds (both verified):

          1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
          2. Pin sagemaker-core<2.15.0, which predates the pre-check.

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

              Description

              @akwoprosper

              PySDK Version

              • PySDK V2 (2.x)
              • PySDK V3 (3.x)

              Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

              Describe the bug

              sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

              Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

              Two observable consequences:

              1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
              2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

              2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

              To reproduce

              Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

              pip install 'sagemaker-core==2.15.0'
              fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
              role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
              role_type="training",
              )

              Result:

              RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
              Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
              ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
              ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
              ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
              ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
              ecr:GetDownloadUrlForLayer
              

              Confirm the verdict is an org-layer artifact rather than a real permission gap:

              aws iam simulate-principal-policy \
              --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
              --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
              {
              "EvalActionName": "cloudwatch:PutMetricData",
              "EvalDecision": "implicitDeny",
              "MatchedStatements": [],
              "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
              // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
              }

              The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

              Expected behavior

              The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

              Screenshots or logs

              .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
              570 # Permission check (definitive denial blocks; unverifiable warns)
              571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
              572 if verdict is False:
              > 573 raise RoleValidationError(
              574 _build_validation_error_message(role_arn, role_type, miss...
              

              System information

              • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
              • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
              • Framework version: N/A
              • Python version: 3.12
              • CPU or GPU: N/A (fails before job submission)
              • Custom Docker image (Y/N): N/A

              Additional context

              Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

              Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

              Relevant docs:

              Workarounds (both verified):

              1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
              2. Pin sagemaker-core<2.15.0, which predates the pre-check.

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

                  Description

                  @akwoprosper

                  PySDK Version

                  • PySDK V2 (2.x)
                  • PySDK V3 (3.x)

                  Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

                  Describe the bug

                  sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

                  Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

                  Two observable consequences:

                  1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
                  2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

                  2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

                  To reproduce

                  Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

                  pip install 'sagemaker-core==2.15.0'
                  fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
                  role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
                  role_type="training",
                  )

                  Result:

                  RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
                  Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
                  ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
                  ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
                  ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
                  ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
                  ecr:GetDownloadUrlForLayer
                  

                  Confirm the verdict is an org-layer artifact rather than a real permission gap:

                  aws iam simulate-principal-policy \
                  --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
                  --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
                  {
                  "EvalActionName": "cloudwatch:PutMetricData",
                  "EvalDecision": "implicitDeny",
                  "MatchedStatements": [],
                  "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
                  // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
                  }

                  The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

                  Expected behavior

                  The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

                  Screenshots or logs

                  .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
                  570 # Permission check (definitive denial blocks; unverifiable warns)
                  571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
                  572 if verdict is False:
                  > 573 raise RoleValidationError(
                  574 _build_validation_error_message(role_arn, role_type, miss...
                  

                  System information

                  • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
                  • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
                  • Framework version: N/A
                  • Python version: 3.12
                  • CPU or GPU: N/A (fails before job submission)
                  • Custom Docker image (Y/N): N/A

                  Additional context

                  Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

                  Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

                  Relevant docs:

                  Workarounds (both verified):

                  1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
                  2. Pin sagemaker-core<2.15.0, which predates the pre-check.

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

                      Description

                      @akwoprosper

                      PySDK Version

                      • PySDK V2 (2.x)
                      • PySDK V3 (3.x)

                      Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

                      Describe the bug

                      sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

                      Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

                      Two observable consequences:

                      1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
                      2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

                      2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

                      To reproduce

                      Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

                      pip install 'sagemaker-core==2.15.0'
                      fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
                      role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
                      role_type="training",
                      )

                      Result:

                      RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
                      Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
                      ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
                      ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
                      ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
                      ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
                      ecr:GetDownloadUrlForLayer
                      

                      Confirm the verdict is an org-layer artifact rather than a real permission gap:

                      aws iam simulate-principal-policy \
                      --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
                      --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
                      {
                      "EvalActionName": "cloudwatch:PutMetricData",
                      "EvalDecision": "implicitDeny",
                      "MatchedStatements": [],
                      "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
                      // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
                      }

                      The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

                      Expected behavior

                      The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

                      Screenshots or logs

                      .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
                      570 # Permission check (definitive denial blocks; unverifiable warns)
                      571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
                      572 if verdict is False:
                      > 573 raise RoleValidationError(
                      574 _build_validation_error_message(role_arn, role_type, miss...
                      

                      System information

                      • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
                      • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
                      • Framework version: N/A
                      • Python version: 3.12
                      • CPU or GPU: N/A (fails before job submission)
                      • Custom Docker image (Y/N): N/A

                      Additional context

                      Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

                      Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

                      Relevant docs:

                      Workarounds (both verified):

                      1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
                      2. Pin sagemaker-core<2.15.0, which predates the pre-check.

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

                          Description

                          @akwoprosper

                          PySDK Version

                          • PySDK V2 (2.x)
                          • PySDK V3 (3.x)

                          Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

                          Describe the bug

                          sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

                          Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

                          Two observable consequences:

                          1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
                          2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

                          2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

                          To reproduce

                          Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

                          pip install 'sagemaker-core==2.15.0'
                          fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
                          role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
                          role_type="training",
                          )

                          Result:

                          RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
                          Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
                          ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
                          ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
                          ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
                          ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
                          ecr:GetDownloadUrlForLayer
                          

                          Confirm the verdict is an org-layer artifact rather than a real permission gap:

                          aws iam simulate-principal-policy \
                          --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
                          --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
                          {
                          "EvalActionName": "cloudwatch:PutMetricData",
                          "EvalDecision": "implicitDeny",
                          "MatchedStatements": [],
                          "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
                          // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
                          }

                          The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

                          Expected behavior

                          The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

                          Screenshots or logs

                          .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
                          570 # Permission check (definitive denial blocks; unverifiable warns)
                          571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
                          572 if verdict is False:
                          > 573 raise RoleValidationError(
                          574 _build_validation_error_message(role_arn, role_type, miss...
                          

                          System information

                          • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
                          • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
                          • Framework version: N/A
                          • Python version: 3.12
                          • CPU or GPU: N/A (fails before job submission)
                          • Custom Docker image (Y/N): N/A

                          Additional context

                          Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

                          Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

                          Relevant docs:

                          Workarounds (both verified):

                          1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
                          2. Pin sagemaker-core<2.15.0, which predates the pre-check.

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              sagemaker-core 2.15.0: role validation raises false-positive RoleValidationError under condition-based SCPs (IAM simulator can't evaluate conditional SCPs) #6019

                              Description

                              @akwoprosper

                              PySDK Version

                              • PySDK V2 (2.x)
                              • PySDK V3 (3.x)

                              Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).

                              Describe the bug

                              sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...)TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role_evaluate_permissionsiam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).

                              Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.

                              Two observable consequences:

                              1. Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
                              2. The same role works fine from a notebook / via a direct create_training_job call, because those paths don't run this client-side pre-check.

                              2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.

                              To reproduce

                              Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.

                              pip install 'sagemaker-core==2.15.0'
                              fromsagemaker.core.helper.iam_role_resolverimportIamRoleResolver# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.IamRoleResolver().resolve_and_validate_role(
                              role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
                              role_type="training",
                              )

                              Result:

                              RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
                              Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
                              ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
                              ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
                              ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
                              ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
                              ecr:GetDownloadUrlForLayer
                              

                              Confirm the verdict is an org-layer artifact rather than a real permission gap:

                              aws iam simulate-principal-policy \
                              --policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
                              --action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
                              {
                              "EvalActionName": "cloudwatch:PutMetricData",
                              "EvalDecision": "implicitDeny",
                              "MatchedStatements": [],
                              "OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
                              // ...same for the other actions, including sagemaker:CreateTrainingJob —// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
                              }

                              The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.

                              Expected behavior

                              The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.

                              Screenshots or logs

                              .../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
                              570 # Permission check (definitive denial blocks; unverifiable warns)
                              571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
                              572 if verdict is False:
                              > 573 raise RoleValidationError(
                              574 _build_validation_error_message(role_arn, role_type, miss...
                              

                              System information

                              • SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag v3.15.0)
                              • Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
                              • Framework version: N/A
                              • Python version: 3.12
                              • CPU or GPU: N/A (fails before job submission)
                              • Custom Docker image (Y/N): N/A

                              Additional context

                              Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.

                              Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.

                              Relevant docs:

                              Workarounds (both verified):

                              1. Attach an explicit Deny on iam:SimulatePrincipalPolicy to the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow).
                              2. Pin sagemaker-core<2.15.0, which predates the pre-check.

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions