bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

Description

@sagarneeldubey

Description

PySDK Version

  • PySDK V2 (2.x)
  • PySDK V3 (3.x)

Describe the bug

When using ModelBuilder with ModelServer.TGI, there is no working way to
load model weights from S3 — the deployed TGI container always pulls weights
from huggingface.co at startup. Both parameters that look like they should
enable S3 loading are silently ignored (neither raises at build time):

  1. model_path="s3://..." — the TGI build path calls
    _create_dir_structure(self.model_path), which runs
    Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
    this does not fail on an S3 URI; it silently creates a literal local
    directory tree named s3:/<bucket>/<prefix> in the working directory and
    continues. The S3 URI never reaches the container, and HF_MODEL_ID still
    resolves to the HF repo id.
    Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

  2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
    upload destination prefix for SDK scaffolding, not as a weight source. The
    build method sets self.s3_upload_path = None (comment: "TGI downloads
    models directly from HuggingFace Hub") and overwrites the attribute via
    self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
    with HF_MODEL_ID=<hf_repo_id> regardless.
    Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
from HF), which makes scale-to-zero async endpoints impractical and creates a
hard runtime dependency on huggingface.co. It also blocks deploying custom
fine-tuned weights that are not on the HF Hub at all.

This is closely related to #5529 (DJL Serving overwrites a user-provided
HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
this report is the TGI path and the required fix differs — see Additional
context
.

To reproduce

Build-only repro (no GPU, no deploy, ungated model so no token needed).
sagemaker==3.12.0.

frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
sample_output=[{"generated_text": "Hi"}],
)
# --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
model="Qwen/Qwen2.5-7B-Instruct",
model_path=S3_URI,
model_server=ModelServer.TGI,
schema_builder=schema,
role_arn=ROLE,
instance_type="ml.g5.2xlarge",
)
mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
model="Qwen/Qwen2.5-7B-Instruct",
s3_model_data_url=S3_URI,
model_server=ModelServer.TGI,
schema_builder=schema,
role_arn=ROLE,
instance_type="ml.g5.2xlarge",
)
mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

Expected behavior

With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
the SDK should not silently ignore an S3 value passed via model_path /
s3_model_data_url (either honor it or raise a clear error).

Screenshots or logs

Output of the repro above (SDK 3.12.0):

Case 1 — local junk dir created: True
Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct

For comparison, the working lower-level deployment logs this within ~3s of
container start (weights served from S3, no HF download):

INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
INFO download: Files are already present on the host. Skipping download.
INFO download: Successfully downloaded weights for /opt/ml/model

System information

  • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
  • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
  • Python version: 3.12
  • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
  • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

Additional context

Working workaround — bypass ModelBuilder and construct the resources
directly, attaching the S3 weights as an uncompressed ModelDataSource and
pointing HF_MODEL_ID at the mount path:

fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
)
fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
framework="huggingface-llm", region="us-east-1",
version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
)
primary=ContainerDefinition(
image=image_uri,
environment={
"HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
"SM_NUM_GPUS": "1",
"MESSAGES_API_ENABLED": "true",
},
model_data_source=ModelDataSource(
s3_data_source=S3ModelDataSource(
s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
s3_data_type="S3Prefix",
compression_type="None",
)
),
)
Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
Endpoint.create(...)

This drops cold start from ~12 min to ~3–5 min for a 7B model.

Relationship to existing issues

#5529#5588This issue
Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

Suggested fix (TGI path) — when an S3 URI is supplied:

  1. Skip _create_dir_structure local-dir creation for S3 inputs.
  2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
    the early branch in
    sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
    already emits the correct shape when model_path is an S3 URI, but it's
    currently unreachable because _create_dir_structure runs first.
  3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

More broadly, consider unifying S3-as-weight-source handling across all
model_server branches so behavior is consistent (ties together #5529, #5588,
and this issue).

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

      Description

      @sagarneeldubey

      Description

      PySDK Version

      • PySDK V2 (2.x)
      • PySDK V3 (3.x)

      Describe the bug

      When using ModelBuilder with ModelServer.TGI, there is no working way to
      load model weights from S3 — the deployed TGI container always pulls weights
      from huggingface.co at startup. Both parameters that look like they should
      enable S3 loading are silently ignored (neither raises at build time):

      1. model_path="s3://..." — the TGI build path calls
        _create_dir_structure(self.model_path), which runs
        Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
        this does not fail on an S3 URI; it silently creates a literal local
        directory tree named s3:/<bucket>/<prefix> in the working directory and
        continues. The S3 URI never reaches the container, and HF_MODEL_ID still
        resolves to the HF repo id.
        Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

      2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
        upload destination prefix for SDK scaffolding, not as a weight source. The
        build method sets self.s3_upload_path = None (comment: "TGI downloads
        models directly from HuggingFace Hub") and overwrites the attribute via
        self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
        with HF_MODEL_ID=<hf_repo_id> regardless.
        Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

      Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
      from HF), which makes scale-to-zero async endpoints impractical and creates a
      hard runtime dependency on huggingface.co. It also blocks deploying custom
      fine-tuned weights that are not on the HF Hub at all.

      This is closely related to #5529 (DJL Serving overwrites a user-provided
      HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
      this report is the TGI path and the required fix differs — see Additional
      context
      .

      To reproduce

      Build-only repro (no GPU, no deploy, ungated model so no token needed).
      sagemaker==3.12.0.

      frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
      sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
      sample_output=[{"generated_text": "Hi"}],
      )
      # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
      model="Qwen/Qwen2.5-7B-Instruct",
      model_path=S3_URI,
      model_server=ModelServer.TGI,
      schema_builder=schema,
      role_arn=ROLE,
      instance_type="ml.g5.2xlarge",
      )
      mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
      model="Qwen/Qwen2.5-7B-Instruct",
      s3_model_data_url=S3_URI,
      model_server=ModelServer.TGI,
      schema_builder=schema,
      role_arn=ROLE,
      instance_type="ml.g5.2xlarge",
      )
      mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

      Expected behavior

      With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
      to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
      the SDK should not silently ignore an S3 value passed via model_path /
      s3_model_data_url (either honor it or raise a clear error).

      Screenshots or logs

      Output of the repro above (SDK 3.12.0):

      Case 1 — local junk dir created: True
      Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
      Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
      

      For comparison, the working lower-level deployment logs this within ~3s of
      container start (weights served from S3, no HF download):

      INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
      INFO download: Files are already present on the host. Skipping download.
      INFO download: Successfully downloaded weights for /opt/ml/model
      

      System information

      • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
      • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
      • Python version: 3.12
      • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
      • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

      Additional context

      Working workaround — bypass ModelBuilder and construct the resources
      directly, attaching the S3 weights as an uncompressed ModelDataSource and
      pointing HF_MODEL_ID at the mount path:

      fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
      ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
      )
      fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
      framework="huggingface-llm", region="us-east-1",
      version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
      )
      primary=ContainerDefinition(
      image=image_uri,
      environment={
      "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
      "SM_NUM_GPUS": "1",
      "MESSAGES_API_ENABLED": "true",
      },
      model_data_source=ModelDataSource(
      s3_data_source=S3ModelDataSource(
      s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
      s3_data_type="S3Prefix",
      compression_type="None",
      )
      ),
      )
      Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
      EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
      Endpoint.create(...)

      This drops cold start from ~12 min to ~3–5 min for a 7B model.

      Relationship to existing issues

      #5529#5588This issue
      Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
      Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
      Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
      Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
      Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

      Suggested fix (TGI path) — when an S3 URI is supplied:

      1. Skip _create_dir_structure local-dir creation for S3 inputs.
      2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
        the early branch in
        sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
        already emits the correct shape when model_path is an S3 URI, but it's
        currently unreachable because _create_dir_structure runs first.
      3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

      More broadly, consider unifying S3-as-weight-source handling across all
      model_server branches so behavior is consistent (ties together #5529, #5588,
      and this issue).

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

          Description

          @sagarneeldubey

          Description

          PySDK Version

          • PySDK V2 (2.x)
          • PySDK V3 (3.x)

          Describe the bug

          When using ModelBuilder with ModelServer.TGI, there is no working way to
          load model weights from S3 — the deployed TGI container always pulls weights
          from huggingface.co at startup. Both parameters that look like they should
          enable S3 loading are silently ignored (neither raises at build time):

          1. model_path="s3://..." — the TGI build path calls
            _create_dir_structure(self.model_path), which runs
            Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
            this does not fail on an S3 URI; it silently creates a literal local
            directory tree named s3:/<bucket>/<prefix> in the working directory and
            continues. The S3 URI never reaches the container, and HF_MODEL_ID still
            resolves to the HF repo id.
            Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

          2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
            upload destination prefix for SDK scaffolding, not as a weight source. The
            build method sets self.s3_upload_path = None (comment: "TGI downloads
            models directly from HuggingFace Hub") and overwrites the attribute via
            self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
            with HF_MODEL_ID=<hf_repo_id> regardless.
            Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

          Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
          from HF), which makes scale-to-zero async endpoints impractical and creates a
          hard runtime dependency on huggingface.co. It also blocks deploying custom
          fine-tuned weights that are not on the HF Hub at all.

          This is closely related to #5529 (DJL Serving overwrites a user-provided
          HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
          this report is the TGI path and the required fix differs — see Additional
          context
          .

          To reproduce

          Build-only repro (no GPU, no deploy, ungated model so no token needed).
          sagemaker==3.12.0.

          frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
          sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
          sample_output=[{"generated_text": "Hi"}],
          )
          # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
          model="Qwen/Qwen2.5-7B-Instruct",
          model_path=S3_URI,
          model_server=ModelServer.TGI,
          schema_builder=schema,
          role_arn=ROLE,
          instance_type="ml.g5.2xlarge",
          )
          mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
          model="Qwen/Qwen2.5-7B-Instruct",
          s3_model_data_url=S3_URI,
          model_server=ModelServer.TGI,
          schema_builder=schema,
          role_arn=ROLE,
          instance_type="ml.g5.2xlarge",
          )
          mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

          Expected behavior

          With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
          to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
          the SDK should not silently ignore an S3 value passed via model_path /
          s3_model_data_url (either honor it or raise a clear error).

          Screenshots or logs

          Output of the repro above (SDK 3.12.0):

          Case 1 — local junk dir created: True
          Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
          Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
          

          For comparison, the working lower-level deployment logs this within ~3s of
          container start (weights served from S3, no HF download):

          INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
          INFO download: Files are already present on the host. Skipping download.
          INFO download: Successfully downloaded weights for /opt/ml/model
          

          System information

          • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
          • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
          • Python version: 3.12
          • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
          • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

          Additional context

          Working workaround — bypass ModelBuilder and construct the resources
          directly, attaching the S3 weights as an uncompressed ModelDataSource and
          pointing HF_MODEL_ID at the mount path:

          fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
          ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
          )
          fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
          framework="huggingface-llm", region="us-east-1",
          version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
          )
          primary=ContainerDefinition(
          image=image_uri,
          environment={
          "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
          "SM_NUM_GPUS": "1",
          "MESSAGES_API_ENABLED": "true",
          },
          model_data_source=ModelDataSource(
          s3_data_source=S3ModelDataSource(
          s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
          s3_data_type="S3Prefix",
          compression_type="None",
          )
          ),
          )
          Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
          EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
          Endpoint.create(...)

          This drops cold start from ~12 min to ~3–5 min for a 7B model.

          Relationship to existing issues

          #5529#5588This issue
          Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
          Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
          Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
          Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
          Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

          Suggested fix (TGI path) — when an S3 URI is supplied:

          1. Skip _create_dir_structure local-dir creation for S3 inputs.
          2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
            the early branch in
            sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
            already emits the correct shape when model_path is an S3 URI, but it's
            currently unreachable because _create_dir_structure runs first.
          3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

          More broadly, consider unifying S3-as-weight-source handling across all
          model_server branches so behavior is consistent (ties together #5529, #5588,
          and this issue).

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

              Description

              @sagarneeldubey

              Description

              PySDK Version

              • PySDK V2 (2.x)
              • PySDK V3 (3.x)

              Describe the bug

              When using ModelBuilder with ModelServer.TGI, there is no working way to
              load model weights from S3 — the deployed TGI container always pulls weights
              from huggingface.co at startup. Both parameters that look like they should
              enable S3 loading are silently ignored (neither raises at build time):

              1. model_path="s3://..." — the TGI build path calls
                _create_dir_structure(self.model_path), which runs
                Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
                this does not fail on an S3 URI; it silently creates a literal local
                directory tree named s3:/<bucket>/<prefix> in the working directory and
                continues. The S3 URI never reaches the container, and HF_MODEL_ID still
                resolves to the HF repo id.
                Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

              2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
                upload destination prefix for SDK scaffolding, not as a weight source. The
                build method sets self.s3_upload_path = None (comment: "TGI downloads
                models directly from HuggingFace Hub") and overwrites the attribute via
                self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
                with HF_MODEL_ID=<hf_repo_id> regardless.
                Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

              Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
              from HF), which makes scale-to-zero async endpoints impractical and creates a
              hard runtime dependency on huggingface.co. It also blocks deploying custom
              fine-tuned weights that are not on the HF Hub at all.

              This is closely related to #5529 (DJL Serving overwrites a user-provided
              HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
              this report is the TGI path and the required fix differs — see Additional
              context
              .

              To reproduce

              Build-only repro (no GPU, no deploy, ungated model so no token needed).
              sagemaker==3.12.0.

              frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
              sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
              sample_output=[{"generated_text": "Hi"}],
              )
              # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
              model="Qwen/Qwen2.5-7B-Instruct",
              model_path=S3_URI,
              model_server=ModelServer.TGI,
              schema_builder=schema,
              role_arn=ROLE,
              instance_type="ml.g5.2xlarge",
              )
              mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
              model="Qwen/Qwen2.5-7B-Instruct",
              s3_model_data_url=S3_URI,
              model_server=ModelServer.TGI,
              schema_builder=schema,
              role_arn=ROLE,
              instance_type="ml.g5.2xlarge",
              )
              mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

              Expected behavior

              With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
              to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
              the SDK should not silently ignore an S3 value passed via model_path /
              s3_model_data_url (either honor it or raise a clear error).

              Screenshots or logs

              Output of the repro above (SDK 3.12.0):

              Case 1 — local junk dir created: True
              Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
              Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
              

              For comparison, the working lower-level deployment logs this within ~3s of
              container start (weights served from S3, no HF download):

              INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
              INFO download: Files are already present on the host. Skipping download.
              INFO download: Successfully downloaded weights for /opt/ml/model
              

              System information

              • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
              • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
              • Python version: 3.12
              • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
              • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

              Additional context

              Working workaround — bypass ModelBuilder and construct the resources
              directly, attaching the S3 weights as an uncompressed ModelDataSource and
              pointing HF_MODEL_ID at the mount path:

              fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
              ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
              )
              fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
              framework="huggingface-llm", region="us-east-1",
              version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
              )
              primary=ContainerDefinition(
              image=image_uri,
              environment={
              "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
              "SM_NUM_GPUS": "1",
              "MESSAGES_API_ENABLED": "true",
              },
              model_data_source=ModelDataSource(
              s3_data_source=S3ModelDataSource(
              s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
              s3_data_type="S3Prefix",
              compression_type="None",
              )
              ),
              )
              Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
              EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
              Endpoint.create(...)

              This drops cold start from ~12 min to ~3–5 min for a 7B model.

              Relationship to existing issues

              #5529#5588This issue
              Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
              Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
              Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
              Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
              Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

              Suggested fix (TGI path) — when an S3 URI is supplied:

              1. Skip _create_dir_structure local-dir creation for S3 inputs.
              2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
                the early branch in
                sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
                already emits the correct shape when model_path is an S3 URI, but it's
                currently unreachable because _create_dir_structure runs first.
              3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

              More broadly, consider unifying S3-as-weight-source handling across all
              model_server branches so behavior is consistent (ties together #5529, #5588,
              and this issue).

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

                  Description

                  @sagarneeldubey

                  Description

                  PySDK Version

                  • PySDK V2 (2.x)
                  • PySDK V3 (3.x)

                  Describe the bug

                  When using ModelBuilder with ModelServer.TGI, there is no working way to
                  load model weights from S3 — the deployed TGI container always pulls weights
                  from huggingface.co at startup. Both parameters that look like they should
                  enable S3 loading are silently ignored (neither raises at build time):

                  1. model_path="s3://..." — the TGI build path calls
                    _create_dir_structure(self.model_path), which runs
                    Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
                    this does not fail on an S3 URI; it silently creates a literal local
                    directory tree named s3:/<bucket>/<prefix> in the working directory and
                    continues. The S3 URI never reaches the container, and HF_MODEL_ID still
                    resolves to the HF repo id.
                    Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

                  2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
                    upload destination prefix for SDK scaffolding, not as a weight source. The
                    build method sets self.s3_upload_path = None (comment: "TGI downloads
                    models directly from HuggingFace Hub") and overwrites the attribute via
                    self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
                    with HF_MODEL_ID=<hf_repo_id> regardless.
                    Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

                  Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
                  from HF), which makes scale-to-zero async endpoints impractical and creates a
                  hard runtime dependency on huggingface.co. It also blocks deploying custom
                  fine-tuned weights that are not on the HF Hub at all.

                  This is closely related to #5529 (DJL Serving overwrites a user-provided
                  HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
                  this report is the TGI path and the required fix differs — see Additional
                  context
                  .

                  To reproduce

                  Build-only repro (no GPU, no deploy, ungated model so no token needed).
                  sagemaker==3.12.0.

                  frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
                  sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
                  sample_output=[{"generated_text": "Hi"}],
                  )
                  # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
                  model="Qwen/Qwen2.5-7B-Instruct",
                  model_path=S3_URI,
                  model_server=ModelServer.TGI,
                  schema_builder=schema,
                  role_arn=ROLE,
                  instance_type="ml.g5.2xlarge",
                  )
                  mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
                  model="Qwen/Qwen2.5-7B-Instruct",
                  s3_model_data_url=S3_URI,
                  model_server=ModelServer.TGI,
                  schema_builder=schema,
                  role_arn=ROLE,
                  instance_type="ml.g5.2xlarge",
                  )
                  mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

                  Expected behavior

                  With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
                  to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
                  the SDK should not silently ignore an S3 value passed via model_path /
                  s3_model_data_url (either honor it or raise a clear error).

                  Screenshots or logs

                  Output of the repro above (SDK 3.12.0):

                  Case 1 — local junk dir created: True
                  Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                  Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                  

                  For comparison, the working lower-level deployment logs this within ~3s of
                  container start (weights served from S3, no HF download):

                  INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
                  INFO download: Files are already present on the host. Skipping download.
                  INFO download: Successfully downloaded weights for /opt/ml/model
                  

                  System information

                  • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
                  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
                  • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
                  • Python version: 3.12
                  • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
                  • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

                  Additional context

                  Working workaround — bypass ModelBuilder and construct the resources
                  directly, attaching the S3 weights as an uncompressed ModelDataSource and
                  pointing HF_MODEL_ID at the mount path:

                  fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
                  ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
                  )
                  fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
                  framework="huggingface-llm", region="us-east-1",
                  version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
                  )
                  primary=ContainerDefinition(
                  image=image_uri,
                  environment={
                  "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
                  "SM_NUM_GPUS": "1",
                  "MESSAGES_API_ENABLED": "true",
                  },
                  model_data_source=ModelDataSource(
                  s3_data_source=S3ModelDataSource(
                  s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
                  s3_data_type="S3Prefix",
                  compression_type="None",
                  )
                  ),
                  )
                  Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
                  EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
                  Endpoint.create(...)

                  This drops cold start from ~12 min to ~3–5 min for a 7B model.

                  Relationship to existing issues

                  #5529#5588This issue
                  Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
                  Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
                  Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
                  Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
                  Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

                  Suggested fix (TGI path) — when an S3 URI is supplied:

                  1. Skip _create_dir_structure local-dir creation for S3 inputs.
                  2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
                    the early branch in
                    sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
                    already emits the correct shape when model_path is an S3 URI, but it's
                    currently unreachable because _create_dir_structure runs first.
                  3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

                  More broadly, consider unifying S3-as-weight-source handling across all
                  model_server branches so behavior is consistent (ties together #5529, #5588,
                  and this issue).

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

                      Description

                      @sagarneeldubey

                      Description

                      PySDK Version

                      • PySDK V2 (2.x)
                      • PySDK V3 (3.x)

                      Describe the bug

                      When using ModelBuilder with ModelServer.TGI, there is no working way to
                      load model weights from S3 — the deployed TGI container always pulls weights
                      from huggingface.co at startup. Both parameters that look like they should
                      enable S3 loading are silently ignored (neither raises at build time):

                      1. model_path="s3://..." — the TGI build path calls
                        _create_dir_structure(self.model_path), which runs
                        Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
                        this does not fail on an S3 URI; it silently creates a literal local
                        directory tree named s3:/<bucket>/<prefix> in the working directory and
                        continues. The S3 URI never reaches the container, and HF_MODEL_ID still
                        resolves to the HF repo id.
                        Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

                      2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
                        upload destination prefix for SDK scaffolding, not as a weight source. The
                        build method sets self.s3_upload_path = None (comment: "TGI downloads
                        models directly from HuggingFace Hub") and overwrites the attribute via
                        self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
                        with HF_MODEL_ID=<hf_repo_id> regardless.
                        Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

                      Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
                      from HF), which makes scale-to-zero async endpoints impractical and creates a
                      hard runtime dependency on huggingface.co. It also blocks deploying custom
                      fine-tuned weights that are not on the HF Hub at all.

                      This is closely related to #5529 (DJL Serving overwrites a user-provided
                      HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
                      this report is the TGI path and the required fix differs — see Additional
                      context
                      .

                      To reproduce

                      Build-only repro (no GPU, no deploy, ungated model so no token needed).
                      sagemaker==3.12.0.

                      frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
                      sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
                      sample_output=[{"generated_text": "Hi"}],
                      )
                      # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
                      model="Qwen/Qwen2.5-7B-Instruct",
                      model_path=S3_URI,
                      model_server=ModelServer.TGI,
                      schema_builder=schema,
                      role_arn=ROLE,
                      instance_type="ml.g5.2xlarge",
                      )
                      mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
                      model="Qwen/Qwen2.5-7B-Instruct",
                      s3_model_data_url=S3_URI,
                      model_server=ModelServer.TGI,
                      schema_builder=schema,
                      role_arn=ROLE,
                      instance_type="ml.g5.2xlarge",
                      )
                      mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

                      Expected behavior

                      With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
                      to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
                      the SDK should not silently ignore an S3 value passed via model_path /
                      s3_model_data_url (either honor it or raise a clear error).

                      Screenshots or logs

                      Output of the repro above (SDK 3.12.0):

                      Case 1 — local junk dir created: True
                      Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                      Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                      

                      For comparison, the working lower-level deployment logs this within ~3s of
                      container start (weights served from S3, no HF download):

                      INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
                      INFO download: Files are already present on the host. Skipping download.
                      INFO download: Successfully downloaded weights for /opt/ml/model
                      

                      System information

                      • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
                      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
                      • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
                      • Python version: 3.12
                      • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
                      • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

                      Additional context

                      Working workaround — bypass ModelBuilder and construct the resources
                      directly, attaching the S3 weights as an uncompressed ModelDataSource and
                      pointing HF_MODEL_ID at the mount path:

                      fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
                      ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
                      )
                      fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
                      framework="huggingface-llm", region="us-east-1",
                      version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
                      )
                      primary=ContainerDefinition(
                      image=image_uri,
                      environment={
                      "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
                      "SM_NUM_GPUS": "1",
                      "MESSAGES_API_ENABLED": "true",
                      },
                      model_data_source=ModelDataSource(
                      s3_data_source=S3ModelDataSource(
                      s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
                      s3_data_type="S3Prefix",
                      compression_type="None",
                      )
                      ),
                      )
                      Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
                      EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
                      Endpoint.create(...)

                      This drops cold start from ~12 min to ~3–5 min for a 7B model.

                      Relationship to existing issues

                      #5529#5588This issue
                      Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
                      Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
                      Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
                      Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
                      Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

                      Suggested fix (TGI path) — when an S3 URI is supplied:

                      1. Skip _create_dir_structure local-dir creation for S3 inputs.
                      2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
                        the early branch in
                        sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
                        already emits the correct shape when model_path is an S3 URI, but it's
                        currently unreachable because _create_dir_structure runs first.
                      3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

                      More broadly, consider unifying S3-as-weight-source handling across all
                      model_server branches so behavior is consistent (ties together #5529, #5588,
                      and this issue).

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

                          Description

                          @sagarneeldubey

                          Description

                          PySDK Version

                          • PySDK V2 (2.x)
                          • PySDK V3 (3.x)

                          Describe the bug

                          When using ModelBuilder with ModelServer.TGI, there is no working way to
                          load model weights from S3 — the deployed TGI container always pulls weights
                          from huggingface.co at startup. Both parameters that look like they should
                          enable S3 loading are silently ignored (neither raises at build time):

                          1. model_path="s3://..." — the TGI build path calls
                            _create_dir_structure(self.model_path), which runs
                            Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
                            this does not fail on an S3 URI; it silently creates a literal local
                            directory tree named s3:/<bucket>/<prefix> in the working directory and
                            continues. The S3 URI never reaches the container, and HF_MODEL_ID still
                            resolves to the HF repo id.
                            Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

                          2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
                            upload destination prefix for SDK scaffolding, not as a weight source. The
                            build method sets self.s3_upload_path = None (comment: "TGI downloads
                            models directly from HuggingFace Hub") and overwrites the attribute via
                            self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
                            with HF_MODEL_ID=<hf_repo_id> regardless.
                            Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

                          Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
                          from HF), which makes scale-to-zero async endpoints impractical and creates a
                          hard runtime dependency on huggingface.co. It also blocks deploying custom
                          fine-tuned weights that are not on the HF Hub at all.

                          This is closely related to #5529 (DJL Serving overwrites a user-provided
                          HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
                          this report is the TGI path and the required fix differs — see Additional
                          context
                          .

                          To reproduce

                          Build-only repro (no GPU, no deploy, ungated model so no token needed).
                          sagemaker==3.12.0.

                          frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
                          sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
                          sample_output=[{"generated_text": "Hi"}],
                          )
                          # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
                          model="Qwen/Qwen2.5-7B-Instruct",
                          model_path=S3_URI,
                          model_server=ModelServer.TGI,
                          schema_builder=schema,
                          role_arn=ROLE,
                          instance_type="ml.g5.2xlarge",
                          )
                          mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
                          model="Qwen/Qwen2.5-7B-Instruct",
                          s3_model_data_url=S3_URI,
                          model_server=ModelServer.TGI,
                          schema_builder=schema,
                          role_arn=ROLE,
                          instance_type="ml.g5.2xlarge",
                          )
                          mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

                          Expected behavior

                          With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
                          to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
                          the SDK should not silently ignore an S3 value passed via model_path /
                          s3_model_data_url (either honor it or raise a clear error).

                          Screenshots or logs

                          Output of the repro above (SDK 3.12.0):

                          Case 1 — local junk dir created: True
                          Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                          Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                          

                          For comparison, the working lower-level deployment logs this within ~3s of
                          container start (weights served from S3, no HF download):

                          INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
                          INFO download: Files are already present on the host. Skipping download.
                          INFO download: Successfully downloaded weights for /opt/ml/model
                          

                          System information

                          • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
                          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
                          • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
                          • Python version: 3.12
                          • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
                          • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

                          Additional context

                          Working workaround — bypass ModelBuilder and construct the resources
                          directly, attaching the S3 weights as an uncompressed ModelDataSource and
                          pointing HF_MODEL_ID at the mount path:

                          fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
                          ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
                          )
                          fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
                          framework="huggingface-llm", region="us-east-1",
                          version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
                          )
                          primary=ContainerDefinition(
                          image=image_uri,
                          environment={
                          "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
                          "SM_NUM_GPUS": "1",
                          "MESSAGES_API_ENABLED": "true",
                          },
                          model_data_source=ModelDataSource(
                          s3_data_source=S3ModelDataSource(
                          s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
                          s3_data_type="S3Prefix",
                          compression_type="None",
                          )
                          ),
                          )
                          Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
                          EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
                          Endpoint.create(...)

                          This drops cold start from ~12 min to ~3–5 min for a 7B model.

                          Relationship to existing issues

                          #5529#5588This issue
                          Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
                          Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
                          Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
                          Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
                          Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

                          Suggested fix (TGI path) — when an S3 URI is supplied:

                          1. Skip _create_dir_structure local-dir creation for S3 inputs.
                          2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
                            the early branch in
                            sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
                            already emits the correct shape when model_path is an S3 URI, but it's
                            currently unreachable because _create_dir_structure runs first.
                          3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

                          More broadly, consider unifying S3-as-weight-source handling across all
                          model_server branches so behavior is consistent (ties together #5529, #5588,
                          and this issue).

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              bug: ModelBuilder TGI path ignores S3 model inputs (model_path and s3_model_data_url), forcing HuggingFace Hub download #5943

                              Description

                              @sagarneeldubey

                              Description

                              PySDK Version

                              • PySDK V2 (2.x)
                              • PySDK V3 (3.x)

                              Describe the bug

                              When using ModelBuilder with ModelServer.TGI, there is no working way to
                              load model weights from S3 — the deployed TGI container always pulls weights
                              from huggingface.co at startup. Both parameters that look like they should
                              enable S3 loading are silently ignored (neither raises at build time):

                              1. model_path="s3://..." — the TGI build path calls
                                _create_dir_structure(self.model_path), which runs
                                Path(model_path).mkdir(parents=True, exist_ok=True). On a POSIX filesystem
                                this does not fail on an S3 URI; it silently creates a literal local
                                directory tree named s3:/<bucket>/<prefix> in the working directory and
                                continues. The S3 URI never reaches the container, and HF_MODEL_ID still
                                resolves to the HF repo id.
                                Source: sagemaker/serve/model_server/tgi/prepare.py::_create_dir_structure

                              2. s3_model_data_url="s3://..." — on the TGI path this is treated as an
                                upload destination prefix for SDK scaffolding, not as a weight source. The
                                build method sets self.s3_upload_path = None (comment: "TGI downloads
                                models directly from HuggingFace Hub") and overwrites the attribute via
                                self.s3_model_data_url, _ = self._prepare_for_mode(). The container ends up
                                with HF_MODEL_ID=<hf_repo_id> regardless.
                                Source: sagemaker/serve/model_builder_servers.py::_build_for_tgi

                              Impact: a 7B model cold-starts in ~10–12 min on every scale-out (re-downloading
                              from HF), which makes scale-to-zero async endpoints impractical and creates a
                              hard runtime dependency on huggingface.co. It also blocks deploying custom
                              fine-tuned weights that are not on the HF Hub at all.

                              This is closely related to #5529 (DJL Serving overwrites a user-provided
                              HF_MODEL_ID) and #5588 (DJL passthrough injects a bad ModelDataUrl), but
                              this report is the TGI path and the required fix differs — see Additional
                              context
                              .

                              To reproduce

                              Build-only repro (no GPU, no deploy, ungated model so no token needed).
                              sagemaker==3.12.0.

                              frompathlibimportPathfromsagemaker.serveimportModelBuilderfromsagemaker.serve.builder.schema_builderimportSchemaBuilderfromsagemaker.serve.utils.typesimportModelServerROLE="arn:aws:iam::123456789012:role/SageMakerRole"# any valid role ARNS3_URI="s3://my-bucket/models/qwen2.5-7b-instruct/"# need not exist for the reproschema=SchemaBuilder(
                              sample_input={"inputs": "Hello", "parameters": {"max_new_tokens": 16}},
                              sample_output=[{"generated_text": "Hi"}],
                              )
                              # --- Case 1: model_path as an S3 URI ---mb=ModelBuilder(
                              model="Qwen/Qwen2.5-7B-Instruct",
                              model_path=S3_URI,
                              model_server=ModelServer.TGI,
                              schema_builder=schema,
                              role_arn=ROLE,
                              instance_type="ml.g5.2xlarge",
                              )
                              mb.build() # does NOT raiseprint("Case 1 — local junk dir created:", Path(S3_URI).exists()) # True (e.g. ./s3:/my-bucket/...)print("Case 1 — HF_MODEL_ID:", mb.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)# --- Case 2: s3_model_data_url as an S3 URI ---mb2=ModelBuilder(
                              model="Qwen/Qwen2.5-7B-Instruct",
                              s3_model_data_url=S3_URI,
                              model_server=ModelServer.TGI,
                              schema_builder=schema,
                              role_arn=ROLE,
                              instance_type="ml.g5.2xlarge",
                              )
                              mb2.build() # does NOT raiseprint("Case 2 — HF_MODEL_ID:", mb2.env_vars.get("HF_MODEL_ID")) # 'Qwen/Qwen2.5-7B-Instruct' (WRONG)

                              Expected behavior

                              With weights pre-staged in S3, a TGI endpoint should load them from S3 (mounted
                              to /opt/ml/model) instead of downloading from huggingface.co — and at minimum
                              the SDK should not silently ignore an S3 value passed via model_path /
                              s3_model_data_url (either honor it or raise a clear error).

                              Screenshots or logs

                              Output of the repro above (SDK 3.12.0):

                              Case 1 — local junk dir created: True
                              Case 1 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                              Case 2 — HF_MODEL_ID: Qwen/Qwen2.5-7B-Instruct
                              

                              For comparison, the working lower-level deployment logs this within ~3s of
                              container start (weights served from S3, no HF download):

                              INFO text_generation_launcher: Args { model_id: "/opt/ml/model", ... }
                              INFO download: Files are already present on the host. Skipping download.
                              INFO download: Successfully downloaded weights for /opt/ml/model
                              

                              System information

                              • SageMaker Python SDK version: 3.12.0 (also present in 3.3.x–3.13.x by source inspection)
                              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Hugging Face TGI (Text Generation Inference)
                              • Framework version: TGI DLC huggingface-pytorch-tgi-inference:2.4.0-tgi2.3.1-gpu-py311-cu124-ubuntu22.04
                              • Python version: 3.12
                              • CPU or GPU: GPU (ml.g5.2xlarge, ml.g5.12xlarge)
                              • Custom Docker image (Y/N): N (AWS-provided HF TGI DLC)

                              Additional context

                              Working workaround — bypass ModelBuilder and construct the resources
                              directly, attaching the S3 weights as an uncompressed ModelDataSource and
                              pointing HF_MODEL_ID at the mount path:

                              fromsagemaker.core.resourcesimportEndpoint, EndpointConfig, Modelfromsagemaker.core.shapes.shapesimport (
                              ContainerDefinition, ModelDataSource, ProductionVariant, S3ModelDataSource,
                              )
                              fromsagemaker.core.image_retriever.image_retrieverimportImageRetrieverimage_uri=ImageRetriever.retrieve(
                              framework="huggingface-llm", region="us-east-1",
                              version="2.3.1", image_scope="inference", instance_type="ml.g5.2xlarge",
                              )
                              primary=ContainerDefinition(
                              image=image_uri,
                              environment={
                              "HF_MODEL_ID": "/opt/ml/model", # mount path, not the HF repo id"HF_HUB_OFFLINE": "1",
                              "SM_NUM_GPUS": "1",
                              "MESSAGES_API_ENABLED": "true",
                              },
                              model_data_source=ModelDataSource(
                              s3_data_source=S3ModelDataSource(
                              s3_uri="s3://my-bucket/models/qwen2.5-7b-instruct/",
                              s3_data_type="S3Prefix",
                              compression_type="None",
                              )
                              ),
                              )
                              Model.create(model_name=..., primary_container=primary, execution_role_arn=ROLE, ...)
                              EndpointConfig.create(..., production_variants=[ProductionVariant(...)])
                              Endpoint.create(...)

                              This drops cold start from ~12 min to ~3–5 min for a 7B model.

                              Relationship to existing issues

                              #5529#5588This issue
                              Model serverDJL Serving (LMI)DJL Serving (LMI)TGI
                              Modemodel + env_varspassthrough (image_uri + env_vars, no model)model + model_path/s3_model_data_url
                              Symptomuser HF_MODEL_ID overwrittenbad ModelDataUrl injected → ValidationExceptionS3 inputs silently ignored; junk local dir created
                              Does the engine accept S3 in HF_MODEL_ID?Yes (LMI option.model_id)YesNo — TGI expects HF repo id or local path
                              Sufficient fixsetdefault on HF_MODEL_IDclear s3_model_data_url in passthroughsetdefault is necessary but not sufficient; TGI also needs weights mounted at /opt/ml/model via ModelDataSource

                              Suggested fix (TGI path) — when an S3 URI is supplied:

                              1. Skip _create_dir_structure local-dir creation for S3 inputs.
                              2. Attach the S3 prefix as an uncompressed ModelDataSource on the container —
                                the early branch in
                                sagemaker/serve/model_server/tgi/server.py::SageMakerTgiServing._upload_tgi_artifacts
                                already emits the correct shape when model_path is an S3 URI, but it's
                                currently unreachable because _create_dir_structure runs first.
                              3. Set HF_MODEL_ID=/opt/ml/model instead of the HF repo id.

                              More broadly, consider unifying S3-as-weight-source handling across all
                              model_server branches so behavior is consistent (ties together #5529, #5588,
                              and this issue).

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions