Cannot create multiple model endpoints in local mode #2020

Description

@pwerth

Describe the bug
In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

To reproduce
I have two images corresponding to two different containers, let's call them image-1 and image-2.

In test_train_model1.py:

def test_train_and_predict(tmp_path):
sagemaker_session = LocalSession()
sagemaker_session.config = {'local': {'local_code': True}}
training_channel = tmp_path / "training"
model_output = tmp_path / "model"
training_output = tmp_path / "output"
training_channel.mkdir(exist_ok=True)
model_output.mkdir(exist_ok=True)
training_output.mkdir(exist_ok=True)
# Move the test dataset into the location expected by train container
train = pd.read_csv(<path_to_local_csv>)
train.to_csv(str(training_channel) + "/data.csv", index=False)
estimator = Estimator(
settings.get("image"), # equals `image_1`
settings.get("iam_role_arn"),
settings.getint('instance_count'),
settings.get("instance_type"),
base_job_name='model-1-train',
volume_size=512,
model_uri="file://" + str(model_output),
output_path="file://" + str(training_output),
sagemaker_session=sagemaker_session,
hyperparameters=settings.get("hyperparameters")
)
estimator.fit({
"training": f"file://{training_channel}"
}, logs="All", wait=True)
assert estimator
# Check that model got saved to correct location
assert estimator.model_data == f'file://{training_output}/model.tar.gz'
# Deploy the model locally
predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
# Check that we got a valid prediction
data = ... # some test input
response = json.loads(predictor.predict(json.dumps(data)))
assert response 

In test_model_2.py (identical logic, different image):

def test_train_and_predict(tmp_path):
sagemaker_session = LocalSession()
sagemaker_session.config = {'local': {'local_code': True}}
training_channel = tmp_path / "training"
model_output = tmp_path / "model"
training_output = tmp_path / "output"
training_channel.mkdir(exist_ok=True)
model_output.mkdir(exist_ok=True)
training_output.mkdir(exist_ok=True)
# Move the test dataset into the location expected by train container
train = pd.read_csv(<path_to_local_csv>)
train.to_csv(str(training_channel) + "/data.csv", index=False)
estimator = Estimator(
settings.get("image"), # equals `image_2`
settings.get("iam_role_arn"),
settings.getint('instance_count'),
settings.get("instance_type"),
base_job_name='model-1-train',
volume_size=512,
model_uri="file://" + str(model_output),
output_path="file://" + str(training_output),
sagemaker_session=sagemaker_session,
hyperparameters=settings.get("hyperparameters")
)
estimator.fit({
"training": f"file://{training_channel}"
}, logs="All", wait=True)
assert estimator
# Check that model got saved to correct location
assert estimator.model_data == f'file://{training_output}/model.tar.gz'
# Deploy the model locally
predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
# Check that we got a valid prediction
data = ... # some test input
response = json.loads(predictor.predict(json.dumps(data)))
assert response 

Expected behavior
I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

Screenshots or logs
Logs from the second test starting to run:

INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
Traceback (most recent call last):
File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
_stream_output(self.process)
File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
raise RuntimeError("Process exited with code: %s" % exit_code)
RuntimeError: Process exited with code: 1
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
self.run()
File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
raise RuntimeError(msg)
RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
Creating tmphvpfolv1_algo-1-n4k0x_1 ... done

I tried running the command in manually in my terminal and got the following :

docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
ERROR: Encountered errors while bringing up the project.

Note: this output is more useful than what comes out of the SDK.

System information
A description of your system. Please provide:

  • SageMaker Python SDK version: 2.18.0
  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
  • Framework version: 3.1.0
  • Python version: 3.7
  • Custom Docker image (Y/N): Y

Additional context
Add any other context about the problem here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      Cannot create multiple model endpoints in local mode #2020

      Description

      @pwerth

      Describe the bug
      In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

      This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

      If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

      To reproduce
      I have two images corresponding to two different containers, let's call them image-1 and image-2.

      In test_train_model1.py:

      def test_train_and_predict(tmp_path):
      sagemaker_session = LocalSession()
      sagemaker_session.config = {'local': {'local_code': True}}
      training_channel = tmp_path / "training"
      model_output = tmp_path / "model"
      training_output = tmp_path / "output"
      training_channel.mkdir(exist_ok=True)
      model_output.mkdir(exist_ok=True)
      training_output.mkdir(exist_ok=True)
      # Move the test dataset into the location expected by train container
      train = pd.read_csv(<path_to_local_csv>)
      train.to_csv(str(training_channel) + "/data.csv", index=False)
      estimator = Estimator(
      settings.get("image"), # equals `image_1`
      settings.get("iam_role_arn"),
      settings.getint('instance_count'),
      settings.get("instance_type"),
      base_job_name='model-1-train',
      volume_size=512,
      model_uri="file://" + str(model_output),
      output_path="file://" + str(training_output),
      sagemaker_session=sagemaker_session,
      hyperparameters=settings.get("hyperparameters")
      )
      estimator.fit({
      "training": f"file://{training_channel}"
      }, logs="All", wait=True)
      assert estimator
      # Check that model got saved to correct location
      assert estimator.model_data == f'file://{training_output}/model.tar.gz'
      # Deploy the model locally
      predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
      # Check that we got a valid prediction
      data = ... # some test input
      response = json.loads(predictor.predict(json.dumps(data)))
      assert response 

      In test_model_2.py (identical logic, different image):

      def test_train_and_predict(tmp_path):
      sagemaker_session = LocalSession()
      sagemaker_session.config = {'local': {'local_code': True}}
      training_channel = tmp_path / "training"
      model_output = tmp_path / "model"
      training_output = tmp_path / "output"
      training_channel.mkdir(exist_ok=True)
      model_output.mkdir(exist_ok=True)
      training_output.mkdir(exist_ok=True)
      # Move the test dataset into the location expected by train container
      train = pd.read_csv(<path_to_local_csv>)
      train.to_csv(str(training_channel) + "/data.csv", index=False)
      estimator = Estimator(
      settings.get("image"), # equals `image_2`
      settings.get("iam_role_arn"),
      settings.getint('instance_count'),
      settings.get("instance_type"),
      base_job_name='model-1-train',
      volume_size=512,
      model_uri="file://" + str(model_output),
      output_path="file://" + str(training_output),
      sagemaker_session=sagemaker_session,
      hyperparameters=settings.get("hyperparameters")
      )
      estimator.fit({
      "training": f"file://{training_channel}"
      }, logs="All", wait=True)
      assert estimator
      # Check that model got saved to correct location
      assert estimator.model_data == f'file://{training_output}/model.tar.gz'
      # Deploy the model locally
      predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
      # Check that we got a valid prediction
      data = ... # some test input
      response = json.loads(predictor.predict(json.dumps(data)))
      assert response 

      Expected behavior
      I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

      Screenshots or logs
      Logs from the second test starting to run:

      INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
      Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
      Traceback (most recent call last):
      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
      _stream_output(self.process)
      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
      raise RuntimeError("Process exited with code: %s" % exit_code)
      RuntimeError: Process exited with code: 1
      During handling of the above exception, another exception occurred:
      Traceback (most recent call last):
      File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
      self.run()
      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
      raise RuntimeError(msg)
      RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
      Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
      

      I tried running the command in manually in my terminal and got the following :

      docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
      Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
      Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
      ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
      ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
      ERROR: Encountered errors while bringing up the project.
      

      Note: this output is more useful than what comes out of the SDK.

      System information
      A description of your system. Please provide:

      • SageMaker Python SDK version: 2.18.0
      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
      • Framework version: 3.1.0
      • Python version: 3.7
      • Custom Docker image (Y/N): Y

      Additional context
      Add any other context about the problem here.

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          Cannot create multiple model endpoints in local mode #2020

          Description

          @pwerth

          Describe the bug
          In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

          This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

          If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

          To reproduce
          I have two images corresponding to two different containers, let's call them image-1 and image-2.

          In test_train_model1.py:

          def test_train_and_predict(tmp_path):
          sagemaker_session = LocalSession()
          sagemaker_session.config = {'local': {'local_code': True}}
          training_channel = tmp_path / "training"
          model_output = tmp_path / "model"
          training_output = tmp_path / "output"
          training_channel.mkdir(exist_ok=True)
          model_output.mkdir(exist_ok=True)
          training_output.mkdir(exist_ok=True)
          # Move the test dataset into the location expected by train container
          train = pd.read_csv(<path_to_local_csv>)
          train.to_csv(str(training_channel) + "/data.csv", index=False)
          estimator = Estimator(
          settings.get("image"), # equals `image_1`
          settings.get("iam_role_arn"),
          settings.getint('instance_count'),
          settings.get("instance_type"),
          base_job_name='model-1-train',
          volume_size=512,
          model_uri="file://" + str(model_output),
          output_path="file://" + str(training_output),
          sagemaker_session=sagemaker_session,
          hyperparameters=settings.get("hyperparameters")
          )
          estimator.fit({
          "training": f"file://{training_channel}"
          }, logs="All", wait=True)
          assert estimator
          # Check that model got saved to correct location
          assert estimator.model_data == f'file://{training_output}/model.tar.gz'
          # Deploy the model locally
          predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
          # Check that we got a valid prediction
          data = ... # some test input
          response = json.loads(predictor.predict(json.dumps(data)))
          assert response 

          In test_model_2.py (identical logic, different image):

          def test_train_and_predict(tmp_path):
          sagemaker_session = LocalSession()
          sagemaker_session.config = {'local': {'local_code': True}}
          training_channel = tmp_path / "training"
          model_output = tmp_path / "model"
          training_output = tmp_path / "output"
          training_channel.mkdir(exist_ok=True)
          model_output.mkdir(exist_ok=True)
          training_output.mkdir(exist_ok=True)
          # Move the test dataset into the location expected by train container
          train = pd.read_csv(<path_to_local_csv>)
          train.to_csv(str(training_channel) + "/data.csv", index=False)
          estimator = Estimator(
          settings.get("image"), # equals `image_2`
          settings.get("iam_role_arn"),
          settings.getint('instance_count'),
          settings.get("instance_type"),
          base_job_name='model-1-train',
          volume_size=512,
          model_uri="file://" + str(model_output),
          output_path="file://" + str(training_output),
          sagemaker_session=sagemaker_session,
          hyperparameters=settings.get("hyperparameters")
          )
          estimator.fit({
          "training": f"file://{training_channel}"
          }, logs="All", wait=True)
          assert estimator
          # Check that model got saved to correct location
          assert estimator.model_data == f'file://{training_output}/model.tar.gz'
          # Deploy the model locally
          predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
          # Check that we got a valid prediction
          data = ... # some test input
          response = json.loads(predictor.predict(json.dumps(data)))
          assert response 

          Expected behavior
          I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

          Screenshots or logs
          Logs from the second test starting to run:

          INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
          Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
          Traceback (most recent call last):
          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
          _stream_output(self.process)
          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
          raise RuntimeError("Process exited with code: %s" % exit_code)
          RuntimeError: Process exited with code: 1
          During handling of the above exception, another exception occurred:
          Traceback (most recent call last):
          File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
          self.run()
          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
          raise RuntimeError(msg)
          RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
          Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
          

          I tried running the command in manually in my terminal and got the following :

          docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
          Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
          Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
          ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
          ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
          ERROR: Encountered errors while bringing up the project.
          

          Note: this output is more useful than what comes out of the SDK.

          System information
          A description of your system. Please provide:

          • SageMaker Python SDK version: 2.18.0
          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
          • Framework version: 3.1.0
          • Python version: 3.7
          • Custom Docker image (Y/N): Y

          Additional context
          Add any other context about the problem here.

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Cannot create multiple model endpoints in local mode #2020

              Description

              @pwerth

              Describe the bug
              In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

              This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

              If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

              To reproduce
              I have two images corresponding to two different containers, let's call them image-1 and image-2.

              In test_train_model1.py:

              def test_train_and_predict(tmp_path):
              sagemaker_session = LocalSession()
              sagemaker_session.config = {'local': {'local_code': True}}
              training_channel = tmp_path / "training"
              model_output = tmp_path / "model"
              training_output = tmp_path / "output"
              training_channel.mkdir(exist_ok=True)
              model_output.mkdir(exist_ok=True)
              training_output.mkdir(exist_ok=True)
              # Move the test dataset into the location expected by train container
              train = pd.read_csv(<path_to_local_csv>)
              train.to_csv(str(training_channel) + "/data.csv", index=False)
              estimator = Estimator(
              settings.get("image"), # equals `image_1`
              settings.get("iam_role_arn"),
              settings.getint('instance_count'),
              settings.get("instance_type"),
              base_job_name='model-1-train',
              volume_size=512,
              model_uri="file://" + str(model_output),
              output_path="file://" + str(training_output),
              sagemaker_session=sagemaker_session,
              hyperparameters=settings.get("hyperparameters")
              )
              estimator.fit({
              "training": f"file://{training_channel}"
              }, logs="All", wait=True)
              assert estimator
              # Check that model got saved to correct location
              assert estimator.model_data == f'file://{training_output}/model.tar.gz'
              # Deploy the model locally
              predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
              # Check that we got a valid prediction
              data = ... # some test input
              response = json.loads(predictor.predict(json.dumps(data)))
              assert response 

              In test_model_2.py (identical logic, different image):

              def test_train_and_predict(tmp_path):
              sagemaker_session = LocalSession()
              sagemaker_session.config = {'local': {'local_code': True}}
              training_channel = tmp_path / "training"
              model_output = tmp_path / "model"
              training_output = tmp_path / "output"
              training_channel.mkdir(exist_ok=True)
              model_output.mkdir(exist_ok=True)
              training_output.mkdir(exist_ok=True)
              # Move the test dataset into the location expected by train container
              train = pd.read_csv(<path_to_local_csv>)
              train.to_csv(str(training_channel) + "/data.csv", index=False)
              estimator = Estimator(
              settings.get("image"), # equals `image_2`
              settings.get("iam_role_arn"),
              settings.getint('instance_count'),
              settings.get("instance_type"),
              base_job_name='model-1-train',
              volume_size=512,
              model_uri="file://" + str(model_output),
              output_path="file://" + str(training_output),
              sagemaker_session=sagemaker_session,
              hyperparameters=settings.get("hyperparameters")
              )
              estimator.fit({
              "training": f"file://{training_channel}"
              }, logs="All", wait=True)
              assert estimator
              # Check that model got saved to correct location
              assert estimator.model_data == f'file://{training_output}/model.tar.gz'
              # Deploy the model locally
              predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
              # Check that we got a valid prediction
              data = ... # some test input
              response = json.loads(predictor.predict(json.dumps(data)))
              assert response 

              Expected behavior
              I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

              Screenshots or logs
              Logs from the second test starting to run:

              INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
              Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
              Traceback (most recent call last):
              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
              _stream_output(self.process)
              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
              raise RuntimeError("Process exited with code: %s" % exit_code)
              RuntimeError: Process exited with code: 1
              During handling of the above exception, another exception occurred:
              Traceback (most recent call last):
              File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
              self.run()
              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
              raise RuntimeError(msg)
              RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
              Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
              

              I tried running the command in manually in my terminal and got the following :

              docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
              Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
              Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
              ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
              ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
              ERROR: Encountered errors while bringing up the project.
              

              Note: this output is more useful than what comes out of the SDK.

              System information
              A description of your system. Please provide:

              • SageMaker Python SDK version: 2.18.0
              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
              • Framework version: 3.1.0
              • Python version: 3.7
              • Custom Docker image (Y/N): Y

              Additional context
              Add any other context about the problem here.

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  Cannot create multiple model endpoints in local mode #2020

                  Description

                  @pwerth

                  Describe the bug
                  In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

                  This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

                  If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

                  To reproduce
                  I have two images corresponding to two different containers, let's call them image-1 and image-2.

                  In test_train_model1.py:

                  def test_train_and_predict(tmp_path):
                  sagemaker_session = LocalSession()
                  sagemaker_session.config = {'local': {'local_code': True}}
                  training_channel = tmp_path / "training"
                  model_output = tmp_path / "model"
                  training_output = tmp_path / "output"
                  training_channel.mkdir(exist_ok=True)
                  model_output.mkdir(exist_ok=True)
                  training_output.mkdir(exist_ok=True)
                  # Move the test dataset into the location expected by train container
                  train = pd.read_csv(<path_to_local_csv>)
                  train.to_csv(str(training_channel) + "/data.csv", index=False)
                  estimator = Estimator(
                  settings.get("image"), # equals `image_1`
                  settings.get("iam_role_arn"),
                  settings.getint('instance_count'),
                  settings.get("instance_type"),
                  base_job_name='model-1-train',
                  volume_size=512,
                  model_uri="file://" + str(model_output),
                  output_path="file://" + str(training_output),
                  sagemaker_session=sagemaker_session,
                  hyperparameters=settings.get("hyperparameters")
                  )
                  estimator.fit({
                  "training": f"file://{training_channel}"
                  }, logs="All", wait=True)
                  assert estimator
                  # Check that model got saved to correct location
                  assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                  # Deploy the model locally
                  predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                  # Check that we got a valid prediction
                  data = ... # some test input
                  response = json.loads(predictor.predict(json.dumps(data)))
                  assert response 

                  In test_model_2.py (identical logic, different image):

                  def test_train_and_predict(tmp_path):
                  sagemaker_session = LocalSession()
                  sagemaker_session.config = {'local': {'local_code': True}}
                  training_channel = tmp_path / "training"
                  model_output = tmp_path / "model"
                  training_output = tmp_path / "output"
                  training_channel.mkdir(exist_ok=True)
                  model_output.mkdir(exist_ok=True)
                  training_output.mkdir(exist_ok=True)
                  # Move the test dataset into the location expected by train container
                  train = pd.read_csv(<path_to_local_csv>)
                  train.to_csv(str(training_channel) + "/data.csv", index=False)
                  estimator = Estimator(
                  settings.get("image"), # equals `image_2`
                  settings.get("iam_role_arn"),
                  settings.getint('instance_count'),
                  settings.get("instance_type"),
                  base_job_name='model-1-train',
                  volume_size=512,
                  model_uri="file://" + str(model_output),
                  output_path="file://" + str(training_output),
                  sagemaker_session=sagemaker_session,
                  hyperparameters=settings.get("hyperparameters")
                  )
                  estimator.fit({
                  "training": f"file://{training_channel}"
                  }, logs="All", wait=True)
                  assert estimator
                  # Check that model got saved to correct location
                  assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                  # Deploy the model locally
                  predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                  # Check that we got a valid prediction
                  data = ... # some test input
                  response = json.loads(predictor.predict(json.dumps(data)))
                  assert response 

                  Expected behavior
                  I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

                  Screenshots or logs
                  Logs from the second test starting to run:

                  INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
                  Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
                  Traceback (most recent call last):
                  File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
                  _stream_output(self.process)
                  File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
                  raise RuntimeError("Process exited with code: %s" % exit_code)
                  RuntimeError: Process exited with code: 1
                  During handling of the above exception, another exception occurred:
                  Traceback (most recent call last):
                  File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
                  self.run()
                  File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
                  raise RuntimeError(msg)
                  RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
                  Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
                  

                  I tried running the command in manually in my terminal and got the following :

                  docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
                  Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
                  Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
                  ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                  ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                  ERROR: Encountered errors while bringing up the project.
                  

                  Note: this output is more useful than what comes out of the SDK.

                  System information
                  A description of your system. Please provide:

                  • SageMaker Python SDK version: 2.18.0
                  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
                  • Framework version: 3.1.0
                  • Python version: 3.7
                  • Custom Docker image (Y/N): Y

                  Additional context
                  Add any other context about the problem here.

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      Cannot create multiple model endpoints in local mode #2020

                      Description

                      @pwerth

                      Describe the bug
                      In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

                      This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

                      If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

                      To reproduce
                      I have two images corresponding to two different containers, let's call them image-1 and image-2.

                      In test_train_model1.py:

                      def test_train_and_predict(tmp_path):
                      sagemaker_session = LocalSession()
                      sagemaker_session.config = {'local': {'local_code': True}}
                      training_channel = tmp_path / "training"
                      model_output = tmp_path / "model"
                      training_output = tmp_path / "output"
                      training_channel.mkdir(exist_ok=True)
                      model_output.mkdir(exist_ok=True)
                      training_output.mkdir(exist_ok=True)
                      # Move the test dataset into the location expected by train container
                      train = pd.read_csv(<path_to_local_csv>)
                      train.to_csv(str(training_channel) + "/data.csv", index=False)
                      estimator = Estimator(
                      settings.get("image"), # equals `image_1`
                      settings.get("iam_role_arn"),
                      settings.getint('instance_count'),
                      settings.get("instance_type"),
                      base_job_name='model-1-train',
                      volume_size=512,
                      model_uri="file://" + str(model_output),
                      output_path="file://" + str(training_output),
                      sagemaker_session=sagemaker_session,
                      hyperparameters=settings.get("hyperparameters")
                      )
                      estimator.fit({
                      "training": f"file://{training_channel}"
                      }, logs="All", wait=True)
                      assert estimator
                      # Check that model got saved to correct location
                      assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                      # Deploy the model locally
                      predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                      # Check that we got a valid prediction
                      data = ... # some test input
                      response = json.loads(predictor.predict(json.dumps(data)))
                      assert response 

                      In test_model_2.py (identical logic, different image):

                      def test_train_and_predict(tmp_path):
                      sagemaker_session = LocalSession()
                      sagemaker_session.config = {'local': {'local_code': True}}
                      training_channel = tmp_path / "training"
                      model_output = tmp_path / "model"
                      training_output = tmp_path / "output"
                      training_channel.mkdir(exist_ok=True)
                      model_output.mkdir(exist_ok=True)
                      training_output.mkdir(exist_ok=True)
                      # Move the test dataset into the location expected by train container
                      train = pd.read_csv(<path_to_local_csv>)
                      train.to_csv(str(training_channel) + "/data.csv", index=False)
                      estimator = Estimator(
                      settings.get("image"), # equals `image_2`
                      settings.get("iam_role_arn"),
                      settings.getint('instance_count'),
                      settings.get("instance_type"),
                      base_job_name='model-1-train',
                      volume_size=512,
                      model_uri="file://" + str(model_output),
                      output_path="file://" + str(training_output),
                      sagemaker_session=sagemaker_session,
                      hyperparameters=settings.get("hyperparameters")
                      )
                      estimator.fit({
                      "training": f"file://{training_channel}"
                      }, logs="All", wait=True)
                      assert estimator
                      # Check that model got saved to correct location
                      assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                      # Deploy the model locally
                      predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                      # Check that we got a valid prediction
                      data = ... # some test input
                      response = json.loads(predictor.predict(json.dumps(data)))
                      assert response 

                      Expected behavior
                      I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

                      Screenshots or logs
                      Logs from the second test starting to run:

                      INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
                      Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
                      Traceback (most recent call last):
                      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
                      _stream_output(self.process)
                      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
                      raise RuntimeError("Process exited with code: %s" % exit_code)
                      RuntimeError: Process exited with code: 1
                      During handling of the above exception, another exception occurred:
                      Traceback (most recent call last):
                      File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
                      self.run()
                      File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
                      raise RuntimeError(msg)
                      RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
                      Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
                      

                      I tried running the command in manually in my terminal and got the following :

                      docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
                      Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
                      Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
                      ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                      ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                      ERROR: Encountered errors while bringing up the project.
                      

                      Note: this output is more useful than what comes out of the SDK.

                      System information
                      A description of your system. Please provide:

                      • SageMaker Python SDK version: 2.18.0
                      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
                      • Framework version: 3.1.0
                      • Python version: 3.7
                      • Custom Docker image (Y/N): Y

                      Additional context
                      Add any other context about the problem here.

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          Cannot create multiple model endpoints in local mode #2020

                          Description

                          @pwerth

                          Describe the bug
                          In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

                          This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

                          If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

                          To reproduce
                          I have two images corresponding to two different containers, let's call them image-1 and image-2.

                          In test_train_model1.py:

                          def test_train_and_predict(tmp_path):
                          sagemaker_session = LocalSession()
                          sagemaker_session.config = {'local': {'local_code': True}}
                          training_channel = tmp_path / "training"
                          model_output = tmp_path / "model"
                          training_output = tmp_path / "output"
                          training_channel.mkdir(exist_ok=True)
                          model_output.mkdir(exist_ok=True)
                          training_output.mkdir(exist_ok=True)
                          # Move the test dataset into the location expected by train container
                          train = pd.read_csv(<path_to_local_csv>)
                          train.to_csv(str(training_channel) + "/data.csv", index=False)
                          estimator = Estimator(
                          settings.get("image"), # equals `image_1`
                          settings.get("iam_role_arn"),
                          settings.getint('instance_count'),
                          settings.get("instance_type"),
                          base_job_name='model-1-train',
                          volume_size=512,
                          model_uri="file://" + str(model_output),
                          output_path="file://" + str(training_output),
                          sagemaker_session=sagemaker_session,
                          hyperparameters=settings.get("hyperparameters")
                          )
                          estimator.fit({
                          "training": f"file://{training_channel}"
                          }, logs="All", wait=True)
                          assert estimator
                          # Check that model got saved to correct location
                          assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                          # Deploy the model locally
                          predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                          # Check that we got a valid prediction
                          data = ... # some test input
                          response = json.loads(predictor.predict(json.dumps(data)))
                          assert response 

                          In test_model_2.py (identical logic, different image):

                          def test_train_and_predict(tmp_path):
                          sagemaker_session = LocalSession()
                          sagemaker_session.config = {'local': {'local_code': True}}
                          training_channel = tmp_path / "training"
                          model_output = tmp_path / "model"
                          training_output = tmp_path / "output"
                          training_channel.mkdir(exist_ok=True)
                          model_output.mkdir(exist_ok=True)
                          training_output.mkdir(exist_ok=True)
                          # Move the test dataset into the location expected by train container
                          train = pd.read_csv(<path_to_local_csv>)
                          train.to_csv(str(training_channel) + "/data.csv", index=False)
                          estimator = Estimator(
                          settings.get("image"), # equals `image_2`
                          settings.get("iam_role_arn"),
                          settings.getint('instance_count'),
                          settings.get("instance_type"),
                          base_job_name='model-1-train',
                          volume_size=512,
                          model_uri="file://" + str(model_output),
                          output_path="file://" + str(training_output),
                          sagemaker_session=sagemaker_session,
                          hyperparameters=settings.get("hyperparameters")
                          )
                          estimator.fit({
                          "training": f"file://{training_channel}"
                          }, logs="All", wait=True)
                          assert estimator
                          # Check that model got saved to correct location
                          assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                          # Deploy the model locally
                          predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                          # Check that we got a valid prediction
                          data = ... # some test input
                          response = json.loads(predictor.predict(json.dumps(data)))
                          assert response 

                          Expected behavior
                          I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

                          Screenshots or logs
                          Logs from the second test starting to run:

                          INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
                          Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
                          Traceback (most recent call last):
                          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
                          _stream_output(self.process)
                          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
                          raise RuntimeError("Process exited with code: %s" % exit_code)
                          RuntimeError: Process exited with code: 1
                          During handling of the above exception, another exception occurred:
                          Traceback (most recent call last):
                          File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
                          self.run()
                          File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
                          raise RuntimeError(msg)
                          RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
                          Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
                          

                          I tried running the command in manually in my terminal and got the following :

                          docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
                          Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
                          Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
                          ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                          ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                          ERROR: Encountered errors while bringing up the project.
                          

                          Note: this output is more useful than what comes out of the SDK.

                          System information
                          A description of your system. Please provide:

                          • SageMaker Python SDK version: 2.18.0
                          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
                          • Framework version: 3.1.0
                          • Python version: 3.7
                          • Custom Docker image (Y/N): Y

                          Additional context
                          Add any other context about the problem here.

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              Cannot create multiple model endpoints in local mode #2020

                              Description

                              @pwerth

                              Describe the bug
                              In local mode, two model endpoints cannot be created in the same session. When the second endpoint is created, it tries to use the same port (8080) that is still being occupied by the first endpoint, which results in any calls to predict being routed to the first container.

                              This is breaking my unit tests, because I have tests for model training and predictions that span multiple models. My hunch is that this is not related to Pytest though.

                              If I run either unit test individually, it passes. But I cannot run them together because whichever runs second always fails, because the response has the wrong format - since when I call predict in the second test, it ends up hitting the /invocations endpoint on the first container.

                              To reproduce
                              I have two images corresponding to two different containers, let's call them image-1 and image-2.

                              In test_train_model1.py:

                              def test_train_and_predict(tmp_path):
                              sagemaker_session = LocalSession()
                              sagemaker_session.config = {'local': {'local_code': True}}
                              training_channel = tmp_path / "training"
                              model_output = tmp_path / "model"
                              training_output = tmp_path / "output"
                              training_channel.mkdir(exist_ok=True)
                              model_output.mkdir(exist_ok=True)
                              training_output.mkdir(exist_ok=True)
                              # Move the test dataset into the location expected by train container
                              train = pd.read_csv(<path_to_local_csv>)
                              train.to_csv(str(training_channel) + "/data.csv", index=False)
                              estimator = Estimator(
                              settings.get("image"), # equals `image_1`
                              settings.get("iam_role_arn"),
                              settings.getint('instance_count'),
                              settings.get("instance_type"),
                              base_job_name='model-1-train',
                              volume_size=512,
                              model_uri="file://" + str(model_output),
                              output_path="file://" + str(training_output),
                              sagemaker_session=sagemaker_session,
                              hyperparameters=settings.get("hyperparameters")
                              )
                              estimator.fit({
                              "training": f"file://{training_channel}"
                              }, logs="All", wait=True)
                              assert estimator
                              # Check that model got saved to correct location
                              assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                              # Deploy the model locally
                              predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                              # Check that we got a valid prediction
                              data = ... # some test input
                              response = json.loads(predictor.predict(json.dumps(data)))
                              assert response 

                              In test_model_2.py (identical logic, different image):

                              def test_train_and_predict(tmp_path):
                              sagemaker_session = LocalSession()
                              sagemaker_session.config = {'local': {'local_code': True}}
                              training_channel = tmp_path / "training"
                              model_output = tmp_path / "model"
                              training_output = tmp_path / "output"
                              training_channel.mkdir(exist_ok=True)
                              model_output.mkdir(exist_ok=True)
                              training_output.mkdir(exist_ok=True)
                              # Move the test dataset into the location expected by train container
                              train = pd.read_csv(<path_to_local_csv>)
                              train.to_csv(str(training_channel) + "/data.csv", index=False)
                              estimator = Estimator(
                              settings.get("image"), # equals `image_2`
                              settings.get("iam_role_arn"),
                              settings.getint('instance_count'),
                              settings.get("instance_type"),
                              base_job_name='model-1-train',
                              volume_size=512,
                              model_uri="file://" + str(model_output),
                              output_path="file://" + str(training_output),
                              sagemaker_session=sagemaker_session,
                              hyperparameters=settings.get("hyperparameters")
                              )
                              estimator.fit({
                              "training": f"file://{training_channel}"
                              }, logs="All", wait=True)
                              assert estimator
                              # Check that model got saved to correct location
                              assert estimator.model_data == f'file://{training_output}/model.tar.gz'
                              # Deploy the model locally
                              predictor = estimator.deploy(initial_instance_count=1, instance_type='local')
                              # Check that we got a valid prediction
                              data = ... # some test input
                              response = json.loads(predictor.predict(json.dumps(data)))
                              assert response 

                              Expected behavior
                              I expect both containers to be built, both models to train, and both predictions to be served. When I call predict on the second estimator, it should hit the /invocations endpoint on the second container.

                              Screenshots or logs
                              Logs from the second test starting to run:

                              INFO sagemaker.local.image:image.py:508 docker command: docker-compose -f /private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmphvpfolv1/docker-compose.yaml up --build --abort-on-container-exit
                              Creating tmphvpfolv1_algo-1-n4k0x_1 ... Exception in thread Thread-1:
                              Traceback (most recent call last):
                              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 627, in run
                              _stream_output(self.process)
                              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 687, in _stream_output
                              raise RuntimeError("Process exited with code: %s" % exit_code)
                              RuntimeError: Process exited with code: 1
                              During handling of the above exception, another exception occurred:
                              Traceback (most recent call last):
                              File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.7/lib/python3.7/threading.py", line 917, in _bootstrap_inner
                              self.run()
                              File "/Users/me/Documents/code/my-repo/venv/lib/python3.7/site-packages/sagemaker-2.18.0-py3.7.egg/sagemaker/local/image.py", line 632, in run
                              raise RuntimeError(msg)
                              RuntimeError: Failed to run: ['docker-compose', '-f', '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml', 'up', '--build', '--abort-on-container-exit'], Process exited with code: 1
                              Creating tmphvpfolv1_algo-1-n4k0x_1 ... done
                              

                              I tried running the command in manually in my terminal and got the following :

                              docker-compose -f '/private/var/folders/yq/nt_pyt5112b702866l2l4zrm0000gn/T/tmpkq9_j1e0/docker-compose.yaml' up --build --abort-on-container-exit
                              Starting tmpkq9_j1eo_algo-1-6cyan_1 ...
                              Starting tmpkq9_j1eo_algo-1-6cyan_1 ... error
                              ERROR: for tmpkq9_j1eo_algo-1-6cyan_1 Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                              ERROR: for algo-1-6cyan Cannot start service algo-1-6cyan: driver failed programming external connectivity on endpoint tmpkq9_j1e0_algo-1-6cyan_1 (7c68cf8e050f0c06aa99532dc52335a6961b298ec19c328068944d3504ed98): Bind for 0.0.0.0:8080 failed: port is already allocated
                              ERROR: Encountered errors while bringing up the project.
                              

                              Note: this output is more useful than what comes out of the SDK.

                              System information
                              A description of your system. Please provide:

                              • SageMaker Python SDK version: 2.18.0
                              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): Lightgbm
                              • Framework version: 3.1.0
                              • Python version: 3.7
                              • Custom Docker image (Y/N): Y

                              Additional context
                              Add any other context about the problem here.

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions