Horovod multi-node fails to connect and hangs indefinitely #1369

Description

@jarednielsen

Describe the bug
Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

To reproduce
Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

from sagemaker.tensorflow import TensorFlow
from sagemaker.inputs import FileSystemInput
role = ...
image_name = "jarednielsen/albert-tf:sagemaker"
fsx_id = ...
hvd_instance_type = "ml.p3.16xlarge"
hvd_processes_per_host = 8
hvd_instance_count = 8
batch_size = 8
distributions = {
"mpi": {
"enabled": True,
"processes_per_host": hvd_processes_per_host,
"custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
}
}
hyperparameters = {
"model_size": "base",
"batch_size": batch_size,
"max_seq_length": 512,
"gradient_accumulation_steps": 1,
"learning_rate": 0.00176,
"optimizer": "lamb",
"fsx_prefix": "/opt/ml/input/data/training",
"name": "sagemaker",
}
estimator_hvd = TensorFlow(
entry_point="/path/to/blank/file.py",
role=role,
framework_version="2.1.0",
py_version="py3",
hyperparameters=hyperparameters,
train_instance_count=hvd_instance_count,
train_instance_type=hvd_instance_type,
distributions=distributions,
image_name=image_name,
subnets=[subnet_id],
security_group_ids=[security_group_id],
enable_sagemaker_metrics=True,
)
fsx_input = FileSystemInput(
file_system_id=fsx_id,
file_system_type="FSxLustre",
directory_path="/fsx",
file_system_access_mode="rw",
)
estimator_hvd.fit(fsx_input)

Expected behavior
It to not hang :)

Screenshots or logs

$ $ python run_sagemaker.py
2020-03-20 00:49:47 Starting - Starting the training job...
2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
2020-03-20 00:56:58 Starting - Preparing the instances for training.........
2020-03-20 00:58:47 Downloading - Downloading input data
2020-03-20 00:58:47 Training - Downloading the training image...............
2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
[repeated 64 times]

System information
A description of your system. Please provide:

  • SageMaker Python SDK version: 1.50.14
  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
  • Framework version: 2.1
  • Python version: 3.7
  • CPU or GPU: GPU
  • Custom Docker image (Y/N): Yes

Additional context
The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py

This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

# Below here is necessary to install SSH on SageMaker machines
RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
RUN mkdir -p /root/.ssh/ && \
ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
# Allow OpenSSH to talk to containers without asking for confirmation
RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
&& echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
&& mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config

Anything more I need to do?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      Horovod multi-node fails to connect and hangs indefinitely #1369

      Description

      @jarednielsen

      Describe the bug
      Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

      To reproduce
      Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

      from sagemaker.tensorflow import TensorFlow
      from sagemaker.inputs import FileSystemInput
      role = ...
      image_name = "jarednielsen/albert-tf:sagemaker"
      fsx_id = ...
      hvd_instance_type = "ml.p3.16xlarge"
      hvd_processes_per_host = 8
      hvd_instance_count = 8
      batch_size = 8
      distributions = {
      "mpi": {
      "enabled": True,
      "processes_per_host": hvd_processes_per_host,
      "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
      }
      }
      hyperparameters = {
      "model_size": "base",
      "batch_size": batch_size,
      "max_seq_length": 512,
      "gradient_accumulation_steps": 1,
      "learning_rate": 0.00176,
      "optimizer": "lamb",
      "fsx_prefix": "/opt/ml/input/data/training",
      "name": "sagemaker",
      }
      estimator_hvd = TensorFlow(
      entry_point="/path/to/blank/file.py",
      role=role,
      framework_version="2.1.0",
      py_version="py3",
      hyperparameters=hyperparameters,
      train_instance_count=hvd_instance_count,
      train_instance_type=hvd_instance_type,
      distributions=distributions,
      image_name=image_name,
      subnets=[subnet_id],
      security_group_ids=[security_group_id],
      enable_sagemaker_metrics=True,
      )
      fsx_input = FileSystemInput(
      file_system_id=fsx_id,
      file_system_type="FSxLustre",
      directory_path="/fsx",
      file_system_access_mode="rw",
      )
      estimator_hvd.fit(fsx_input)
      

      Expected behavior
      It to not hang :)

      Screenshots or logs

      $ $ python run_sagemaker.py
      2020-03-20 00:49:47 Starting - Starting the training job...
      2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
      2020-03-20 00:56:58 Starting - Preparing the instances for training.........
      2020-03-20 00:58:47 Downloading - Downloading input data
      2020-03-20 00:58:47 Training - Downloading the training image...............
      2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
      2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
      2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
      2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
      2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
      2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
      [repeated 64 times]
      

      System information
      A description of your system. Please provide:

      • SageMaker Python SDK version: 1.50.14
      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
      • Framework version: 2.1
      • Python version: 3.7
      • CPU or GPU: GPU
      • Custom Docker image (Y/N): Yes

      Additional context
      The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

      ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
      

      This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

      I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

      # Below here is necessary to install SSH on SageMaker machines
      RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
      RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
      RUN mkdir -p /root/.ssh/ && \
      ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
      cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
      printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
      # Allow OpenSSH to talk to containers without asking for confirmation
      RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
      && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
      && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
      

      Anything more I need to do?

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          Horovod multi-node fails to connect and hangs indefinitely #1369

          Description

          @jarednielsen

          Describe the bug
          Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

          To reproduce
          Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

          from sagemaker.tensorflow import TensorFlow
          from sagemaker.inputs import FileSystemInput
          role = ...
          image_name = "jarednielsen/albert-tf:sagemaker"
          fsx_id = ...
          hvd_instance_type = "ml.p3.16xlarge"
          hvd_processes_per_host = 8
          hvd_instance_count = 8
          batch_size = 8
          distributions = {
          "mpi": {
          "enabled": True,
          "processes_per_host": hvd_processes_per_host,
          "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
          }
          }
          hyperparameters = {
          "model_size": "base",
          "batch_size": batch_size,
          "max_seq_length": 512,
          "gradient_accumulation_steps": 1,
          "learning_rate": 0.00176,
          "optimizer": "lamb",
          "fsx_prefix": "/opt/ml/input/data/training",
          "name": "sagemaker",
          }
          estimator_hvd = TensorFlow(
          entry_point="/path/to/blank/file.py",
          role=role,
          framework_version="2.1.0",
          py_version="py3",
          hyperparameters=hyperparameters,
          train_instance_count=hvd_instance_count,
          train_instance_type=hvd_instance_type,
          distributions=distributions,
          image_name=image_name,
          subnets=[subnet_id],
          security_group_ids=[security_group_id],
          enable_sagemaker_metrics=True,
          )
          fsx_input = FileSystemInput(
          file_system_id=fsx_id,
          file_system_type="FSxLustre",
          directory_path="/fsx",
          file_system_access_mode="rw",
          )
          estimator_hvd.fit(fsx_input)
          

          Expected behavior
          It to not hang :)

          Screenshots or logs

          $ $ python run_sagemaker.py
          2020-03-20 00:49:47 Starting - Starting the training job...
          2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
          2020-03-20 00:56:58 Starting - Preparing the instances for training.........
          2020-03-20 00:58:47 Downloading - Downloading input data
          2020-03-20 00:58:47 Training - Downloading the training image...............
          2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
          2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
          2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
          2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
          2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
          2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
          [repeated 64 times]
          

          System information
          A description of your system. Please provide:

          • SageMaker Python SDK version: 1.50.14
          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
          • Framework version: 2.1
          • Python version: 3.7
          • CPU or GPU: GPU
          • Custom Docker image (Y/N): Yes

          Additional context
          The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

          ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
          

          This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

          I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

          # Below here is necessary to install SSH on SageMaker machines
          RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
          RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
          RUN mkdir -p /root/.ssh/ && \
          ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
          cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
          printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
          # Allow OpenSSH to talk to containers without asking for confirmation
          RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
          && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
          && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
          

          Anything more I need to do?

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Horovod multi-node fails to connect and hangs indefinitely #1369

              Description

              @jarednielsen

              Describe the bug
              Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

              To reproduce
              Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

              from sagemaker.tensorflow import TensorFlow
              from sagemaker.inputs import FileSystemInput
              role = ...
              image_name = "jarednielsen/albert-tf:sagemaker"
              fsx_id = ...
              hvd_instance_type = "ml.p3.16xlarge"
              hvd_processes_per_host = 8
              hvd_instance_count = 8
              batch_size = 8
              distributions = {
              "mpi": {
              "enabled": True,
              "processes_per_host": hvd_processes_per_host,
              "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
              }
              }
              hyperparameters = {
              "model_size": "base",
              "batch_size": batch_size,
              "max_seq_length": 512,
              "gradient_accumulation_steps": 1,
              "learning_rate": 0.00176,
              "optimizer": "lamb",
              "fsx_prefix": "/opt/ml/input/data/training",
              "name": "sagemaker",
              }
              estimator_hvd = TensorFlow(
              entry_point="/path/to/blank/file.py",
              role=role,
              framework_version="2.1.0",
              py_version="py3",
              hyperparameters=hyperparameters,
              train_instance_count=hvd_instance_count,
              train_instance_type=hvd_instance_type,
              distributions=distributions,
              image_name=image_name,
              subnets=[subnet_id],
              security_group_ids=[security_group_id],
              enable_sagemaker_metrics=True,
              )
              fsx_input = FileSystemInput(
              file_system_id=fsx_id,
              file_system_type="FSxLustre",
              directory_path="/fsx",
              file_system_access_mode="rw",
              )
              estimator_hvd.fit(fsx_input)
              

              Expected behavior
              It to not hang :)

              Screenshots or logs

              $ $ python run_sagemaker.py
              2020-03-20 00:49:47 Starting - Starting the training job...
              2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
              2020-03-20 00:56:58 Starting - Preparing the instances for training.........
              2020-03-20 00:58:47 Downloading - Downloading input data
              2020-03-20 00:58:47 Training - Downloading the training image...............
              2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
              2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
              2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
              2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
              2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
              2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
              [repeated 64 times]
              

              System information
              A description of your system. Please provide:

              • SageMaker Python SDK version: 1.50.14
              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
              • Framework version: 2.1
              • Python version: 3.7
              • CPU or GPU: GPU
              • Custom Docker image (Y/N): Yes

              Additional context
              The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

              ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
              

              This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

              I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

              # Below here is necessary to install SSH on SageMaker machines
              RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
              RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
              RUN mkdir -p /root/.ssh/ && \
              ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
              cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
              printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
              # Allow OpenSSH to talk to containers without asking for confirmation
              RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
              && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
              && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
              

              Anything more I need to do?

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  Horovod multi-node fails to connect and hangs indefinitely #1369

                  Description

                  @jarednielsen

                  Describe the bug
                  Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

                  To reproduce
                  Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

                  from sagemaker.tensorflow import TensorFlow
                  from sagemaker.inputs import FileSystemInput
                  role = ...
                  image_name = "jarednielsen/albert-tf:sagemaker"
                  fsx_id = ...
                  hvd_instance_type = "ml.p3.16xlarge"
                  hvd_processes_per_host = 8
                  hvd_instance_count = 8
                  batch_size = 8
                  distributions = {
                  "mpi": {
                  "enabled": True,
                  "processes_per_host": hvd_processes_per_host,
                  "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
                  }
                  }
                  hyperparameters = {
                  "model_size": "base",
                  "batch_size": batch_size,
                  "max_seq_length": 512,
                  "gradient_accumulation_steps": 1,
                  "learning_rate": 0.00176,
                  "optimizer": "lamb",
                  "fsx_prefix": "/opt/ml/input/data/training",
                  "name": "sagemaker",
                  }
                  estimator_hvd = TensorFlow(
                  entry_point="/path/to/blank/file.py",
                  role=role,
                  framework_version="2.1.0",
                  py_version="py3",
                  hyperparameters=hyperparameters,
                  train_instance_count=hvd_instance_count,
                  train_instance_type=hvd_instance_type,
                  distributions=distributions,
                  image_name=image_name,
                  subnets=[subnet_id],
                  security_group_ids=[security_group_id],
                  enable_sagemaker_metrics=True,
                  )
                  fsx_input = FileSystemInput(
                  file_system_id=fsx_id,
                  file_system_type="FSxLustre",
                  directory_path="/fsx",
                  file_system_access_mode="rw",
                  )
                  estimator_hvd.fit(fsx_input)
                  

                  Expected behavior
                  It to not hang :)

                  Screenshots or logs

                  $ $ python run_sagemaker.py
                  2020-03-20 00:49:47 Starting - Starting the training job...
                  2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
                  2020-03-20 00:56:58 Starting - Preparing the instances for training.........
                  2020-03-20 00:58:47 Downloading - Downloading input data
                  2020-03-20 00:58:47 Training - Downloading the training image...............
                  2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
                  2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
                  2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
                  2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                  2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
                  2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
                  [repeated 64 times]
                  

                  System information
                  A description of your system. Please provide:

                  • SageMaker Python SDK version: 1.50.14
                  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
                  • Framework version: 2.1
                  • Python version: 3.7
                  • CPU or GPU: GPU
                  • Custom Docker image (Y/N): Yes

                  Additional context
                  The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

                  ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
                  

                  This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

                  I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

                  # Below here is necessary to install SSH on SageMaker machines
                  RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
                  RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
                  RUN mkdir -p /root/.ssh/ && \
                  ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
                  cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
                  printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
                  # Allow OpenSSH to talk to containers without asking for confirmation
                  RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
                  && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
                  && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
                  

                  Anything more I need to do?

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      Horovod multi-node fails to connect and hangs indefinitely #1369

                      Description

                      @jarednielsen

                      Describe the bug
                      Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

                      To reproduce
                      Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

                      from sagemaker.tensorflow import TensorFlow
                      from sagemaker.inputs import FileSystemInput
                      role = ...
                      image_name = "jarednielsen/albert-tf:sagemaker"
                      fsx_id = ...
                      hvd_instance_type = "ml.p3.16xlarge"
                      hvd_processes_per_host = 8
                      hvd_instance_count = 8
                      batch_size = 8
                      distributions = {
                      "mpi": {
                      "enabled": True,
                      "processes_per_host": hvd_processes_per_host,
                      "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
                      }
                      }
                      hyperparameters = {
                      "model_size": "base",
                      "batch_size": batch_size,
                      "max_seq_length": 512,
                      "gradient_accumulation_steps": 1,
                      "learning_rate": 0.00176,
                      "optimizer": "lamb",
                      "fsx_prefix": "/opt/ml/input/data/training",
                      "name": "sagemaker",
                      }
                      estimator_hvd = TensorFlow(
                      entry_point="/path/to/blank/file.py",
                      role=role,
                      framework_version="2.1.0",
                      py_version="py3",
                      hyperparameters=hyperparameters,
                      train_instance_count=hvd_instance_count,
                      train_instance_type=hvd_instance_type,
                      distributions=distributions,
                      image_name=image_name,
                      subnets=[subnet_id],
                      security_group_ids=[security_group_id],
                      enable_sagemaker_metrics=True,
                      )
                      fsx_input = FileSystemInput(
                      file_system_id=fsx_id,
                      file_system_type="FSxLustre",
                      directory_path="/fsx",
                      file_system_access_mode="rw",
                      )
                      estimator_hvd.fit(fsx_input)
                      

                      Expected behavior
                      It to not hang :)

                      Screenshots or logs

                      $ $ python run_sagemaker.py
                      2020-03-20 00:49:47 Starting - Starting the training job...
                      2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
                      2020-03-20 00:56:58 Starting - Preparing the instances for training.........
                      2020-03-20 00:58:47 Downloading - Downloading input data
                      2020-03-20 00:58:47 Training - Downloading the training image...............
                      2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
                      2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
                      2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
                      2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                      2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
                      2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
                      [repeated 64 times]
                      

                      System information
                      A description of your system. Please provide:

                      • SageMaker Python SDK version: 1.50.14
                      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
                      • Framework version: 2.1
                      • Python version: 3.7
                      • CPU or GPU: GPU
                      • Custom Docker image (Y/N): Yes

                      Additional context
                      The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

                      ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
                      

                      This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

                      I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

                      # Below here is necessary to install SSH on SageMaker machines
                      RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
                      RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
                      RUN mkdir -p /root/.ssh/ && \
                      ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
                      cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
                      printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
                      # Allow OpenSSH to talk to containers without asking for confirmation
                      RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
                      && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
                      && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
                      

                      Anything more I need to do?

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          Horovod multi-node fails to connect and hangs indefinitely #1369

                          Description

                          @jarednielsen

                          Describe the bug
                          Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

                          To reproduce
                          Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

                          from sagemaker.tensorflow import TensorFlow
                          from sagemaker.inputs import FileSystemInput
                          role = ...
                          image_name = "jarednielsen/albert-tf:sagemaker"
                          fsx_id = ...
                          hvd_instance_type = "ml.p3.16xlarge"
                          hvd_processes_per_host = 8
                          hvd_instance_count = 8
                          batch_size = 8
                          distributions = {
                          "mpi": {
                          "enabled": True,
                          "processes_per_host": hvd_processes_per_host,
                          "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
                          }
                          }
                          hyperparameters = {
                          "model_size": "base",
                          "batch_size": batch_size,
                          "max_seq_length": 512,
                          "gradient_accumulation_steps": 1,
                          "learning_rate": 0.00176,
                          "optimizer": "lamb",
                          "fsx_prefix": "/opt/ml/input/data/training",
                          "name": "sagemaker",
                          }
                          estimator_hvd = TensorFlow(
                          entry_point="/path/to/blank/file.py",
                          role=role,
                          framework_version="2.1.0",
                          py_version="py3",
                          hyperparameters=hyperparameters,
                          train_instance_count=hvd_instance_count,
                          train_instance_type=hvd_instance_type,
                          distributions=distributions,
                          image_name=image_name,
                          subnets=[subnet_id],
                          security_group_ids=[security_group_id],
                          enable_sagemaker_metrics=True,
                          )
                          fsx_input = FileSystemInput(
                          file_system_id=fsx_id,
                          file_system_type="FSxLustre",
                          directory_path="/fsx",
                          file_system_access_mode="rw",
                          )
                          estimator_hvd.fit(fsx_input)
                          

                          Expected behavior
                          It to not hang :)

                          Screenshots or logs

                          $ $ python run_sagemaker.py
                          2020-03-20 00:49:47 Starting - Starting the training job...
                          2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
                          2020-03-20 00:56:58 Starting - Preparing the instances for training.........
                          2020-03-20 00:58:47 Downloading - Downloading input data
                          2020-03-20 00:58:47 Training - Downloading the training image...............
                          2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
                          2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
                          2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
                          2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                          2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
                          2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
                          [repeated 64 times]
                          

                          System information
                          A description of your system. Please provide:

                          • SageMaker Python SDK version: 1.50.14
                          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
                          • Framework version: 2.1
                          • Python version: 3.7
                          • CPU or GPU: GPU
                          • Custom Docker image (Y/N): Yes

                          Additional context
                          The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

                          ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
                          

                          This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

                          I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

                          # Below here is necessary to install SSH on SageMaker machines
                          RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
                          RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
                          RUN mkdir -p /root/.ssh/ && \
                          ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
                          cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
                          printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
                          # Allow OpenSSH to talk to containers without asking for confirmation
                          RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
                          && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
                          && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
                          

                          Anything more I need to do?

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              Horovod multi-node fails to connect and hangs indefinitely #1369

                              Description

                              @jarednielsen

                              Describe the bug
                              Running a horovod tensorflow job with multiple nodes gives a "Cannot connect to host algo-1" error and hangs indefinitely. The horovod job runs successfully if I specify a single node and multiple processes. I have been able to run horovod multi-node training with the same script outside of SageMaker.

                              To reproduce
                              Bit complex to add all the scaffolding, but I can put together an reproducible example if necessary. The gist of it is

                              from sagemaker.tensorflow import TensorFlow
                              from sagemaker.inputs import FileSystemInput
                              role = ...
                              image_name = "jarednielsen/albert-tf:sagemaker"
                              fsx_id = ...
                              hvd_instance_type = "ml.p3.16xlarge"
                              hvd_processes_per_host = 8
                              hvd_instance_count = 8
                              batch_size = 8
                              distributions = {
                              "mpi": {
                              "enabled": True,
                              "processes_per_host": hvd_processes_per_host,
                              "custom_mpi_options": "-verbose --NCCL_DEBUG=INFO -x OMPI_MCA_btl_vader_single_copy_mechanism=none",
                              }
                              }
                              hyperparameters = {
                              "model_size": "base",
                              "batch_size": batch_size,
                              "max_seq_length": 512,
                              "gradient_accumulation_steps": 1,
                              "learning_rate": 0.00176,
                              "optimizer": "lamb",
                              "fsx_prefix": "/opt/ml/input/data/training",
                              "name": "sagemaker",
                              }
                              estimator_hvd = TensorFlow(
                              entry_point="/path/to/blank/file.py",
                              role=role,
                              framework_version="2.1.0",
                              py_version="py3",
                              hyperparameters=hyperparameters,
                              train_instance_count=hvd_instance_count,
                              train_instance_type=hvd_instance_type,
                              distributions=distributions,
                              image_name=image_name,
                              subnets=[subnet_id],
                              security_group_ids=[security_group_id],
                              enable_sagemaker_metrics=True,
                              )
                              fsx_input = FileSystemInput(
                              file_system_id=fsx_id,
                              file_system_type="FSxLustre",
                              directory_path="/fsx",
                              file_system_access_mode="rw",
                              )
                              estimator_hvd.fit(fsx_input)
                              

                              Expected behavior
                              It to not hang :)

                              Screenshots or logs

                              $ $ python run_sagemaker.py
                              2020-03-20 00:49:47 Starting - Starting the training job...
                              2020-03-20 00:49:50 Starting - Launching requested ML instances.......................................
                              2020-03-20 00:56:58 Starting - Preparing the instances for training.........
                              2020-03-20 00:58:47 Downloading - Downloading input data
                              2020-03-20 00:58:47 Training - Downloading the training image...............
                              2020-03-20 01:01:21 Training - Training image download completed. Training in progress.2020-03-20 01:01:22,880 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:22,880 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:23,049 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:23,049 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:22,603 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:22,603 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:23,981 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:23,982 sagemaker-containers INFO Creating SSH daemon.
                              2020-03-20 01:01:23,986 sagemaker-containers INFO Waiting for MPI workers to establish their SSH connections
                              2020-03-20 01:01:23,361 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:23,361 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:25,144 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:25,144 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:22,417 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:22,417 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:01:24,060 sagemaker-containers INFO Starting MPI run as worker node.
                              2020-03-20 01:01:24,061 sagemaker-containers INFO Waiting for MPI Master to create SSH daemon.
                              2020-03-20 01:03:32,823 sagemaker-containers INFO Cannot connect to host algo-1
                              2020-03-20 01:03:32,824 sagemaker-containers INFO Connection failed with exception: [Errno 110] Connection timed out
                              [repeated 64 times]
                              

                              System information
                              A description of your system. Please provide:

                              • SageMaker Python SDK version: 1.50.14
                              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): TensorFlow,
                              • Framework version: 2.1
                              • Python version: 3.7
                              • CPU or GPU: GPU
                              • Custom Docker image (Y/N): Yes

                              Additional context
                              The one lead I can think of is that I'm doing something a little different with the entrypoint script. Instead of specifying it in the Tensorflow() constructor, I specify it in the Dockerfile with

                              ENV SAGEMAKER_PROGRAM /opt/ml/input/data/training/myscript.py
                              

                              This is a quirk specific to my situation, but works fine on single-node. Anything I should dive into to investigate the Horovod hanging issue?

                              I have the following in my Dockerfile, following the lead of https://github.com/aws/sagemaker-tensorflow-container/blob/master/docker/1.15.2/py3/Dockerfile.gpu

                              # Below here is necessary to install SSH on SageMaker machines
                              RUN apt-get update && apt-get install -y --no-install-recommends openssh-server && mkdir -p /var/run/sshd
                              RUN sed 's@session\s*required\s*pam_loginuid.so@session optional pam_loginuid.so@g' -i /etc/pam.d/sshd
                              RUN mkdir -p /root/.ssh/ && \
                              ssh-keygen -q -t rsa -N '' -f /root/.ssh/id_rsa && \
                              cp /root/.ssh/id_rsa.pub /root/.ssh/authorized_keys && \
                              printf "Host * StrictHostKeyChecking no" >> /root/.ssh/config
                              # Allow OpenSSH to talk to containers without asking for confirmation
                              RUN cat /etc/ssh/ssh_config | grep -v StrictHostKeyChecking > /etc/ssh/ssh_config.new \
                              && echo " StrictHostKeyChecking no" >> /etc/ssh/ssh_config.new \
                              && mv /etc/ssh/ssh_config.new /etc/ssh/ssh_config
                              

                              Anything more I need to do?

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions