[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode - #529

Closed
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod
Closed

[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script Mode#529
icywang86rui wants to merge 16 commits into
aws:masterfrom
icywang86rui:horovod

Conversation

@icywang86rui

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Dec 6, 2018

Copy link
Copy Markdown

Codecov Report

Merging #529 into master will increase coverage by 0.01%.
The diff coverage is 100%.

Impacted file tree graph

@@ Coverage Diff @@## master #529 +/- ##
==========================================
+ Coverage 92.79% 92.81% +0.01% 
==========================================
Files 71 71 Lines 5373 5386 +13 ==========================================
+ Hits 4986 4999 +13 
Misses 387 387
Impacted FilesCoverage Δ
src/sagemaker/tensorflow/defaults.py100% <100%> (ø)⬆️
src/sagemaker/tensorflow/estimator.py94.92% <100%> (+0.27%)⬆️
src/sagemaker/estimator.py90.35% <100%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 4ffdeda...3cfdc57. Read the comment docs.

@laurenyu

Copy link
Copy Markdown
Contributor

please make the PR title an imperative statement

@icywang86ruiicywang86rui changed the title HorovodEnable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skip(reason='The containers have not been updated in Prod yet.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume we're not going to merge the PR until the containers are released? let's remove this skip

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point. I will remove it.

@icywang86ruiicywang86rui changed the title Enable distributed training with Horovod for TensorFlow Script Mode[DO NOT MERGE] Enable distributed training with Horovod for TensorFlow Script ModeDec 6, 2018
Comment threadsrc/sagemaker/estimator.py Outdated

__framework_name__ = None
LAUNCH_PS_ENV_NAME = 'sagemaker_parameter_server_enabled'
USE_MPI_ENV_NAME = 'sagemaker_mpi_enabled'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: what about LAUNCH_MPI_ENV_NAME to be consistent with the other name?

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
''''''''''''''''''''

To run your training job in a distributed fashion you need to set ``train_instance_count`` to a number larger than 1.
We support two different types of distributed training, parameter server and MPI. The ``distributions`` parameter is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We support more than these two distributed training types. I guess the difference is that these two types require additional setup.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's possible for user to run other types of distributed training. But I wouldn't say those are supported. These two types are setup by our code and we are going to support and maintain that code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we need to provide details about the 2 configurations options custom_mpi_options and processes_per_host, here is the draft that I wrote:

Please see distribution and Training with Horovod sections of https://github.com/uditbhatia/sagemaker-python-sdk/blob/horovod-documentation/src/sagemaker/tensorflow/README.rst

PLease note couple of links are broken as it is still a draft. But I hope this helps you.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re: Marcio's initial concern - I think this would read better if the "supporting two types of distributed training" were attributed to distributions rather than "us" (aka SageMaker). So maybe change this to:

To run your training job in a distributed fashion you need to set train_instance_count to a number larger than 1. In addition, you will need to ensure that the correct processes are started during training. You can either do this yourself or use the distributions parameter.

The distributions parameter can be used for:

  • launching parameter server: blah blah blah explanation
  • using MPI: other explanation blah blah blah

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={'mpi': {'enabled': True}})
tf_estimator.fit('s3://bucket/path/to/training/data')

If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If MPI is enabled the container will construct and run MPI commands which executes your training script. You can find
If MPI is enabled the container will configure and execute `mpirun` with your training script. You can find

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ['SM_OUTPUT_DATA_DIR'])
parser.add_argument('--model_dir', type=str, default=os.environ['SM_MODEL_DIR'])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's remove the default here given that is passed through the hyperparameters.
Use os.environ.get instead to avoid errors running the script outside SageMaker.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

k

hvd.init()

# Download and load MNIST dataset.
mnist = learn.datasets.mnist.read_data_sets('MNIST-data-%d' % hvd.rank())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the size of the dataset?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the training data is 164M. With the eval data and the label it's about 200M.

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is running a script that uses MPI but not Horovod a use case?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is but if you use the tensorflow container it uses Horovod. I could be wrong.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if that's the case, then I think this should be changed to something like "We support two different ways of handling distributed training: parameter servers and MPI. The use of MPI can be with or without Horovod." maybe include a link to Horovod documentation as well.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated

Training with ``MPI`` is configured by specifying following fields in ``distributions``:

- ``enabled (bool)``: If set to `True`, the MPI setup is performed and ``mpirun`` command is executed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double backticks for True

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
tf_estimator = TensorFlow(entry_point='tf-train.py', role='SageMakerRole',
train_instance_count=1, train_instance_type='ml.p2.xlarge',
framework_version='1.11', py_version='py3',
distributions: {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

line up the arguments, and also s/: /=

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
"mpi":{
"enabled":True,
"processes_per_host":2,
"custom_mpi_options": "--NCCL_DEBUG INFO"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
distributions={
"mpi":{
"enabled":True,
"processes_per_host":2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

spaces after the colons

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/server/servers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

if self.sagemaker_session.local_mode and local_code:
return '/opt/ml/shared/{}'.format(directory)
elif mpi:
return '/opt/ml/model'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make this a constant?

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for strings

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Distributed Training
''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...in a distributed fashion, set...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with parameter servers
"""""""""""""""""""""""""""""""

If parameter server is enabled, the container will launch a parameter server thread in each instance first then execute

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server thread on each instance in the training cluster, and then executes your training code. You can..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

Comment threadsrc/sagemaker/tensorflow/README.rst Outdated
Training with Horovod
"""""""""""""""""""""

Horovod is a distributed training framework based on MPI. You can find more details in `Horovod README <https://github.com/uber/horovod>`__.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"...more details at..."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated

@laurenyulaurenyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all small comments. otherwise lgtm.

Training with parameter servers
"""""""""""""""""""""""""""""""

If you specify parameter_server as the value of the distributions parameter, the container launches a parameter server

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

backticks around parameter_server

use the following setup:
distributions (dict): A dictionary with information on how to run distributed training
(default: None). Currently we support distributed training with parameter servers and MPI. To enable
parameter server use the following setup:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for "To enable parameter server" - s/server/servers

Comment threadtests/integ/test_tf_script_mode.py Outdated
['graph.pbtxt', 'model.ckpt-0.index', 'model.ckpt-0.meta', 'saved_model.pb'])


@pytest.mark.skipif(integ.PYTHON_VERSION != 'py3', reason="Script Mode tests are only configured to run with Python 3")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

single quotes for the reason string

parser = argparse.ArgumentParser()
# Data, model, and output directories
parser.add_argument('--output-data-dir', type=str, default=os.environ.get('SM_OUTPUT_DATA_DIR'))
parser.add_argument('--model_dir', type=str)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: it's strange to me that we would mix underscores and hyphens in our examples like this

''''''''''''''''''''

To run your training job with multiple instances in a distributed fashion you need to set ``train_instance_count``
to a number larger than 1. We support two different types of distributed training, parameter server and Horovod.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MPI without Horovod is a valid use case.

@mvsuspmvsusp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I previously approved this PR by mistake

@mvsuspmvsusp closed this Dec 20, 2018
metrizable pushed a commit to metrizable/sagemaker-python-sdk that referenced this pull request Dec 1, 2020
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@icywang86rui@codecov-io@laurenyu@mvsusp@uditbhatia@yangaws@eslesar-aws