Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add APIs to export transform and deploy config - #497

Merged
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference
Nov 18, 2018
Merged

Add APIs to export transform and deploy config#497
yangaws merged 11 commits into
aws:masterfrom
yangaws:airflow_inference

Conversation

@yangaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

  1. Add APIs to export transform config from a transformer or an estimator.
  2. Add APIs to export deploy config from a model or an estimator.
  3. Add related unit tests.

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

  • I have read the CONTRIBUTING doc
  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have updated the changelog with a description of my changes (if appropriate)
  • I have updated any necessary documentation (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov-io

codecov-io commented Nov 16, 2018

Copy link
Copy Markdown

Codecov Report

Merging #497 into master will decrease coverage by <.01%.
The diff coverage is 95.23%.

Impacted file tree graph

@@ Coverage Diff @@## master #497 +/- ##
==========================================
- Coverage 94.28% 94.28% -0.01% 
==========================================
Files 59 59 Lines 4551 4603 +52 ==========================================
+ Hits 4291 4340 +49 - Misses 260 263 +3
Impacted FilesCoverage Δ
src/sagemaker/transformer.py100% <ø> (ø)⬆️
src/sagemaker/estimator.py90.47% <100%> (+0.16%)⬆️
src/sagemaker/workflow/airflow.py92.09% <93.61%> (+0.55%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update e37ac12...d1a005e. Read the comment docs.

model (sagemaker.model.FrameworkModel): The framework model
instance_type (str): The EC2 instance type to deploy this Model to. For example, 'ml.p2.xlarge'.
s3_operations (dict): The dict to specify S3 operations (upload `source_dir`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these newlines should stay

If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.
If None, server will use one worker per vCPU. Only effective when estimator is
SageMaker framework.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

s/is/is a

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated.

Comment threadsrc/sagemaker/workflow/airflow.py Outdated
vpc_config_override=vpc_config_override)
else:
raise TypeError('Estimator must be one of BYO estimator, framework estimator or amazon algorithm'
'estimator.')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might be more helpful here to put the paths to the classes themselves, i.e. sagemaker.estimator.Estimator, etc.


if transformer.output_path is None:
transformer.output_path = 's3://{}/{}'.format(
transformer.sagemaker_session.default_bucket(), transformer._current_job_name)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this logic that's also in Transformer.transform? if so, I wonder if it'd be worth refactoring it into a private method that you can call on transformer here (like EstimatorBase._prepare_for_training)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep it's similar except the naming method is different. One is airflow_name_from_base and the other is name_from_base. So if I wrap them in a method, that method needs to take a method as input arg which I don't want to do for just these small amounts of codes for now. If we apply the different methods to get name first, then there's just one line left which I guess no need to wrap. I am still targeting at what I mentioned in other comments that we could reformat the codes in session/estimator/transformer/etc in a way that we could wrap a lot codes in methods and couple airflow with them.

return model_config(instance_type, model, role, image)


def transform_config(transformer, data, data_type='S3Prefix', content_type=None, compression_type=None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are these kinds of methods potentially reusable for Session? (this is sort of out of scope of this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep I do want it to be reused in session. Now I cannot because I need to bypass all the validations/describe calls/etc to use Jinja templating.

If in session, we can construct the config first, and then apply all manipulations afterwards. That would be awesome and these two parts can be coupled. (which is really good since for now I am worried if session part got updated, for example, new entries in boto config introduced, airflow part cannot catch it)

@iquinteroiquintero left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments.

transform_env = model.env.copy()
if env is not None:
transform_env.update(env)
if self.latest_training_job is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can just do

if self.latest_training_job:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep that looks nice. But the problem is we won't get into
if var
if var is things like [], or {}. Not just None. Hence I kind of want to explicitly say the only thing I don't want is None. Especially sometimes I do have some var = {} and need to do var.update() after. If there's an if in between, I do need to deal with the {} case.
Yep we probably should do
if var
most of times and
if var is not None
only when needed. But using the second one for now is kind of consistent to what we have in the existed classes (like estimator, model, etc). I prefer not to change for now. If later when we introduce style guide and can force it everywhere, we probably could do the change for all.

Comment threadsrc/sagemaker/estimator.py Outdated
logging.warning('No finished training job found associated with this estimator. Please make sure'
'this estimator is only used for building workflow config')
model_name = self._current_job_name
transform_env = env if env is not None else {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar here,

transform_env = env or {}

vpc_config = model.vpc_config
self.sagemaker_session.create_model(model_name, role, container_def, vpc_config)
transform_env = model.env.copy()
if env is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if env:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

transformer._current_job_name = job_name
else:
base_name = transformer.base_transform_job_name
transformer._current_job_name = utils.airflow_name_from_base(base_name) \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transformer._current_job_name=utils.airflow_name_from_base(base_name) ifbase_nameelsetransformer.model_name

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

'TransformResources': job_config['resource_config'],
}

if transformer.strategy is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar to my other comments, do all these as

iftransformer.strategy:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

production_variant = sagemaker.production_variant(model.name, instance_type, initial_instance_count)
name = model.name
config_options = {'EndpointConfigName': name, 'ProductionVariants': [production_variant]}
if tags is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iftags:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.


# if there is s3 operations needed for model, move it to root level of config
s3_operations = model_base_config.pop('S3Operations', None)
if s3_operations is not None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

^^

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Answered above.

def test_transformer_config(sagemaker_session):
tf_transformer = transformer.Transformer(
model_name="tensorflow-model",
instance_count="{{ instance_count }}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Im not sure I understand what this syntax is doing?

"{{ instance_count }}"

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is Jinja templates. It will be evaluated during Airflow runtime. For example, I can put something in database (using xcom in ariflow) and do "{{ task_instance.xcom_pull(task_id='task', key='key') }}" to get the record in the table 'task' with key 'key'. Here in unit tests I just make some random Jinja templating strings for testing purpose.

assert config == expected_config


def test_transform_config_from_amazon_alg_estimator(sagemaker_session):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you include in the test name that this also tests the code paths where you don't supply the optional args?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Talked offline. No need to change.

if env is not None:
transform_env.update(env)
else:
logging.warning('No finished training job found associated with this estimator. Please make sure'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there unit tests for this change?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep the transformer from estimator tests will cover this.

@yangaws
yangaws merged commit 53a43f6 into aws:masterNov 18, 2018
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Evan-W-ang added a commit to Evan-W-ang/sagemaker-python-sdk that referenced this pull request Jun 8, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@yangaws@codecov-io@iquintero@laurenyu