Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add Public facing APIs to display benchmark metrics and list deployme… - #4589

Closed
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master
Closed

Add Public facing APIs to display benchmark metrics and list deployme…#4589
makungaj1 wants to merge 26 commits into
aws:masterfrom
makungaj1:master

Conversation

@makungaj1

@makungaj1makungaj1 commented Apr 17, 2024

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:
Introduce ReadOnly APIs, list deployment configurations and display benchmark metrics. These APIs are used to quickly select and try different configuration that best serve the use case.

Sample output for list_deployment_configs api

[{'ConfigName': 'neuron-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'neuron-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference-budget',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}},
{'ConfigName': 'gpu-inference',
'BenchmarkMetrics': [{'name': 'Latency', 'value': '100', 'unit': 'Tokens/S'},
{'name': 'Throughput', 'value': '100', 'unit': 'Tokens/Second'}],
'DeploymentConfig': {'ModelDataDownloadTimeout': None,
'ContainerStartupHealthCheckTimeout': None,
'ImageUri': '763104351884.dkr.ecr.us-west-2.amazonaws.com/pytorch-inference:1.10.0-gpu-py38',
'ModelData': {'S3DataSource': {'S3Uri': 's3://jumpstart-cache-prod-us-west-2/pytorch-ic/pytorch-ic-mobilenet-v2/artifacts/inference-prepack/v1.0.0/',
'S3DataType': 'S3Prefix',
'CompressionType': 'None'}},
'InstanceType': 'ml.g5.2xlarge',
'Environment': {'SAGEMAKER_PROGRAM': 'inference.py',
'ENDPOINT_SERVER_TIMEOUT': '3600',
'MODEL_CACHE_ROOT': '/opt/ml/model',
'SAGEMAKER_ENV': '1',
'SAGEMAKER_MODEL_SERVER_WORKERS': '1'}}}]

Sample output for display_benchmark_metrics api

| | Config Name | Instance Type | Latency (Tokens/S) | Throughput (Tokens/Second) |
|---:|:------------------------|:----------------|---------------------:|-----------------------------:|
| 0 | neuron-inference | ml.g5.2xlarge | 100 | 100 |
| 1 | neuron-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 2 | gpu-inference-budget | ml.g5.2xlarge | 100 | 100 |
| 3 | gpu-inference | ml.g5.2xlarge | 100 | 100 |

Testing done:

Merge Checklist

Put an x in the boxes that apply. You can also fill these out after creating the PR. If you're unsure about any of them, don't hesitate to ask. We're here to help! This is simply a reminder of what we are going to look for before merging your pull request.

General

  • I have read the CONTRIBUTING doc
  • I certify that the changes I am introducing will be backward compatible, and I have discussed concerns about this, if any, with the Python SDK team
  • I used the commit message format described in CONTRIBUTING
  • I have passed the region in to all S3 and STS clients that I've initialized as part of this change.
  • I have updated any necessary documentation, including READMEs and API docs (if appropriate)

Tests

  • I have added tests that prove my fix is effective or that my feature works (if appropriate)
  • I have added unit and/or integration tests as appropriate to ensure backward compatibility of the changes
  • I have checked that my tests are not configured for a specific region or account (if appropriate)
  • I have used unique_name_from_base to create resource names in integ tests (if appropriate)

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The instance rate is not available as part of the benchmark metrics data object, we need to use AWS pricing API to query the rate using current region info.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated

def display_benchmark_metrics(self):
"""Display Benchmark Metrics for deployment configs."""
config_names = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should consider making this a data class instead of stand along variables

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsetup.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment on lines +823 to +825
df["Instance Rate<br>($/Hour)"] = df["Instance Rate<br>($/Hour)"].apply(
lambda rate: "${0:.2f}".format(float(rate))
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can this be a fast follow to this PR?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
instance_rates = []
for deployment_config in self.list_deployment_configs():
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a chance that a deployment config does not have a benchark metric? In which case do we need validation here?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
config_names.append(deployment_config.get("ConfigName"))
instance_types.append(deployment_config.get("BenchmarkMetrics").key()[0])

for benchmark_metric in deployment_config.get("BenchmarkMetrics").values():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If any of the fields are potentially missing in the deployment configuration, there is a chance we lose order in lists

DC = { {l=A, t=A1, c=A2}, {l=B t=B1, c=}, {l=C, t=C1, c=C2}}

[A, B, C] -> latency
[A1, B1, C1] -> throughput
[A2, C2] -> cost

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment on lines +2185 to +2187
self.container_startup_health_check_timeout = (
deploy_kwargs.container_startup_health_check_timeout
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not just assign the value?

Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Comment threadsrc/sagemaker/jumpstart/types.py
@samruds

Copy link
Copy Markdown
Collaborator

Current structure looks good. Will tests be added in separate PR?

Comment threadsrc/sagemaker/jumpstart/model.py
Comment threadsrc/sagemaker/jumpstart/model.py Outdated
@makungaj1

Copy link
Copy Markdown
ContributorAuthor

Current structure looks good. Will tests be added in separate PR?

Will add tests in the current PR after #4583 is merged as some low levels APIs, bug fixes are not in the main branch yet.

@samruds

samruds commented Apr 19, 2024

Copy link
Copy Markdown
Collaborator

Two follow ups, tentative approval for logic. Not ok to merge to master branch.

  1. Add UT
  2. Integrating with Pricing API for cost.

Unit Tests
Comment threadsrc/sagemaker/jumpstart/utils.py Outdated
df.style.set_caption("Benchmark Metrics").set_table_styles([headers]).set_properties(
**{"text-align": "left"}
)
return df

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if the function has return value, do you need to update the method signature?

Comment threadsrc/sagemaker/jumpstart/model.py Outdated
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@makungaj1@samruds@qiyunz@gwang111@jiapinw