TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

Description

@dillon-odonovan

Describe the bug
When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

The TorchServe inference.py script was implemented in #4662

With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

To reproduce

import boto3
import io
import mlflow
from mlflow import MlflowClient
from mlflow.models import infer_signature
import numpy as np
import pandas as pd
from sklearn import datasets
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
from sklearn.model_selection import train_test_split
X, y = datasets.load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
params = {
"solver": "lbfgs",
"max_iter": 1000,
"multi_class": "auto",
"random_state": 8888
}
lr = LogisticRegression(**params)
lr.fit(X_train, y_train)
y_pred = lr.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
with mlflow.start_run() as run:
mlflow.log_params(params)
mlflow.log_metric('accuracy', accuracy)
mlflow.set_tag('Training Info', 'Basic LR model for iris data')
signature = infer_signature(X_train, lr.predict(X_train))
model_info = mlflow.sklearn.log_model(
sk_model=lr,
artifact_path='sklearn-model',
signature=signature,
input_example=X_train,
registered_model_name='tracking-quickstart'
)
model_uri = f'runs:/{run.info.run_id}/sklearn-model'
schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
model_builder = ModelBuilder(
mode=Mode.SAGEMAKER_ENDPOINT,
schema_builder=schema_builder,
role_arn=os.environ['ROLE_ARN'],
model_metadata={
"MLFLOW_MODEL_PATH": model_uri,
"MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
}
)
model = model_builder.build()
predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
predictor.predict(X_test) # works as expected
sagemaker_runtime_client = boto3.client('sagemaker-runtime')
# works as expected:
buffer = io.BytesIO()
np.save(buffer, X_test)
sagemaker_runtime_client.invoke_endpoint(
EndpointName=predictor.endpoint_name,
Body=buffer.getvalue(),
ContentType='application/x-npy'
)
predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
# does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
invoke_response = sagemaker_runtime_client.invoke_endpoint(
EndpointName=predictor.endpoint_name,
Body=json_body,
ContentType='application/json'
)
ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
<title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
<h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
<p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.

Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

import codecs
invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
EndpointName=predictor.endpoint_name,
Body=buffer.getvalue(),
ContentType='application/x-npy',
Accept='application/json'
)
reader = codecs.getreader('utf-8')
json_response = reader(invoke_response_json_resp['Body'])
json.load(json_response)
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte

Whereas the below works:

invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
EndpointName=predictor.endpoint_name,
Body=buffer.getvalue(),
ContentType='application/x-npy',
Accept='application/json'
)
np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON

In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

Expected behavior
I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

Screenshots or logs

2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.

System information
A description of your system. Please provide:

  • SageMaker Python SDK version: 2.231.0
  • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
  • Framework version: 1.2.1
  • Python version: 3.8.17
  • CPU or GPU: CPU
  • Custom Docker image (Y/N): N

Additional context
Relevant guides / documentation used to generate example:

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
     blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

    Description

    @dillon-odonovan

    Describe the bug
    When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

    The TorchServe inference.py script was implemented in #4662

    With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

    This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
    or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

    To reproduce

    import boto3
    import io
    import mlflow
    from mlflow import MlflowClient
    from mlflow.models import infer_signature
    import numpy as np
    import pandas as pd
    from sklearn import datasets
    from sklearn.linear_model import LogisticRegression
    from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
    from sklearn.model_selection import train_test_split
    X, y = datasets.load_iris(return_X_y=True)
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    params = {
    "solver": "lbfgs",
    "max_iter": 1000,
    "multi_class": "auto",
    "random_state": 8888
    }
    lr = LogisticRegression(**params)
    lr.fit(X_train, y_train)
    y_pred = lr.predict(X_test)
    accuracy = accuracy_score(y_test, y_pred)
    mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
    with mlflow.start_run() as run:
    mlflow.log_params(params)
    mlflow.log_metric('accuracy', accuracy)
    mlflow.set_tag('Training Info', 'Basic LR model for iris data')
    signature = infer_signature(X_train, lr.predict(X_train))
    model_info = mlflow.sklearn.log_model(
    sk_model=lr,
    artifact_path='sklearn-model',
    signature=signature,
    input_example=X_train,
    registered_model_name='tracking-quickstart'
    )
    model_uri = f'runs:/{run.info.run_id}/sklearn-model'
    schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
    model_builder = ModelBuilder(
    mode=Mode.SAGEMAKER_ENDPOINT,
    schema_builder=schema_builder,
    role_arn=os.environ['ROLE_ARN'],
    model_metadata={
    "MLFLOW_MODEL_PATH": model_uri,
    "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
    }
    )
    model = model_builder.build()
    predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
    predictor.predict(X_test) # works as expected
    sagemaker_runtime_client = boto3.client('sagemaker-runtime')
    # works as expected:
    buffer = io.BytesIO()
    np.save(buffer, X_test)
    sagemaker_runtime_client.invoke_endpoint(
    EndpointName=predictor.endpoint_name,
    Body=buffer.getvalue(),
    ContentType='application/x-npy'
    )
    predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
    # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
    invoke_response = sagemaker_runtime_client.invoke_endpoint(
    EndpointName=predictor.endpoint_name,
    Body=json_body,
    ContentType='application/json'
    )
    ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
    <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
    <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
    <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
    ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
    

    Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

    import codecs
    invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
    EndpointName=predictor.endpoint_name,
    Body=buffer.getvalue(),
    ContentType='application/x-npy',
    Accept='application/json'
    )
    reader = codecs.getreader('utf-8')
    json_response = reader(invoke_response_json_resp['Body'])
    json.load(json_response)
    UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
    

    Whereas the below works:

    invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
    EndpointName=predictor.endpoint_name,
    Body=buffer.getvalue(),
    ContentType='application/x-npy',
    Accept='application/json'
    )
    np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
    

    In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

    Expected behavior
    I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

    Screenshots or logs

    2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
    2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
    2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
    2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
    2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
    2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
    2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
    2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
    2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
    2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
    2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
    

    System information
    A description of your system. Please provide:

    • SageMaker Python SDK version: 2.231.0
    • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
    • Framework version: 1.2.1
    • Python version: 3.8.17
    • CPU or GPU: CPU
    • Custom Docker image (Y/N): N

    Additional context
    Relevant guides / documentation used to generate example:

    Metadata

    Metadata

    Assignees

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

      Description

      @dillon-odonovan

      Describe the bug
      When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

      The TorchServe inference.py script was implemented in #4662

      With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

      This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
      or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

      To reproduce

      import boto3
      import io
      import mlflow
      from mlflow import MlflowClient
      from mlflow.models import infer_signature
      import numpy as np
      import pandas as pd
      from sklearn import datasets
      from sklearn.linear_model import LogisticRegression
      from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
      from sklearn.model_selection import train_test_split
      X, y = datasets.load_iris(return_X_y=True)
      X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
      params = {
      "solver": "lbfgs",
      "max_iter": 1000,
      "multi_class": "auto",
      "random_state": 8888
      }
      lr = LogisticRegression(**params)
      lr.fit(X_train, y_train)
      y_pred = lr.predict(X_test)
      accuracy = accuracy_score(y_test, y_pred)
      mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
      with mlflow.start_run() as run:
      mlflow.log_params(params)
      mlflow.log_metric('accuracy', accuracy)
      mlflow.set_tag('Training Info', 'Basic LR model for iris data')
      signature = infer_signature(X_train, lr.predict(X_train))
      model_info = mlflow.sklearn.log_model(
      sk_model=lr,
      artifact_path='sklearn-model',
      signature=signature,
      input_example=X_train,
      registered_model_name='tracking-quickstart'
      )
      model_uri = f'runs:/{run.info.run_id}/sklearn-model'
      schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
      model_builder = ModelBuilder(
      mode=Mode.SAGEMAKER_ENDPOINT,
      schema_builder=schema_builder,
      role_arn=os.environ['ROLE_ARN'],
      model_metadata={
      "MLFLOW_MODEL_PATH": model_uri,
      "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
      }
      )
      model = model_builder.build()
      predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
      predictor.predict(X_test) # works as expected
      sagemaker_runtime_client = boto3.client('sagemaker-runtime')
      # works as expected:
      buffer = io.BytesIO()
      np.save(buffer, X_test)
      sagemaker_runtime_client.invoke_endpoint(
      EndpointName=predictor.endpoint_name,
      Body=buffer.getvalue(),
      ContentType='application/x-npy'
      )
      predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
      # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
      invoke_response = sagemaker_runtime_client.invoke_endpoint(
      EndpointName=predictor.endpoint_name,
      Body=json_body,
      ContentType='application/json'
      )
      ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
      <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
      <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
      <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
      ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
      

      Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

      import codecs
      invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
      EndpointName=predictor.endpoint_name,
      Body=buffer.getvalue(),
      ContentType='application/x-npy',
      Accept='application/json'
      )
      reader = codecs.getreader('utf-8')
      json_response = reader(invoke_response_json_resp['Body'])
      json.load(json_response)
      UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
      

      Whereas the below works:

      invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
      EndpointName=predictor.endpoint_name,
      Body=buffer.getvalue(),
      ContentType='application/x-npy',
      Accept='application/json'
      )
      np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
      

      In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

      Expected behavior
      I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

      Screenshots or logs

      2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
      2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
      2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
      2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
      2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
      2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
      2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
      2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
      2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
      2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
      2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
      

      System information
      A description of your system. Please provide:

      • SageMaker Python SDK version: 2.231.0
      • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
      • Framework version: 1.2.1
      • Python version: 3.8.17
      • CPU or GPU: CPU
      • Custom Docker image (Y/N): N

      Additional context
      Relevant guides / documentation used to generate example:

      Metadata

      Metadata

      Assignees

      Type

      No type

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

        Description

        @dillon-odonovan

        Describe the bug
        When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

        The TorchServe inference.py script was implemented in #4662

        With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

        This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
        or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

        To reproduce

        import boto3
        import io
        import mlflow
        from mlflow import MlflowClient
        from mlflow.models import infer_signature
        import numpy as np
        import pandas as pd
        from sklearn import datasets
        from sklearn.linear_model import LogisticRegression
        from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
        from sklearn.model_selection import train_test_split
        X, y = datasets.load_iris(return_X_y=True)
        X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
        params = {
        "solver": "lbfgs",
        "max_iter": 1000,
        "multi_class": "auto",
        "random_state": 8888
        }
        lr = LogisticRegression(**params)
        lr.fit(X_train, y_train)
        y_pred = lr.predict(X_test)
        accuracy = accuracy_score(y_test, y_pred)
        mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
        with mlflow.start_run() as run:
        mlflow.log_params(params)
        mlflow.log_metric('accuracy', accuracy)
        mlflow.set_tag('Training Info', 'Basic LR model for iris data')
        signature = infer_signature(X_train, lr.predict(X_train))
        model_info = mlflow.sklearn.log_model(
        sk_model=lr,
        artifact_path='sklearn-model',
        signature=signature,
        input_example=X_train,
        registered_model_name='tracking-quickstart'
        )
        model_uri = f'runs:/{run.info.run_id}/sklearn-model'
        schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
        model_builder = ModelBuilder(
        mode=Mode.SAGEMAKER_ENDPOINT,
        schema_builder=schema_builder,
        role_arn=os.environ['ROLE_ARN'],
        model_metadata={
        "MLFLOW_MODEL_PATH": model_uri,
        "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
        }
        )
        model = model_builder.build()
        predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
        predictor.predict(X_test) # works as expected
        sagemaker_runtime_client = boto3.client('sagemaker-runtime')
        # works as expected:
        buffer = io.BytesIO()
        np.save(buffer, X_test)
        sagemaker_runtime_client.invoke_endpoint(
        EndpointName=predictor.endpoint_name,
        Body=buffer.getvalue(),
        ContentType='application/x-npy'
        )
        predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
        # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
        invoke_response = sagemaker_runtime_client.invoke_endpoint(
        EndpointName=predictor.endpoint_name,
        Body=json_body,
        ContentType='application/json'
        )
        ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
        <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
        <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
        <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
        ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
        

        Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

        import codecs
        invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
        EndpointName=predictor.endpoint_name,
        Body=buffer.getvalue(),
        ContentType='application/x-npy',
        Accept='application/json'
        )
        reader = codecs.getreader('utf-8')
        json_response = reader(invoke_response_json_resp['Body'])
        json.load(json_response)
        UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
        

        Whereas the below works:

        invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
        EndpointName=predictor.endpoint_name,
        Body=buffer.getvalue(),
        ContentType='application/x-npy',
        Accept='application/json'
        )
        np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
        

        In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

        Expected behavior
        I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

        Screenshots or logs

        2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
        2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
        2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
        2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
        2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
        2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
        2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
        2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
        2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
        2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
        2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
        

        System information
        A description of your system. Please provide:

        • SageMaker Python SDK version: 2.231.0
        • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
        • Framework version: 1.2.1
        • Python version: 3.8.17
        • CPU or GPU: CPU
        • Custom Docker image (Y/N): N

        Additional context
        Relevant guides / documentation used to generate example:

        Metadata

        Metadata

        Assignees

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

          Description

          @dillon-odonovan

          Describe the bug
          When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

          The TorchServe inference.py script was implemented in #4662

          With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

          This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
          or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

          To reproduce

          import boto3
          import io
          import mlflow
          from mlflow import MlflowClient
          from mlflow.models import infer_signature
          import numpy as np
          import pandas as pd
          from sklearn import datasets
          from sklearn.linear_model import LogisticRegression
          from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
          from sklearn.model_selection import train_test_split
          X, y = datasets.load_iris(return_X_y=True)
          X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
          params = {
          "solver": "lbfgs",
          "max_iter": 1000,
          "multi_class": "auto",
          "random_state": 8888
          }
          lr = LogisticRegression(**params)
          lr.fit(X_train, y_train)
          y_pred = lr.predict(X_test)
          accuracy = accuracy_score(y_test, y_pred)
          mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
          with mlflow.start_run() as run:
          mlflow.log_params(params)
          mlflow.log_metric('accuracy', accuracy)
          mlflow.set_tag('Training Info', 'Basic LR model for iris data')
          signature = infer_signature(X_train, lr.predict(X_train))
          model_info = mlflow.sklearn.log_model(
          sk_model=lr,
          artifact_path='sklearn-model',
          signature=signature,
          input_example=X_train,
          registered_model_name='tracking-quickstart'
          )
          model_uri = f'runs:/{run.info.run_id}/sklearn-model'
          schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
          model_builder = ModelBuilder(
          mode=Mode.SAGEMAKER_ENDPOINT,
          schema_builder=schema_builder,
          role_arn=os.environ['ROLE_ARN'],
          model_metadata={
          "MLFLOW_MODEL_PATH": model_uri,
          "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
          }
          )
          model = model_builder.build()
          predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
          predictor.predict(X_test) # works as expected
          sagemaker_runtime_client = boto3.client('sagemaker-runtime')
          # works as expected:
          buffer = io.BytesIO()
          np.save(buffer, X_test)
          sagemaker_runtime_client.invoke_endpoint(
          EndpointName=predictor.endpoint_name,
          Body=buffer.getvalue(),
          ContentType='application/x-npy'
          )
          predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
          # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
          invoke_response = sagemaker_runtime_client.invoke_endpoint(
          EndpointName=predictor.endpoint_name,
          Body=json_body,
          ContentType='application/json'
          )
          ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
          <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
          <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
          <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
          ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
          

          Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

          import codecs
          invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
          EndpointName=predictor.endpoint_name,
          Body=buffer.getvalue(),
          ContentType='application/x-npy',
          Accept='application/json'
          )
          reader = codecs.getreader('utf-8')
          json_response = reader(invoke_response_json_resp['Body'])
          json.load(json_response)
          UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
          

          Whereas the below works:

          invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
          EndpointName=predictor.endpoint_name,
          Body=buffer.getvalue(),
          ContentType='application/x-npy',
          Accept='application/json'
          )
          np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
          

          In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

          Expected behavior
          I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

          Screenshots or logs

          2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
          2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
          2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
          2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
          2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
          2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
          2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
          2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
          2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
          2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
          2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
          

          System information
          A description of your system. Please provide:

          • SageMaker Python SDK version: 2.231.0
          • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
          • Framework version: 1.2.1
          • Python version: 3.8.17
          • CPU or GPU: CPU
          • Custom Docker image (Y/N): N

          Additional context
          Relevant guides / documentation used to generate example:

          Metadata

          Metadata

          Assignees

          Type

          No type

          Projects

          No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

            Description

            @dillon-odonovan

            Describe the bug
            When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

            The TorchServe inference.py script was implemented in #4662

            With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

            This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
            or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

            To reproduce

            import boto3
            import io
            import mlflow
            from mlflow import MlflowClient
            from mlflow.models import infer_signature
            import numpy as np
            import pandas as pd
            from sklearn import datasets
            from sklearn.linear_model import LogisticRegression
            from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
            from sklearn.model_selection import train_test_split
            X, y = datasets.load_iris(return_X_y=True)
            X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
            params = {
            "solver": "lbfgs",
            "max_iter": 1000,
            "multi_class": "auto",
            "random_state": 8888
            }
            lr = LogisticRegression(**params)
            lr.fit(X_train, y_train)
            y_pred = lr.predict(X_test)
            accuracy = accuracy_score(y_test, y_pred)
            mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
            with mlflow.start_run() as run:
            mlflow.log_params(params)
            mlflow.log_metric('accuracy', accuracy)
            mlflow.set_tag('Training Info', 'Basic LR model for iris data')
            signature = infer_signature(X_train, lr.predict(X_train))
            model_info = mlflow.sklearn.log_model(
            sk_model=lr,
            artifact_path='sklearn-model',
            signature=signature,
            input_example=X_train,
            registered_model_name='tracking-quickstart'
            )
            model_uri = f'runs:/{run.info.run_id}/sklearn-model'
            schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
            model_builder = ModelBuilder(
            mode=Mode.SAGEMAKER_ENDPOINT,
            schema_builder=schema_builder,
            role_arn=os.environ['ROLE_ARN'],
            model_metadata={
            "MLFLOW_MODEL_PATH": model_uri,
            "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
            }
            )
            model = model_builder.build()
            predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
            predictor.predict(X_test) # works as expected
            sagemaker_runtime_client = boto3.client('sagemaker-runtime')
            # works as expected:
            buffer = io.BytesIO()
            np.save(buffer, X_test)
            sagemaker_runtime_client.invoke_endpoint(
            EndpointName=predictor.endpoint_name,
            Body=buffer.getvalue(),
            ContentType='application/x-npy'
            )
            predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
            # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
            invoke_response = sagemaker_runtime_client.invoke_endpoint(
            EndpointName=predictor.endpoint_name,
            Body=json_body,
            ContentType='application/json'
            )
            ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
            <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
            <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
            <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
            ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
            

            Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

            import codecs
            invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
            EndpointName=predictor.endpoint_name,
            Body=buffer.getvalue(),
            ContentType='application/x-npy',
            Accept='application/json'
            )
            reader = codecs.getreader('utf-8')
            json_response = reader(invoke_response_json_resp['Body'])
            json.load(json_response)
            UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
            

            Whereas the below works:

            invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
            EndpointName=predictor.endpoint_name,
            Body=buffer.getvalue(),
            ContentType='application/x-npy',
            Accept='application/json'
            )
            np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
            

            In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

            Expected behavior
            I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

            Screenshots or logs

            2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
            2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
            2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
            2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
            2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
            2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
            2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
            2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
            2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
            2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
            2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
            

            System information
            A description of your system. Please provide:

            • SageMaker Python SDK version: 2.231.0
            • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
            • Framework version: 1.2.1
            • Python version: 3.8.17
            • CPU or GPU: CPU
            • Custom Docker image (Y/N): N

            Additional context
            Relevant guides / documentation used to generate example:

            Metadata

            Metadata

            Assignees

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

              Description

              @dillon-odonovan

              Describe the bug
              When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

              The TorchServe inference.py script was implemented in #4662

              With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

              This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
              or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

              To reproduce

              import boto3
              import io
              import mlflow
              from mlflow import MlflowClient
              from mlflow.models import infer_signature
              import numpy as np
              import pandas as pd
              from sklearn import datasets
              from sklearn.linear_model import LogisticRegression
              from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
              from sklearn.model_selection import train_test_split
              X, y = datasets.load_iris(return_X_y=True)
              X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
              params = {
              "solver": "lbfgs",
              "max_iter": 1000,
              "multi_class": "auto",
              "random_state": 8888
              }
              lr = LogisticRegression(**params)
              lr.fit(X_train, y_train)
              y_pred = lr.predict(X_test)
              accuracy = accuracy_score(y_test, y_pred)
              mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
              with mlflow.start_run() as run:
              mlflow.log_params(params)
              mlflow.log_metric('accuracy', accuracy)
              mlflow.set_tag('Training Info', 'Basic LR model for iris data')
              signature = infer_signature(X_train, lr.predict(X_train))
              model_info = mlflow.sklearn.log_model(
              sk_model=lr,
              artifact_path='sklearn-model',
              signature=signature,
              input_example=X_train,
              registered_model_name='tracking-quickstart'
              )
              model_uri = f'runs:/{run.info.run_id}/sklearn-model'
              schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
              model_builder = ModelBuilder(
              mode=Mode.SAGEMAKER_ENDPOINT,
              schema_builder=schema_builder,
              role_arn=os.environ['ROLE_ARN'],
              model_metadata={
              "MLFLOW_MODEL_PATH": model_uri,
              "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
              }
              )
              model = model_builder.build()
              predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
              predictor.predict(X_test) # works as expected
              sagemaker_runtime_client = boto3.client('sagemaker-runtime')
              # works as expected:
              buffer = io.BytesIO()
              np.save(buffer, X_test)
              sagemaker_runtime_client.invoke_endpoint(
              EndpointName=predictor.endpoint_name,
              Body=buffer.getvalue(),
              ContentType='application/x-npy'
              )
              predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
              # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
              invoke_response = sagemaker_runtime_client.invoke_endpoint(
              EndpointName=predictor.endpoint_name,
              Body=json_body,
              ContentType='application/json'
              )
              ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
              <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
              <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
              <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
              ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
              

              Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

              import codecs
              invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
              EndpointName=predictor.endpoint_name,
              Body=buffer.getvalue(),
              ContentType='application/x-npy',
              Accept='application/json'
              )
              reader = codecs.getreader('utf-8')
              json_response = reader(invoke_response_json_resp['Body'])
              json.load(json_response)
              UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
              

              Whereas the below works:

              invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
              EndpointName=predictor.endpoint_name,
              Body=buffer.getvalue(),
              ContentType='application/x-npy',
              Accept='application/json'
              )
              np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
              

              In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

              Expected behavior
              I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

              Screenshots or logs

              2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
              2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
              2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
              2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
              2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
              2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
              2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
              2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
              2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
              2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
              2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
              

              System information
              A description of your system. Please provide:

              • SageMaker Python SDK version: 2.231.0
              • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
              • Framework version: 1.2.1
              • Python version: 3.8.17
              • CPU or GPU: CPU
              • Custom Docker image (Y/N): N

              Additional context
              Relevant guides / documentation used to generate example:

              Metadata

              Metadata

              Assignees

              Type

              No type

              Projects

              No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                TorchServe inference.py incompatible with default PyTorch inference Transformer for UTF-8 Content Types #4869

                Description

                @dillon-odonovan

                Describe the bug
                When running a SageMaker container using the default PyTorch inference Transformer, when specifying a UTF-8 Content-Type (application/json, text/csv), the TorchServe inference.py implementation will throw an error during de-serialization within input_fn. This is because the TorchServe inference input_fn function expects the input_data to be a bytes-like object, but it has already been decoded to a str by the Transformer. The NumpyDeserializer does support de-serializing from UTF-8 Content Types, but the code is effectively unreachable for input processing (can still be reached for output) without overriding the default Inference Handler / Handler Service / Transformer (transformer can't be specified if input_fn is specified).

                The TorchServe inference.py script was implemented in #4662

                With Python clients, or using the Predictor class from the SageMaker SDK, this is easily worked around. However, if trying to make predictions from other languages, such as Java, this is much more difficult as a JSON representation of the inference input cannot be provided, and custom serialization to match the NPY format would be necessary.

                This is just one example use case - the issue may be applicable for different input beyond Numpy arrays / scikit-learn algorithms. Ownership of fix could lie either in the SageMaker python SDK, within the sagemaker-pytorch-inference repository,
                or elsewhere. A change to any of these components could run the risk of impacting production behavior which clients may be reliant on.

                To reproduce

                import boto3
                import io
                import mlflow
                from mlflow import MlflowClient
                from mlflow.models import infer_signature
                import numpy as np
                import pandas as pd
                from sklearn import datasets
                from sklearn.linear_model import LogisticRegression
                from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
                from sklearn.model_selection import train_test_split
                X, y = datasets.load_iris(return_X_y=True)
                X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
                params = {
                "solver": "lbfgs",
                "max_iter": 1000,
                "multi_class": "auto",
                "random_state": 8888
                }
                lr = LogisticRegression(**params)
                lr.fit(X_train, y_train)
                y_pred = lr.predict(X_test)
                accuracy = accuracy_score(y_test, y_pred)
                mlflow.set_tracking_uri(os.environ['MLFLOW_URI'])
                with mlflow.start_run() as run:
                mlflow.log_params(params)
                mlflow.log_metric('accuracy', accuracy)
                mlflow.set_tag('Training Info', 'Basic LR model for iris data')
                signature = infer_signature(X_train, lr.predict(X_train))
                model_info = mlflow.sklearn.log_model(
                sk_model=lr,
                artifact_path='sklearn-model',
                signature=signature,
                input_example=X_train,
                registered_model_name='tracking-quickstart'
                )
                model_uri = f'runs:/{run.info.run_id}/sklearn-model'
                schema_builder = SchemaBuilder(sample_input=X_train, sample_output=y_pred)
                model_builder = ModelBuilder(
                mode=Mode.SAGEMAKER_ENDPOINT,
                schema_builder=schema_builder,
                role_arn=os.environ['ROLE_ARN'],
                model_metadata={
                "MLFLOW_MODEL_PATH": model_uri,
                "MLFLOW_TRACKING_ARN": os.environ['MLFLOW_TRACKING_SERVER_ARN']
                }
                )
                model = model_builder.build()
                predictor = model.deploy(initial_instance_count=1, instance_type="ml.t2.medium")
                predictor.predict(X_test) # works as expected
                sagemaker_runtime_client = boto3.client('sagemaker-runtime')
                # works as expected:
                buffer = io.BytesIO()
                np.save(buffer, X_test)
                sagemaker_runtime_client.invoke_endpoint(
                EndpointName=predictor.endpoint_name,
                Body=buffer.getvalue(),
                ContentType='application/x-npy'
                )
                predictions = np.load(io.BytesIO(invoke_response['Body'].read()))
                # does not work as expected; json_body = json.dumps(X_test.tolist()).encode('utf-8')
                invoke_response = sagemaker_runtime_client.invoke_endpoint(
                EndpointName=predictor.endpoint_name,
                Body=json_body,
                ContentType='application/json'
                )
                ModelError: An error occurred (ModelError) when calling the InvokeEndpoint operation: Received server error (500) from primary with message "<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 3.2 Final//EN">
                <title>500 Internal Server Error<[/title](https://###REDACTED###.studio.us-west-2.sagemaker.aws/title)>
                <h1>Internal Server Error<[/h1](https://###REDACTED###.studio.us-west-2.sagemaker.aws/h1)>
                <p>The server encountered an internal error and was unable to complete your request. Either the server is overloaded or there is an error in the application.<[/p](https://###REDACTED###.studio.us-west-2.sagemaker.aws/p)>
                ". See https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logEventViewer:group=###REDACTED### in account ###REDACTED### for more information.
                

                Along similar lines (can open separate issue if applicable) - it doesn't seem as though requesting the response as JSON via the Accept header works. This is perhaps expected, though is only evident upon attempting to de-serialize the returned stream:

                import codecs
                invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
                EndpointName=predictor.endpoint_name,
                Body=buffer.getvalue(),
                ContentType='application/x-npy',
                Accept='application/json'
                )
                reader = codecs.getreader('utf-8')
                json_response = reader(invoke_response_json_resp['Body'])
                json.load(json_response)
                UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0: invalid start byte
                

                Whereas the below works:

                invoke_response_json_resp = sagemaker_runtime_client.invoke_endpoint(
                EndpointName=predictor.endpoint_name,
                Body=buffer.getvalue(),
                ContentType='application/x-npy',
                Accept='application/json'
                )
                np.load(io.BytesIO(invoke_response_json_resp['Body'].read())) # evidently the response stream is not JSON
                

                In general, the error messaging during serialization/de-serialization is unhelpful/misleading, as it suggests the (de-)serialization failed for pickled data, which is not always the case.

                Expected behavior
                I expect to be able to invoke the SageMaker endpoints with a JSON-serialized Numpy array and receive NPY response.

                Screenshots or logs

                2024-09-11T23:52:57.493Z IP - - [11/Sep/2024:23:52:55 +0000] "POST /invocations HTTP/1.1" 200 368 "-" "AHC/2.0"
                2024-09-11T23:55:30.402Z 2024-09-11 23:55:30,305 ERROR - inference - Exception on /invocations [POST]
                2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/opt/ml/code/inference.py", line 74, in input_fn io.BytesIO(input_data), content_type[0]
                2024-09-11T23:55:30.402Z TypeError: a bytes-like object is required, not 'str'
                2024-09-11T23:55:30.402Z The above exception was the direct cause of the following exception:
                2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
                2024-09-11T23:55:30.402Z Exception: Encountered error in deserialize_request.
                2024-09-11T23:55:30.402Z IP - - [11/Sep/2024:23:55:30 +0000] "POST /invocations HTTP/1.1" 500 290 "-" "AHC/2.0"
                2024-09-11T23:55:30.402Z During handling of the above exception, another exception occurred:
                2024-09-11T23:55:30.402Z Traceback (most recent call last): File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 2446, in wsgi_app response = self.full_dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1951, in full_dispatch_request rv = self.handle_user_exception(e) File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1820, in handle_user_exception reraise(exc_type, exc_value, tb) File "/miniconda3/lib/python3.8/site-packages/flask/_compat.py", line 39, in reraise raise value File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1949, in full_dispatch_request rv = self.dispatch_request() File "/miniconda3/lib/python3.8/site-packages/flask/app.py", line 1935, in dispatch_request return self.view_functions[rule.endpoint](**req.view_args) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 199, in transform result = self._transform_fn( File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_transformer.py", line 227, in _default_transform_fn data = self._input_fn(content, content_type) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 95, in wrapper six.reraise(error_class, error_class(e), sys.exc_info()[2]) File "/miniconda3/lib/python3.8/site-packages/six.py", line 702, in reraise raise value.with_traceback(tb) File "/miniconda3/lib/python3.8/site-packages/sagemaker_containers/_functions.py", line 93, in wrapper return fn(*args, **kwargs) File "/opt/ml/code/inference.py", line 77, in input_fn raise Exception("Encountered error in deserialize_request.") from e
                2024-09-11T23:55:32.657Z sagemaker_containers._errors.ClientError: Encountered error in deserialize_request.
                

                System information
                A description of your system. Please provide:

                • SageMaker Python SDK version: 2.231.0
                • Framework name (eg. PyTorch) or algorithm (eg. KMeans): PyTorch / SKLearn LogisticRegression
                • Framework version: 1.2.1
                • Python version: 3.8.17
                • CPU or GPU: CPU
                • Custom Docker image (Y/N): N

                Additional context
                Relevant guides / documentation used to generate example:

                Metadata

                Metadata

                Assignees

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions