| InvokeEndpoint | Inference | Supported | Common | Forwards a request body of up to 6 MiB to one configured model container and returns its response body and content type. |
| InvokeEndpointAsync | Inference | Out of scope | Full surface | Asynchronous inference is unsupported; requests return an error. |
| InvokeEndpointWithResponseStream / InvokeEndpointWithBidirectionalStream | Inference | Out of scope | Full surface | Streaming inference is outside this adapter’s subset. |
| ContentType, Accept, CustomAttributes and InferenceId | Inference headers | Supported | Common | These values reach the model container. CustomAttributes in its response are returned to the caller. |
| Explanation, inference-component, session and prefix-cache controls | Inference options | Out of scope | Full surface | These request controls are not implemented or forwarded to the model. Do not rely on them taking effect. |
| Batch transform, training jobs, tuning and notebooks | ML platform | Out of scope | Full surface | These are separate workloads, not synchronous endpoint inference; they remain on AWS. |
| TargetVariant / TargetModel / TargetContainerHostname | Model routing | Out of scope | Full surface | Explicit variant, multi-model and inference-pipeline container selection are refused. Each configured endpoint serves one model container. |
| Model, endpoint configuration and endpoint Terraform resources | Provisioning | Partial | Most usage | Deploys the model image as a Kubernetes Deployment and Service. Model artifacts, AWS instance sizing and weighted production variants need separate handling. |
| CreateModel / CreateEndpointConfig / CreateEndpoint / UpdateEndpoint / DeleteEndpoint | Runtime management | Out of scope | Full surface | Provision serving resources through the origin stack. Live SageMaker management API calls are not covered by this inference adapter. |