- Introduces SageMaker HyperPod Inference Gateway, a Kubernetes‑native GPU‑aware routing add‑on for SageMaker HyperPod clusters.
- Replaces round‑robin load balancing with real‑time inference‑signal routing, reducing first‑token latency up to 82% and p99 TTFT by 97‑98% in mixed‑hardware and burst traffic scenarios.
- Provides Envoy endpoint, body‑based router, and endpoint picker that score pods on six inference signals, supporting any OpenAI‑compatible model server (e.g., vLLM, SGLang) with zero application code changes.