Amazon SageMaker AI inference ... Note

Amazon SageMaker AI inference now supports G7 instances

Amazon SageMaker AI inference now offers enhanced performance with G7 instances, powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. These new instances deliver up to 4.6 times better AI inference performance than previous-generation G6 instances. This advancement addresses the growing need for high GPU throughput and memory capacity for deploying generative AI models in production. Previously, users often had to over-provision compute or quantize models due to memory limitations. G7 instances provide 32 GB of GPU memory per GPU and feature 5th Generation Tensor Cores. They also offer significantly improved networking capabilities with up to 700 Gbps of EFA-enabled networking, seven times faster than G6. Additionally, these instances include up to 7.6 TB of local NVMe SSD storage for rapid access to large models. These features make G7 instances ideal for serving models with 7B to 30B parameters, as well as image and video generation tasks. Deployment is streamlined through the SageMaker AI Inference console, API, or SDK, by selecting G7 instance types. G7 instances are currently available in US East (N. Virginia, Ohio) and US West (Oregon) regions.