Announcing region expansion of... Note

Announcing region expansion of G7e instances on SageMaker AI inference

Amazon EC2 G7e instances are now available on Amazon SageMaker for AI inference in Seoul, London, and Tokyo. These instances offer up to eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, each with 96 GB of memory. They also feature 5th Generation Intel Xeon processors and high networking bandwidth. G7e instances provide up to 2.3 times better inference performance than the previous G6e generation. This expansion allows for deploying inference endpoints closer to users in Asia and Europe, reducing latency for generative AI. With up to 768 GB of total GPU memory, G7e instances can handle LLMs up to 70 billion parameters using FP8 precision without multi-node setups. These instances are ideal for LLM inference, image and video generation, spatial computing, and scientific computing. The new regions join previously supported locations for G7e instances on SageMaker. Pricing details are available on the Amazon SageMaker pricing page.