G7e Instances on SageMaker AI Expand to Seoul, London, Tokyo
AWS adds three regions for G7e inference, bringing NVIDIA Blackwell GPUs closer to users in Asia and Europe.
Can I run G7e instances on SageMaker AI inference in Asia or Europe?
Yes. As of July 23, 2026, AWS has extended G7e instance availability on Amazon SageMaker AI inference to Asia Pacific (Seoul), Asia Pacific (Tokyo), and Europe (London). The expansion adds to the regions where G7e was already supported, giving teams in those geographies a local option for GPU-backed inference endpoints.
What hardware do G7e instances include?
Each G7e instance can carry up to eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, with 96 GB of memory per GPU. The total GPU memory across a fully configured instance reaches 768 GB. Compute is handled by 5th Generation Intel Xeon processors, and networking tops out at 1,600 Gbps via Elastic Fabric Adapter.
| Specification | G7e |
|---|---|
| Max GPUs per instance | 8 |
| GPU memory per GPU | 96 GB |
| Total GPU memory (max) | 768 GB |
| CPU generation | 5th Gen Intel Xeon |
| Max EFA networking bandwidth | 1,600 Gbps |
| Inference performance vs. G6e | Up to 2.3x |
Why does the regional expansion matter for latency?
Deploying inference endpoints in the same region as end users reduces round-trip time. For real-time generative AI applications, that difference is material. Teams serving users in Japan, South Korea, or the United Kingdom can now route requests to a local endpoint rather than crossing to a more distant region.
What model sizes fit on a single G7e instance?
AWS states that the 768 GB total GPU memory is sufficient to serve models up to 70 billion parameters at FP8 precision on a single instance, without requiring multi-node configurations. That simplifies deployment operations: no need to coordinate across nodes for models in that size range.
What workloads are G7e suited for?
According to AWS, the primary use cases are large language model inference, image and video generation, spatial computing, and scientific computing. The common thread is high GPU memory capacity and bandwidth. If a workload fits within those categories and previously required either a different instance family or a multi-node setup, G7e is worth evaluating.
How does G7e compare to the previous generation?
AWS cites up to 2.3x inference performance relative to G6e instances. No additional benchmark methodology or test conditions are specified in the announcement, so teams should treat that figure as a starting point and run their own workload-specific comparisons.
Where can I find pricing?
AWS directs users to its SageMaker AI pricing page for cost details on G7e in the newly added regions. Pricing was not included in the announcement document.