G6 Instances on SageMaker AI Inference Expand to GovCloud
AWS brings NVIDIA L4-powered inference to its GovCloud (US-East) region, giving government workloads access to GPU-accelerated endpoints with compliance guardrails.
Can government workloads on SageMaker AI now use G6 instances?
Yes. AWS has made G6 instances available on SageMaker AI inference in the AWS GovCloud (US-East) region as of July 23, 2026. The expansion lets agencies and organizations operating under strict compliance and data residency requirements run GPU-accelerated inference endpoints without leaving the GovCloud environment.
What G6 instances bring to the table
G6 instances are built around NVIDIA L4 Tensor Core GPUs and third-generation AMD EPYC processors. Each GPU carries 24 GB of memory, and configurations scale up to eight GPUs per instance.
AWS claims the instances deliver up to twice the deep learning inference performance of G4dn instances. That improvement targets production workloads that fit within a single GPU's 24 GB memory ceiling.
Workloads the instances are designed for
AWS lists small-to-medium language models, image generation, and computer vision tasks as the primary use cases. The 24 GB per-GPU ceiling means very large models requiring multi-GPU tensor parallelism across many devices will need a different instance family, but the per-GPU capacity is sufficient for a wide range of current generative AI inference tasks.
| Specification | G6 Instance |
|---|---|
| GPU | NVIDIA L4 Tensor Core |
| GPU memory per card | 24 GB |
| Max GPUs per instance | 8 |
| CPU | 3rd-gen AMD EPYC |
| Inference performance vs. G4dn | Up to 2x |
| SageMaker AI inference region (new) | AWS GovCloud (US-East) |
Why GovCloud availability matters
GovCloud regions operate under additional controls to help customers meet US government compliance frameworks, including restrictions on who can administer the environment. Agencies that previously had to run inference on less capable instance types, or forego SageMaker AI entirely for GPU-intensive workloads, now have a more capable option within that compliance boundary.
The announcement does not specify which other regions already support G6 instances on SageMaker AI inference, referring only to "previously supported regions."
What to check before deploying
Pricing is listed on the AWS SageMaker AI pricing page rather than in the announcement itself, so teams evaluating cost should consult that page directly. Organizations will also want to confirm instance quotas in GovCloud before building deployment pipelines, as new region availability sometimes launches with limited default quota allocations.