Amazon Web Services has made its EC2 G7 instances generally available, bringing NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, configurations of up to eight GPUs and up to 700 Gbps of Elastic Fabric Adapter networking to production AI inference workloads.
Amazon Web Services has made Amazon EC2 G7 instances generally available, adding a Blackwell-based GPU option for customers running AI inference workloads in production.
According to AWS’s general-availability announcement, G7 instances are powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs and can be configured with up to eight GPUs. AWS is positioning the family for generative AI inference, including deployments where throughput and response latency are important considerations.
The launch expands AWS’s range of accelerated compute offerings rather than replacing its existing GPU instance families. For organizations selecting infrastructure, the fit will depend on the model being served, memory requirements, target response times, traffic patterns and software configuration.
AWS said EC2 G7 instances offer up to 700 Gbps of Elastic Fabric Adapter, or EFA, networking. EFA is AWS’s network interface technology for workloads that require fast communication among compute resources.
That capability can be relevant for demanding AI serving deployments, particularly where workloads must coordinate across multiple GPUs or systems. The practical benefit will vary according to model architecture, inference framework, batching strategy and the overall design of the application.
AWS also said its new G7 instances deliver up to 4.6 times the AI inference performance of the previous-generation G6 instances. The figure is AWS’s own performance comparison, and it should not be treated as a universal result: model precision, batch size, runtime software and deployment settings can materially affect measured performance.
NVIDIA described the G7 offering as infrastructure intended for production AI inference in a post on its collaboration with AWS. The company highlighted the use of Blackwell-generation RTX PRO 4500 server GPUs, which are designed to support a broad set of AI and accelerated-computing workloads.
NVIDIA’s post also pointed to GPU-accelerated vector indexing through NVIDIA cuVS in Amazon OpenSearch Serverless. Vector indexing is commonly used in retrieval systems that identify relevant documents or records before an AI model produces a response. Such retrieval components are frequently used in retrieval-augmented generation applications.
The OpenSearch reference is separate from the EC2 G7 launch, but it illustrates the wider software and infrastructure stack AWS and NVIDIA are presenting for production AI applications.
General availability means EC2 G7 instances are available for customer use through AWS’s standard commercial rollout. The new family gives teams an additional choice for inference deployments that need Blackwell-based RTX PRO GPUs, substantial GPU capacity and high-bandwidth networking.
AWS’s announcement centers on inference rather than training. Customers evaluating G7 against other AWS accelerated instances will still need to test their own models and serving configurations, since the most economical option may differ by workload. AWS and NVIDIA have framed G7 as an addition to the cloud provider’s GPU portfolio, aimed at helping customers run AI models in production at scale.
According to AWS’s general availability announcement, G7 instances are powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs and can be configured with up to eight GPUs.
AWS is positioning the family for generative AI inference, including deployments where throughput and response latency are important considerations.
The launch expands AWS’s range of accelerated compute offerings rather than replacing its existing GPU instance families.
Continue reading