# AWS and NVIDIA plan 2 million more GPUs as their AI stack expands beyond accelerators

AWS and NVIDIA plan to deploy two million additional Blackwell Ultra, Rubin and Rubin Ultra GPUs in 2027–2028 while bringing Vera CPUs, NVLink Fusion, Nemotron models and tighter data-processing integrations into AWS.

The scale of the AWS–NVIDIA expansion is the headline, but the builder consequence is broader: AWS is co-engineering more of the NVIDIA stack, from CPUs and interconnects to models, vector indexing and physical-AI infrastructure, rather than merely adding another GPU instance family.

- Status: Active
- Published: 2026-08-28T11:05:38+12:00
- Updated: 2026-08-28T11:05:38+12:00
- Categories: Artificial Intelligence, Cloud & Infrastructure, Cloud Platforms, Compute & AI Infrastructure
- Tags: AI accelerators, AWS, inference infrastructure, NVIDIA, Vera
- Canonical HTML: https://beyondthe.news/dossiers/aws-nvidia-two-million-gpus-vera-nvlink-ai-infrastructure

## What changed

On August 26, 2026, AWS and NVIDIA announced that AWS plans to deploy two million additional NVIDIA GPUs across its global infrastructure in 2027 and 2028. The capacity will span Blackwell Ultra, Rubin and Rubin Ultra generations and comes on top of AWS's earlier plan to add more than one million NVIDIA GPUs beginning in 2026. The companies are also working to bring NVIDIA Vera CPU-based infrastructure to AWS, extend NVLink Fusion with NVIDIA high-bandwidth memory, integrate the NVIDIA platform with Nitro and Elastic Fabric Adapter, accelerate EMR and OpenSearch workloads with cuDF and cuVS, continue Nemotron availability through Bedrock and SageMaker, and build U.S. government AI factories including 100,000 GPUs on secure AWS infrastructure.

## Why it matters

For builders, this is a capacity and architecture signal rather than a benchmark story. AWS is committing to another very large block of NVIDIA capacity while making more of the surrounding NVIDIA stack native to its cloud. That can broaden future managed access to Rubin-era infrastructure, reduce friction for heterogeneous CPU/GPU workloads, and make NVIDIA-accelerated data processing and vector search more available inside ordinary AWS services. It also reinforces a strategic concentration risk: teams that choose AWS for managed AI can gain unusually deep NVIDIA integration while becoming more exposed to the roadmap and pricing decisions of both vendors.

## The capacity plan is materially larger than AWS's earlier 2026 commitment

AWS says demand exceeded the expectations behind its earlier plan for more than one million NVIDIA GPUs. The new agreement adds two million more GPUs in 2027–2028 across Blackwell Ultra, Rubin and Rubin Ultra systems. The companies have not published customer pricing or a region-by-region availability schedule, so the commitment establishes future capacity direction rather than immediately purchasable inventory.

## Vera brings CPU work for agents into the same partnership

AWS and NVIDIA are working to bring Vera CPU-based infrastructure to AWS. NVIDIA positions Vera for CPU-heavy parts of agentic AI and reinforcement-learning systems such as code execution, tool use, sandboxing, analytics and orchestration. AWS says it has already received its first Vera CPU server and Vera Rubin GPU, but the announcement does not give a general-availability date for customer instances.

## The partnership now reaches networking, data processing and open models

The expansion includes NVLink Fusion work, Nitro and EFA integration, Nemotron open models in Bedrock and SageMaker, and NVIDIA cuDF/cuVS acceleration for EMR and OpenSearch. The practical direction is a more integrated stack in which compute, networking, model access and data preparation are increasingly co-designed rather than exposed as isolated products.

## Physical AI and government workloads add two new demand pools

Amazon Robotics is adopting NVIDIA's physical-AI platform, and the companies say they will build AI factories for U.S. government workloads including a 100,000-GPU secure AWS deployment. Those projects do not directly change ordinary developer access today, but they add large competing demand pools for the same accelerator ecosystem and show where AWS expects future AI infrastructure growth.

## Key details

- AWS plans to deploy two million additional NVIDIA GPUs across its global infrastructure in 2027–2028.
- The additional fleet is planned to include Blackwell Ultra, Rubin and Rubin Ultra GPUs.
- AWS had previously announced plans for more than one million additional NVIDIA GPUs beginning in 2026.
- AWS and NVIDIA are working to bring Vera CPU-based infrastructure to AWS; AWS says it has received its first Vera CPU server and Vera Rubin GPU.
- The companies plan deeper integration across NVLink Fusion, AWS Nitro, EFA, Nemotron models, EMR, OpenSearch and NVIDIA CUDA-X libraries.
- The U.S. government AI-factory work includes a planned 100,000-GPU deployment on secure AWS infrastructure.
- Customer pricing, specific regions and general-availability dates for much of the 2027–2028 capacity have not been published.

## Builder takeaways

- Treat the announcement as a roadmap and capacity signal, not as immediately available Rubin capacity; wait for instance families, regions, quotas and pricing before modeling production economics.
- If your workload mixes GPU inference with CPU-heavy agent orchestration or sandbox execution, watch the eventual Vera instance design rather than assuming today's GPU-instance architecture will remain the right comparison.
- For data-heavy AI systems, test whether future cuDF/cuVS integrations in EMR and OpenSearch materially reduce preprocessing and vector-indexing cost before adding separate GPU data infrastructure.
- Keep cloud and accelerator abstraction realistic. Deeper AWS–NVIDIA integration can improve operational simplicity, but it also increases dependence on two coupled vendor roadmaps.
- Capacity-sensitive teams should track the gap between announced GPU counts and generally available customer capacity, including regional quotas and reserved-capacity terms.

## What to watch

- The first AWS instance families using NVIDIA Vera CPUs and Rubin-generation GPUs.
- Regional rollout, quotas and pricing for the announced 2027–2028 GPU capacity.
- Production details for NVLink Fusion with AWS infrastructure and whether it changes cluster topology or scaling economics.
- Benchmarks and pricing for cuDF/cuVS acceleration inside EMR and OpenSearch.
- Whether AWS's expanded NVIDIA capacity changes availability or pricing pressure for Trainium-based alternatives.

## Uncertainties

- The two-million-GPU figure is a forward deployment plan spanning 2027–2028 rather than current installed capacity.
- AWS and NVIDIA have not disclosed the financial value of the expanded agreement or customer pricing.
- Many of the co-engineered integrations are roadmap items without public general-availability dates.
- Vendor performance claims for new EC2 hardware should be validated on workload-specific configurations when products become available.

## Sources

- [AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI](https://press.aboutamazon.com/aws/2026/8/aws-and-nvidia-to-deliver-2-million-additional-gpus-and-next-generation-infrastructure-for-agentic-and-physical-ai) — AWS · primary · 2026-08-26T00:00:00+12:00. Primary joint announcement for the two-million-GPU plan, Vera CPU work, networking, government AI factories, Nemotron, data-processing integrations and robotics.
- [AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI](https://investor.nvidia.com/news/press-release-details/2026/AWS-and-NVIDIA-to-Deliver-2-Million-Additional-GPUs-and-Next-Generation-Infrastructure-for-Agentic-and-Physical-AI/default.aspx) — NVIDIA · primary · 2026-08-26T00:00:00+12:00. NVIDIA's primary version of the joint infrastructure announcement.
- [Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now](https://blogs.nvidia.com/blog/vera-cpu-delivery/) — NVIDIA · primary · 2026-08-27T00:00:00+12:00. Current primary evidence that AWS has received its first Vera CPU server and Vera Rubin GPU, and NVIDIA's stated target workloads for Vera.

