Key Takeaways

  • Nvidia's Vera Rubin architecture pairs the Rubin GPU with a Vera CPU, an inference accelerator, and dedicated storage and networking racks — parts built to keep everything around the GPU efficient.
  • Nvidia's VP of storage technology says the Vera CPU delivered "upwards of 3x improvement" on data operations, letting flash storage run without bottlenecking the GPU.
  • Amazon added 2 million more Nvidia GPUs for 2027–2028 delivery, five months after committing to 1 million — while its own custom-chip business passed a $25B annualized run rate.
  • Nvidia Q2: $96.2B revenue, $89B from data centers (+117% YoY), $108B guided for Q3.

For three years the bear case on Nvidia has been simple: the hyperscalers will build their own chips, and the GPU monopoly ends. That case is mostly right, and it hasn’t mattered. After Wednesday’s earnings call, a different story took shape — Nvidia’s advantage has quietly migrated off the GPU and into everything that surrounds it.

The evidence arrived in the same week from two directions. Nvidia posted $96.2 billion in second-quarter sales, with $89 billion of it from data centers, up 117% year over year, and guided to $108 billion for Q3. And Amazon — a company actively building chips to reduce its dependence on Nvidia — tripled its order.

The bottleneck moved

The old framing treated compute as the scarce thing. As deployments push toward gigawatt scale, TechCrunch’s reporting argues, the scarce thing is orchestration — getting the right data to the right place at the right moment across thousands of components.

That is not a GPU problem. It’s a memory, storage and networking problem, and it’s the layer Nvidia has been quietly building out while everyone watched the accelerators.

You can see it in what the company is actually selling. The Vera Rubin architecture pairs the Rubin GPU with a set of other units: the Vera CPU, the Groq 3 LPX inference accelerator, and dedicated racks for storage and networking. These are as specialized as the GPU itself, but instead of churning through tokens they exist to make sure nothing outside the GPU becomes the limiting factor. TechCrunch’s metaphor: if the GPU is the engine, these are the rest of the car.

What ships alongside the Rubin GPU in Nvidia's Vera Rubin architecture
The parts of Vera Rubin that never touch a token.

Jason Hardy, Nvidia’s VP of storage technology, framed the Vera CPU’s job as data orchestration, and the constraint behind it as physical: “Vera is important because there’s only so much memory that you can put in a single server or any sort of compute platform,” he told TechCrunch.

Memory capacity has scaled alongside compute — which is why memory suppliers have had an extremely good couple of years — but getting that data to the GPU at the right time is a separate discipline. As operators push tokens-per-watt lower, traffic direction starts to dominate the efficiency math.

Nvidia VP of storage technology Jason Hardy on the Vera CPU
Hardy on what the Vera CPU unlocked for flash storage.

OpenAI is solving the same problem the opposite way

The most useful evidence that this is a real bottleneck, and not Nvidia marketing, is that a competitor built an entirely different answer to it.

When OpenAI developed its Jalapeño chip, the stated design goal was to avoid data movement rather than accelerate it. “We designed Jalapeño to minimize data movement and communication delays,” the company wrote in a blog post earlier this month, per TechCrunch. “Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end.”

Two opposite architectures, one shared premise: efficiency now comes from smarter traffic control, not more processor cycles.

Nvidia and OpenAI approaches to the data movement problem
Accelerate the movement, or design so there's less of it.

This does not automatically hand Nvidia the win. As TechCrunch notes, the company will have to compete with rival chipmakers and hyperscalers at this layer exactly as it did on GPUs. What’s changed is the terms: building a competitive GPU no longer buys you a seat, because the question is whether you can make the whole system efficient.

Amazon is the test case

Amazon is simultaneously Nvidia’s biggest counterexample and its biggest customer, which makes this week’s deal unusually informative.

The two companies announced an expanded partnership adding 2 million Nvidia GPUs — Blackwell Ultra, Rubin and Rubin Ultra — to AWS data centers in 2027 and 2028. That came five months after Amazon agreed to deploy more than 1 million Nvidia GPUs starting this year. Nvidia’s statement: since then, “demand has exceeded those expectations.” Neither company disclosed terms, but TechCrunch estimates it at tens of billions of dollars based on unit costs.

Amazon's Nvidia GPU commitments five months apart
The commitment tripled in five months.

Meanwhile Amazon’s own silicon effort is not slowing down. Its custom chip business crossed a $25 billion annualized revenue run rate, driven by $225 billion in total commitments from AI labs including Anthropic and OpenAI. AWS is reportedly in talks to sell its Trainium accelerators — a direct alternative to Nvidia’s Blackwell-class parts — to other companies, and its Arm-based Graviton CPUs already challenge Intel and AMD in general-purpose server silicon.

Both things being true at once is the whole point. Amazon can build a credible accelerator and still buy 2 million Nvidia GPUs, because what it’s buying isn’t only the GPU.

The rest of the deal makes that explicit. Nvidia said its networking hardware — the gear that stitches thousands of GPUs into one system — plus its open models, CPUs, data processing software and robotics platform will be integrated across AWS. Nvidia CFO Colette Kress said Vera CPUs are going out too, “some integrated with Rubin, others standalone,” and that Nvidia expects Vera to reach “every major hyperscaler, neocloud, AI lab, and system OEM.” Jensen Huang has claimed a “brand-new $200 billion TAM” for the CPU line.

Amazon is also adopting Nvidia’s full physical AI stack for its warehouse robots — Omniverse for simulation and digital twins, Cosmos for world models, Isaac for robotics development, and Jetson for on-robot compute — and will serve Nvidia’s Nemotron open models on Bedrock and SageMaker.

The other end of the market

While the hyperscalers argue about racks, the opposite end of the AI compute market had its own week. Apple announced the M5 Ultra and M6, aimed at running models locally rather than in a data center.

The M5 Ultra fuses two dual-die M5 Max chips into Apple’s first quad-die design, with up to a 36-core CPU, up to an 80-core GPU, and 1.2 TB/s of unified memory bandwidth — 50% more than the M3 Ultra from about eighteen months earlier. The M6, built on a 2nm process, adds two CPU and GPU cores over the M5 and delivers nearly 30% more peak GPU compute for AI, which Apple frames in terms of faster prompt processing for on-device LLMs. Mac Minis start at $899 and ship after September 22.

It’s a different bet on the same constraint. Apple’s answer to data movement is to keep the workload on one machine entirely — closer to OpenAI’s logic than Nvidia’s, at a radically different scale.

What happens next

Three things to watch.

Whether “system-level” survives as a defensible moat. Networking and orchestration are harder to replicate than a die, but they’re not magic. Hyperscalers building their own interconnect is the obvious counter-move, and they have the volume to justify it.

Whether Rubin ships on time. Nvidia said production shipments of Rubin began this quarter, and part of the $108 billion Q3 guide depends on them. That’s the number that will actually settle the argument.

Whether the tokens-per-watt framing sticks. If efficiency becomes the metric buyers optimize — rather than raw FLOPs — the companies selling complete systems have a structural advantage over the ones selling parts. That shift is what the whole thesis rests on, and it’s the piece that’s least proven.

For anyone building on top of this stack rather than buying it, the practical read is narrower: capacity keeps growing, but the cost of moving data is what determines what your inference actually costs. That’s the same wall teams hit in production — as Twilio found when its LLM gateway P99 hit 60 seconds — and the same reason multi-GPU kernel work remains hard for LLMs themselves. Nvidia’s bet is that it can sell you out of that problem. This quarter, Amazon bought it.

Quick poll

Where does the AI infrastructure advantage actually live in two years?

Nvidia's Q2: $89B of $96.2B in revenue came from data centers, up 117% year over year.

FAQ

What is Nvidia’s Vera Rubin architecture? Nvidia’s current platform rollout, pairing the Rubin GPU with the Vera CPU, the Groq 3 LPX inference accelerator, and specialized racks for storage and networking — components designed to keep data flowing rather than to process tokens.

What does the Vera CPU actually do? Data orchestration. Nvidia’s VP of storage technology Jason Hardy said it delivered “upwards of 3x improvement” on data operations, which lets flash storage run at full speed without bottlenecking the GPU.

How many Nvidia GPUs is Amazon buying? An additional 2 million Blackwell Ultra, Rubin and Rubin Ultra GPUs for AWS data centers in 2027 and 2028, on top of the more than 1 million it committed to five months earlier.

Isn’t Amazon building chips to replace Nvidia? Yes, and both are happening at once. Amazon’s custom chip business passed a $25 billion annualized run rate and it is reportedly in talks to sell Trainium to other companies — while still tripling its Nvidia order.

What were Nvidia’s Q2 2026 results? $96.2 billion in revenue, beating analyst estimates, with $89 billion from data centers — up 117% year over year. The company guided to $108 billion for Q3.