Nvidia Boosts AI Efficiency
Nvidia is positioning its upcoming Vera Rubin platform as a decisive leap in AI infrastructure efficiency, claiming a 10-fold improvement in tokens per watt over the prior Grace Blackwell generation when running DeepSeek’s R1 model at CoreWeave. The announcement arrives days before AMD’s Advancing AI conference, where the rival is expected to unveil its own Helios rack-scale system built around 72 Instinct MI455X GPUs.
This timing underscores a broader shift in the AI hardware race: performance per watt and per dollar now matter as much as raw throughput, especially as power-constrained data centers confront the demands of always-on agentic workloads. Nvidia’s move to detail both its new Rubin GPU and Vera CPU simultaneously signals an intent to lock in the full system stack—from silicon to networking—before competitors can close the gap.
Nvidia Positions Vera Rubin Against AMD’s Rack-Scale Ambitions
Nvidia’s performance claims center on real-world inference economics rather than peak theoretical metrics. CoreWeave reported 10 times more tokens per watt on DeepSeek R1 with the Vera Rubin NVL72 configuration compared with Grace Blackwell NVL72. The company also stated that its standalone Vera CPU delivers up to 1.9 times better agentic AI performance than AMD’s Epyc Turin processor in internal benchmarks. These figures arrive as AMD prepares to showcase its Helios platform, which pairs 72 Instinct MI455X GPUs with Epyc Venice CPUs and the ROCm software stack.
The competitive stakes extend beyond a single conference. Nvidia generated $215.9 billion in fiscal 2026 data-center revenue, dwarfing AMD’s $34.6 billion in fiscal 2025. Yet hyperscalers including Google, Amazon, Meta, and Microsoft continue to develop custom accelerators, while OpenAI and Broadcom recently announced their own AI processors. Nvidia’s strategy of releasing new platforms on a predictable cadence aims to keep rivals perpetually behind on both hardware and the CUDA software ecosystem that binds them together. Nvidia touts Vera Rubin performance ahead of rival AMD’s Advancing AI event
Domestic Manufacturing Scales to Meet Rubin Demand
Wistron’s new 324,000-square-foot facility in Fort Worth, Texas, represents Nvidia’s largest single-site commitment to U.S. AI system production to date. The plant is already running two manufacturing cells—one for the GB300 Grace Blackwell Ultra Superchip and another for the Vera Rubin Superchip—with capacity scaling to tens of thousands of boards per month. The $700 million investment has created more than 500 jobs, with plans to reach 1,000 by year-end.
Nvidia CEO Jensen Huang framed the opening as part of a broader reindustrialization effort, noting that building chip, packaging, and system plants across the United States enables the country to regain manufacturing muscle lost over recent decades. The Fort Worth site joins a supply chain spanning 350-plus factory locations in 30 countries and supports Nvidia’s stated goal of manufacturing up to $500 billion in advanced AI platforms domestically. By colocating production with design teams and leveraging digital-twin simulation of the entire facility before ground was broken, Wistron and Nvidia have compressed the timeline from concept to volume output. Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems
Vera CPU Challenges Intel and AMD in Agentic Workloads
Nvidia has begun shipping its Vera CPU to customers including OpenAI, Anthropic, and SpaceX. The chip’s Olympus cores are optimized for the irregular, branch-heavy execution paths that dominate agentic AI—tasks that require sustained single-thread performance, high core-to-core bandwidth, and low memory latency under heavy concurrency. Nvidia claims the design delivers twice the single-threaded performance, three times the core-to-core bandwidth, and 40 percent lower memory latency versus competing chiplet-based CPUs.
This focus reflects a reversal in server architecture priorities. Early AI servers paired eight GPUs to a single CPU; agentic systems now require the CPU to orchestrate continuous reasoning loops, tool calls, and verification steps. AMD and Intel have benefited from this shift, with their shares rising 128 percent and 149 percent respectively in 2026, outpacing Nvidia’s more modest gain. Yet Nvidia’s vertical integration of CPU, GPU, NVLink, and networking inside a single rack-scale offering gives it leverage that discrete CPU vendors cannot easily match. Nvidia details its next-generation Vera CPU for AI, setting up challenge to AMD and Intel
Rubin GPU and Spectrum-6 Target Intelligence per Dollar
The Rubin GPU itself integrates two reticle-limited dies connected by Nvidia’s High-Bandwidth Interface, yielding 336 billion transistors and 224 streaming multiprocessors. A third-generation Transformer Engine supports up to 50 petaflops of NVFP4 inference while preserving model accuracy across varying precision formats. When paired with the new Spectrum-6 Ethernet switch delivering 102.4 terabits per second—twice the capacity of the prior generation—the platform claims 1.6 times higher RDMA bandwidth and significantly lower latency than standard Ethernet fabrics.
Early adopters including CoreWeave, Microsoft, Nebius, SpaceXAI, and Tesla are already installing Spectrum-6 systems. The networking improvements directly address the coordination overhead that arises when hundreds of thousands of GPUs must operate as a single logical compute unit across gigascale AI factories. By extending the same co-design philosophy from silicon to the network fabric, Nvidia aims to convert raw hardware gains into measurable reductions in cost per token—the metric that ultimately determines intelligence per dollar in continuous post-training loops. Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
Broader Implications for AI Infrastructure Competition
The simultaneous rollout of the Vera Rubin platform, Vera CPU, Spectrum-6 networking, and expanded U.S. manufacturing capacity illustrates Nvidia’s determination to control every layer of the AI stack. While AMD and custom-accelerator developers will continue to press on specific components, Nvidia’s ability to optimize the entire rack—from power delivery to software scheduling—creates compounding advantages that are difficult to replicate piecemeal.
As agentic workloads shift more execution burden onto the CPU and demand sustained efficiency at extreme scale, the industry’s focus is moving from peak FLOPS to verifiable gains in tokens per watt and intelligence per dollar. Nvidia’s latest disclosures suggest it intends to set the pace for that transition, forcing competitors to respond not only with faster chips but with equally integrated systems.