Tech

Nvidia shows off Vera Rubin platform for tokenmaxxing

Published

on

Last week, Nvidia invited a handful of journalists to a briefing about AI infrastructure to help convey the physical reality behind simulated intelligence.

“Infrastructure is physical, it’s real,” said Ian Buck, general manager of Nvidia’s hyperscale and HPC computing business. “You can touch it, you can see it, and it’s what helps bring AI to life. So that is part of the goal here.”

The briefing included a tour of a working Nvidia lab – a mini-datacenter – nestled in a residential district of Sunnyvale, a Silicon Valley suburb. Attendees were asked not to reveal the location – which is one of four such facilities – though they’re easy enough to discover with a bit of online sleuthing.

AI infrastructure is indeed real – or at least realized as revenue when equipment is booked under a bill-and-hold arrangement – and you can touch it if your job involves handling tech gear. 

Advertisement

It’s also controversial as the AI tsunami crests over a society that’s uncertain if the technology will bring bounty or harm. Just one day after the press event, Dutch activists peppered a data center serving Microsoft workloads with chemical-filled balloons to protest climate and political concerns. Nvidia presumably would prefer not to have neighbors grousing about power and water consumption at its facilities, though such concerns may be muted in an area so in thrall to the tech sector.

In any event, Nvidia’s latest data center hardware should require significantly less human intervention than its older kit. During the lab tour, Andrew Bell, a senior VP of hardware engineering, showed off how the company’s Vera Rubin NVL72 compute tray is vastly faster to install than prior hardware. “With automation assembly, the compute tray inside the Vera Rubin NVL72 can be assembled in one minute, compared to GB200 compute tray assembly that typically takes 90 minutes. This represents a 90x improvement,” he said.

The Nvidia event focused on how the company’s Vera Rubin platform, announced at Computex 2024 and detailed at CES in January, advances its vision of AI Factories, a type of computing infrastructure dedicated to running AI workloads.

Inside the Nvidia lab

The platform consists of six chips: the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 network interface, BlueField-4 data processing unit (DPU), and Spectrum-6 Ethernet switch.

Advertisement

Nvidia claims Vera Rubin is said to deliver 10x more tokens per watt than GB200 NVL72 in initial CoreWeave results on DeepSeek-R1.

These appear in five rack-mountable systems designed for data centers: the Vera Rubin NVL72 compute tray, Vera Rubin NVL72 NVLink switch tray, the Groq 3 LPX inference accelerator tray, the Vera CPU tray, the BlueField-4 STX storage tray, and the Spectrum-6 SPX switch tray.

Nvidia’s pitch for all this hardware is that agentic workloads – running AI agents – require a different approach. AI agents work best, the company claims, with sustained inference and low latency across multiple reasoning steps, high decoding throughput, efficient reasoning over long contexts, large key-value cache capacity, and the ability to scale models across linked GPU domains. The company’s AI Factory calls for unified compute capability that spans the data center.

It’s an architectural gambit based on the business of selling tokens. And while GPUs are the star of the show for training and inference, CPUs play a critical role in data processing, networking, scheduling, and storage. Vera represents a bet that making the core faster is better for agentic workloads than increasing the number of cores.

Advertisement

Nvidia claims that the Vera CPU Olympus Core accelerates agents 2x, offers 3x core-to-core bandwidth, and delivers 40 percent lower latency via LPDDR5X memory. And to the extent Nvidia’s data center customers can move more tokens for their AI inference operations, they should be able to capture more revenue.

Nvidia rack system

“Every AI factory is power-constrained,” said Buck. “And frankly the most important metric is your delivered performance in a fixed watt data center, your perf-per watt. All that comes together in your AI factory revenue. You may rent infrastructure … for dollars per hour but you generate revenue with profitable tokens now, in the hundreds of dollars per hour. Your AI Factory revenue is a function of how many tokens you can generate in a fixed watt data center.”

It’s also a function of the market price of tokens. As the supply of tokens increases – through new hardware that makes token generation more efficient and competing AI providers – there’s downward pressure on the cost of tokens. The hope at Nvidia, and more acutely at the frontier AI labs, is that demand for tokens will increase as prices decline, at a rate that leaves enough of a profit margin to recoup the capital expenditures funding the current datacenter construction boom.

But given reports of companies capping employee token budgets, it’s not clear that every organization is seeing the productivity increases touted by Nvidia and its peers.

Advertisement

Buck offered a back-of-the-envelope calculation to suggest that AI makes developers more productive, but it was more of a thought experiment than a verifiable claim.

Pointing to the surge in GitHub code commits – which haven’t been great for GitHub’s stability – he posited there are about 30-40 million software developers active in the world, which translates to about $3 trillion in salaries.

“They’re now producing three times the output or effectively nine trillion dollars of productivity,” Buck said. At Nvidia, he added, “the number of [code] check-ins and developer productivity has tripled as a result” of AI coding tools.

Buck said that Nvidia uses AI agents for pretty much all of its software development processes. “We use agents also for our chip development,” he added. “Our entire chip bug database is all watched and reviewed by agents.”

Advertisement

The leading AI company uses AI to make its AI software and hardware. Flog AI enough and you can bring it to life. ®

Source link

You must be logged in to post a comment Login

Leave a Reply

Cancel reply

Trending

Exit mobile version