Last updated: September 2026. Editorial Team — researched using primary announcements from Nvidia and reporting from The Daily Star and Nvidia’s own investor materials. See “Sources & Methodology” for our full source list.
Quick Answer
Nvidia officially kicked off its next generation of AI hardware at CES on January 5, 2026, launching the Rubin platform: six new chips designed to work as one integrated AI supercomputer, delivering up to a 10x reduction in inference token cost and a 4x reduction in the number of GPUs needed to train mixture-of-experts (MoE) models compared to the current Blackwell platform. CEO Jensen Huang also announced a major strategic shift: Nvidia now plans to release a new family of AI chips every year, roughly double its prior two-year release cadence. Rubin-based products are entering full production and will reach cloud partners including AWS, Google Cloud, Microsoft, and Oracle Cloud Infrastructure in the second half of 2026.
What’s Actually in the Rubin Platform
Rubin isn’t a single chip — it’s a coordinated family of six components, each handling a different piece of the AI infrastructure puzzle, according to Nvidia’s official announcement. The lineup includes the Rubin GPU itself, a new custom CPU called Vera, the NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet switch. Nvidia’s engineering approach, which it calls “extreme codesign,” involves optimizing all six components to work together as a single system rather than as separate parts bolted together — the company credits this integrated design specifically for the dramatic efficiency gains over Blackwell.

Photo by ClickerHappy via Pexels
The Specific Performance Claims
Nvidia’s official press materials put concrete numbers behind the platform’s headline claims: Rubin delivers up to a 10x reduction in the cost of generating inference tokens (the actual computational output of a running AI model) and a 4x reduction in the number of GPUs required to train mixture-of-experts models, both measured against the current-generation Blackwell platform. On the networking side, the new NVIDIA Spectrum-X Ethernet Photonics switch systems deliver 5x improved power efficiency and uptime — a detail that matters because networking, not just raw compute, has increasingly become a bottleneck in large-scale AI training clusters. Nvidia also introduced a new Inference Context Memory Storage Platform built around the BlueField-4 storage processor, specifically designed to accelerate the kind of multi-step reasoning that agentic AI applications require.
An Accelerated Release Schedule
Perhaps the most consequential announcement wasn’t about Rubin’s specs at all, but about Nvidia’s future release cadence. Speaking at National Taiwan University in Taipei as part of the Computex trade show, Huang said the company now plans to release a new family of AI chips every year, a significant acceleration from its prior schedule of roughly once every two years, according to The Daily Star’s coverage. The new CPU in the lineup, Versa, and the Rubin-family GPUs will bundle next-generation high-bandwidth memory manufactured by SK Hynix, Micron, and Samsung — naming all three major memory suppliers rather than relying on a single vendor, a detail worth noting given how tight memory supply chains have been amid surging AI demand.
Who’s Deploying It First
Nvidia’s official announcement names the specific early adopters: Microsoft’s next-generation Fairwater AI superfactories will feature NVIDIA Vera Rubin NVL72 rack-scale systems, scaling to hundreds of thousands of Vera Rubin Superchips as part of future Fairwater AI superfactory sites. Among the first cloud providers deploying Vera Rubin-based instances in 2026 will be AWS, Google Cloud, Microsoft, and Oracle Cloud Infrastructure, alongside Nvidia Cloud Partners CoreWeave, Lambda, Nebius, and Nscale — a roster that spans both the largest hyperscalers and the newer, AI-specialized cloud infrastructure providers that have emerged specifically to serve GPU-hungry AI workloads.
What Comes After Rubin
Nvidia’s roadmap, as outlined at its earlier GTC 2025 conference, extends well beyond the initial Rubin launch. Rubin Ultra — a package combining four GPUs into a single unit delivering up to 100 petaflops of inference performance — is planned for the second half of 2027. Beyond that sits Feynman, named after physicist Richard Feynman, expected sometime in 2028 as Rubin’s eventual successor. Huang offered few architectural details about Feynman at the time, saying only that it would also pair with a Vera-family CPU.
Why This Matters Beyond Nvidia
The accelerated annual release cadence has implications reaching well past Nvidia’s own product roadmap. A faster chip refresh cycle means cloud providers, AI labs, and enterprises building on Nvidia hardware need to plan for more frequent infrastructure upgrade cycles than the two-year cadence they’d become accustomed to. It also intensifies competitive pressure on rival chipmakers, who now face a moving target that resets annually rather than biennially — a genuine strategic challenge for any competitor trying to plan a multi-year product roadmap against Nvidia’s accelerating pace.
Frequently Asked Questions
What is Nvidia’s Rubin platform?
Rubin is Nvidia’s next-generation AI chip platform, launched January 5, 2026, comprising six coordinated chips including a new GPU, the Vera CPU, and supporting networking and storage components, designed to work together as a single AI supercomputer system.
How much faster is Rubin than Blackwell?
Nvidia claims up to a 10x reduction in inference token generation cost and a 4x reduction in the number of GPUs needed to train mixture-of-experts models, compared to the current Blackwell platform.
When will Rubin-based products be available?
Rubin is in full production, with Rubin-based products from cloud partners including AWS, Google Cloud, Microsoft, and Oracle Cloud Infrastructure becoming available in the second half of 2026.
How often does Nvidia plan to release new AI chips now?
Nvidia CEO Jensen Huang announced the company now plans to release a new family of AI chips annually, roughly double its prior release cadence of once every two years.
Sources & Methodology
This article draws on primary sources including: Nvidia’s official investor press release, “NVIDIA Kicks Off the Next Generation of AI With Rubin” (January 5, 2026); Nvidia’s Hot Chips 2026 conference materials; The Daily Star’s coverage of Huang’s Computex remarks on the accelerated chip release schedule; and TechCrunch’s March 2025 reporting on the original Rubin roadmap announcement at GTC 2025. Figures reflect the most recently published information as of this article’s last-updated date.
This article is for informational purposes and does not constitute investment advice.
