Computing & Semiconductors
Supermicro Ships NVIDIA Vera Rubin NVL72 Racks With DLC-2 Cooling

Super Micro Computer, Inc. (SMCI ) announced on September 23, 2026 that it is now shipping NVIDIA Vera Rubin NVL72 racks integrated with the company’s Data Center Building Block Solutions (DCBBS) and its DLC-2 direct liquid cooling stack.
The San Jose, California-based company said DCBBS now supports all critical computing infrastructure to speed time-to-online for Vera Rubin NVL72 deployments, with one accountable team covering site survey, project design, integration, L11 and L12 testing, on-site deployment, and ongoing support.
Rack Configuration and Liquid Cooling
Each NVL72 rack combines 72 NVIDIA Rubin GPUs and 36 NVIDIA Vera CPUs in a single liquid-cooled rack operating as one machine. Eighteen 1U compute trays, each carrying four Rubin GPUs and two Vera CPUs, connect through nine sixth-generation NVIDIA NVLink switch trays delivering 216 TB/s of scale-up bandwidth. Each rack provides 20.7 TB of HBM4 memory and up to 54 TB of LPDDR5X memory.
The shipping configuration includes Supermicro in-row cooling distribution units rated at 1.8 MW per CDU and deployed with N+1 redundancy, plus optional rear door heat exchangers to capture residual heat load. Supermicro is also providing networking integration and cabling services based on NVIDIA Reference Architecture, covering the AI compute fabric, the converged fabric, and out-of-band management cabling.
According to the release, the Vera Rubin platform was designed in conjunction with direct liquid cooling because greater AI operations per second create heat loads that must move through the full fluid distribution loop, from the cold plates through manifolds to the cooling tower. Supermicro states it tests and validates every rack with the full liquid cooling stack, which it says speeds time-to-online when deployed. The company’s in-house liquid-cooling portfolio spans cold plates, manifolds, hose kits, rack power shelves, in-row and in-rack CDUs, Liquid-to-Air sidecar CDUs, rear door heat exchangers, and facility-side cooling towers.
“We have spent years building the liquid-cooling stack, the manufacturing capacity, and the deployment teams for exactly this moment,” said Charles Liang, president and CEO of Supermicro. “Our customers can now order a Scalable Unit and receive production-ready systems with end-to-end integration, because we design and build every layer between the cold plate and the cooling tower.”
Scalable Units and Deployment Blueprints
Supermicro’s DCBBS Blueprint defines a balanced bill-of-materials for a given power envelope, from 5 MW to gigawatt scale. The Blueprint based on one Vera Rubin NVL72 Scalable Unit provides 1,152 NVIDIA Rubin GPUs with 331 TB of HBM4 memory across 16 compute racks, together with balanced cooling capacity, power delivery, high-performance storage, context memory storage, and networking. A dedicated Supermicro team manages each project across site survey, design, integration, testing, delivery, deployment, and ongoing support.
On its Vera Rubin product page, Supermicro states that each Rubin GPU carries 288 GB of HBM4 memory and that its DLC-2 stack achieves 98+% heat capture. The company sizes the NVL72 Scalable Unit at 227 kW per rack with 1 MW cooling towers, and pairs the 1,152 GPUs with 864 TB of LPDDR5X CPU memory per 5 MW power envelope. The page states the configuration supports NVIDIA’s Context Memory Storage Platform, Spectrum-X Ethernet, and Quantum-X800 InfiniBand, and is managed through Supermicro’s SuperCloud software suite for infrastructure control, deployment automation, and multi-tenant GPU cloud management. Supermicro also lists related Vera Rubin systems on the page, including a 2U HGX Rubin NVL8 system and Vera CPU racks carrying up to 256 liquid-cooled Vera CPUs.
On its Vera Rubin NVL72 page, NVIDIA describes the system as unifying 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs on the third-generation MGX NVL72 rack design, with support from more than 80 MGX ecosystem partners. NVIDIA’s specification table lists 20.7 TB of HBM4 memory at 1,400 TB/s, 65 TB/s of NVLink-C2C bandwidth, 3,168 custom Olympus CPU cores, up to 54 TB of LPDDR5X, 32.4 TB/s of scale-out networking bandwidth, and a 45°C inlet temperature.
NVIDIA claims the NVL72 delivers one-tenth the cost per million tokens and up to 10x more tokens per megawatt than its GB200 NVL72 for agentic inference workloads, and trains mixture-of-experts models with one-fourth the number of GPUs. The comparison bases are published on the page, and NVIDIA notes the performance figures are subject to change.
NVIDIA announced on May 31, 2026 (NVDA ) at GTC Taipei that Vera Rubin was ramping into full production, naming Supermicro among the system builders in full-scale production and among adopters of its DSX AI factory platform, and stating that production shipments were set to begin starting in the fall. NVIDIA described Vera Rubin as a five-rack POD-scale platform ramping across more than 350 factories in 30 countries, and claimed 10x agent throughput at scale versus the previous-generation Grace Blackwell platform.












