GPU Fleet

NVIDIA B200, GB300 & VR200. Next Up

A liquid-cooled fleet built on NVIDIA Blackwell and Rubin, from the B200 GPU to rack-scale GB300 and VR200 systems, engineered for large-scale AI training, inference, and HPC.

NVIDIA B200 GPU

B · Blackwell GPU

NVIDIA B200

A single Blackwell GPU you pair with your own CPUs, for teams training frontier models and serving them in real time.

01 15x Higher Inference

Native FP4 and the second-generation Transformer Engine push real-time throughput on large language and mixture-of-experts models far past the Hopper generation.

02 3x Faster Training

New FP8 precisions shorten training runs for trillion-parameter models, so experiments finish in a fraction of the time.

03 12x More Efficient

More work per watt means the same cluster does far more, at far lower cost, than a comparable H100 fleet.

04 192 GB HBM3e

2.4x an H100's memory at 8 TB/s of bandwidth, keeping even the largest models resident on a single GPU.

GB · Grace CPU + Blackwell GPU

NVIDIA GB300 Next Up

A liquid-cooled rack that behaves like one enormous GPU, built for reasoning and test-time scaling.

01 50x AI Factory Output

Seventy-two Blackwell Ultra GPUs and thirty-six Grace CPUs turn a single rack into an AI factory with up to 50x the output of Hopper.

02 10x More Responsive

Higher per-user throughput keeps reasoning models and agents answering in real time, even under heavy concurrency.

03 5x Better Efficiency

Every megawatt does five times the useful work, so scaling up doesn't mean scaling power linearly.

04 1.1 ExaFLOPS FP4

20 TB of HBM3e and 130 TB/s of fifth-generation NVLink give the rack the memory and bandwidth reasoning workloads live on.

VR · Vera CPU + Rubin GPU

NVIDIA VR200 Next Up

The next generation of NVIDIA rack-scale compute, built on HBM4 and coming to Firebird.

01 3.6 ExaFLOPS Inference

A single rack of 144 Rubin GPUs reaches 3.6 NVFP4 exaFLOPS for inference and 1.2 FP8 exaFLOPS for training.

02 HBM4 Memory

288 GB per GPU and half again the fast memory of GB200 NVL72, headroom for the longest context windows.

03 36 Vera CPUs

Purpose-built Vera CPUs sit beside every Rubin GPU in one NVLink domain, tightly coupling compute and memory.

04 Rubin-Ready Today

Rubin slots into the same racks as GB200 and GB300, so the infrastructure you build now carries straight into the next generation.

Ready to build on Firebird infrastructure?

Secure early access to on-demand GPU compute or dedicated bare metal servers designed to scale.