GPU Fleet
NVIDIA B200, GB300 & VR200. Next Up
A liquid-cooled fleet built on NVIDIA Blackwell and Rubin, from the B200 GPU to rack-scale GB300 and VR200 systems, engineered for large-scale AI training, inference, and HPC.
B · Blackwell GPU
NVIDIA B200
A single Blackwell GPU you pair with your own CPUs, for teams training frontier models and serving them in real time.
Native FP4 and the second-generation Transformer Engine push real-time throughput on large language and mixture-of-experts models far past the Hopper generation.
New FP8 precisions shorten training runs for trillion-parameter models, so experiments finish in a fraction of the time.
More work per watt means the same cluster does far more, at far lower cost, than a comparable H100 fleet.
2.4x an H100's memory at 8 TB/s of bandwidth, keeping even the largest models resident on a single GPU.
GB · Grace CPU + Blackwell GPU
NVIDIA GB300 Next Up
A liquid-cooled rack that behaves like one enormous GPU, built for reasoning and test-time scaling.
Seventy-two Blackwell Ultra GPUs and thirty-six Grace CPUs turn a single rack into an AI factory with up to 50x the output of Hopper.
Higher per-user throughput keeps reasoning models and agents answering in real time, even under heavy concurrency.
Every megawatt does five times the useful work, so scaling up doesn't mean scaling power linearly.
20 TB of HBM3e and 130 TB/s of fifth-generation NVLink give the rack the memory and bandwidth reasoning workloads live on.
VR · Vera CPU + Rubin GPU
NVIDIA VR200 Next Up
The next generation of NVIDIA rack-scale compute, built on HBM4 and coming to Firebird.
A single rack of 144 Rubin GPUs reaches 3.6 NVFP4 exaFLOPS for inference and 1.2 FP8 exaFLOPS for training.
288 GB per GPU and half again the fast memory of GB200 NVL72, headroom for the longest context windows.
Purpose-built Vera CPUs sit beside every Rubin GPU in one NVLink domain, tightly coupling compute and memory.
Rubin slots into the same racks as GB200 and GB300, so the infrastructure you build now carries straight into the next generation.
Ready to build on Firebird infrastructure?
Secure early access to on-demand GPU compute or dedicated bare metal servers designed to scale.