
Dev Log / · updated
The NVIDIA DGX B300: An AI Factory in Ten Rack Units
- Infrastructure
NVIDIA calls the DGX B300 a building block for the AI factory — its name for a data center designed to manufacture intelligence at scale rather than simply host applications.
The machine underneath deserves attention on its own terms: eight Blackwell Ultra accelerators, fused into one memory domain by an enormous switching fabric, with an operations layer that runs the system like a production line.
Every design choice serves two workloads: training models that reason through long, multi-step problems and serving those models to real users in real time.
Eight GPUs, one very large GPU
The compute core is 8 × NVIDIA Blackwell Ultra SXM GPUs. Each carries 288 GB of HBM3e, giving the system 2.304 TB of GPU memory and roughly 62 TB/s of aggregate bandwidth.
Memory is the resource reasoning models exhaust first. A model that thinks for thousands of tokens before answering holds every intermediate step in GPU memory, so capacity and bandwidth per query matter as much as raw compute.
Two NVLink Switch Systems connect the eight GPUs at 14.4 TB/s of aggregate bandwidth. Any GPU can reach any other GPU’s memory at near-local speed.
That is what lets the 2.304 TB pool work as one pool, whether a training run spans all eight accelerators or a long-context inference request outgrows one card.
Compute shaped for the reasoning era
NVIDIA quotes up to 144 petaFLOPS of FP4 Tensor Core performance and 72 petaFLOPS of FP8/FP6. That is 1.5× the dense FP4 throughput and 2× the attention performance of the previous-generation DGX B200.
The attention figure is the one to watch. Attention grows more expensive with context length, and long-horizon agents consume context by the yard: whole codebases, accumulated tool output, and session histories thousands of turns deep.
Doubling attention performance changes how much of that work can remain interactive.
Low-precision FP4 and FP6 are the currency of serving. Frontier models at full precision are too expensive to leave resident; quantization makes them economical to keep loaded all day.
The B300 is balanced for that reality: enormous low-precision throughput first, with FP8/FP6 headroom where a workload needs the extra bits.
From peak compute to requests served
Peak petaFLOPS do not tell you how many real requests a node can carry. Request shape, prefix reuse, output length, and the serving scheduler all sit between the hardware specification and useful throughput.
Public GLM-5.2 results provide two anchors. SGLang measured more than 500 output tokens per second for one user on an 8×B300 system. A separate vLLM and DaoCloud deployment measured production latency under high concurrency across three 8-GPU B300 servers.
The explorer turns those measurements into a workload model. An API client sends the complete prompt; the server’s usage response reports afterward how many input tokens came from cache and how many did not.
The Coding preset comes from OpenCode’s published GLM-5.2/5.1 request data: 52,000 cached input tokens, 700 uncached input tokens, and 150 output tokens.
The Heavy coding preset comes from SGLang’s OpenHands agentic replay. Its requests average roughly 80,000 input tokens and 220 output tokens, with a 92% aggregate prefix-cache hit. That produces the preset’s 73,600 cached and 6,400 uncached input tokens.
The Standard preset is an illustrative 50/50 cache split, not an observed public profile. Custom mode lets you replace every token count.
Incoming requests per second describe traffic pressure. Independent active sessions describe how many distinct contexts compete to remain in GPU memory. Scheduler efficiency accounts for batching gaps and idle time inside the disclosed assumptions.
Workload model
What could one 8x B300 node serve?
Describe one request and the traffic around it. The model estimates whether active contexts remain warm and how incoming demand affects queueing latency.
Request profile
Traffic and contention
Model assumptions
- Model
- GLM-5.2-NVFP4
- KV-cache budget
- 1,624 GB
- FP8 KV footprint
- 51.3 kB/token
- Prefill rate
- 10,000 tok/s/GPU
- Decode rate
- 1,923 tok/s/GPU
Serving rates and KV footprint are fixed extrapolation inputs. Scheduler efficiency models queueing, batching gaps, and idle time rather than a hardware setting.
Request queueing layer
Incoming demand versus service rate
10.0 / 54.1 req/s
18.5% of service rate
Requests that queue
0%
Average queue wait
0.0 s
One-second cohort estimate. Sustained overload keeps extending the queue.
Hardware serving capacity
Tokens / second
2.9M
Tokens / month
7.5T
Requests / second
54.1
Requests / month
142.1M
Effective cache hit
98.7%
Observed baseline 98.7%
GPU-seconds / request
0.148
Context residency
About 599 contexts fit warm at 52.9K tokens each.
Where the GPU time goes
700 input tokens computed; 150 output tokens decoded.
Measured anchors
vLLM: 1,923 output tok/s/GPU derived from its dedicated 8-GPU decode node; 11,950 prefill tok/s/GPU at 8K inputs.
Modeled behavior
The profile split is observed response telemetry. Contexts are treated as equal-sized and similarly active; beyond the KV budget, cached hits fall with the resident fraction.
The default Coding profile has a 98.7% observed cache hit. With 300 independent sessions, its active working set fits inside the estimated KV-cache budget, so the model preserves that split.
At 100% scheduler efficiency, one request consumes about 0.148 GPU-seconds. Input processing and output decoding are nearly even. The modeled service rate is about 54.1 requests per second, or 142.1 million in a 730-hour month.
Incoming demand is not stopped at that rate. The scheduler smooths bursts by queueing work, so contention appears as higher latency. The queue figures model one second of arrivals; demand sustained above the service rate keeps growing the backlog.
More independent sessions can also overflow the memory budget. Cached prefixes then evict, old context must be recomputed, and the service rate falls sharply.
For agentic coding, the B300’s 2.304 TB memory pool matters twice: it holds the model and keeps more working context warm.
Everything around the GPUs
A DGX system is a data center in miniature. Two Intel Xeon 6776P processors with up to 4 TB of system memory handle tokenization, agent orchestration, and data movement.
The network side is built for cluster scale. Eight ConnectX-8 VPI ports carry up to 800 Gb/s of InfiniBand or Ethernet each. BlueField-3 DPUs add another 400 Gb/s.
That is enough fabric to make a row of DGX systems behave like one machine.
Power and cooling make their own engineering argument. The system draws roughly 14.5 kW through twelve 3.2 kW power supplies arranged in N+N redundancy. It moves up to 1,500 CFM of air through a 10–30 °C inlet envelope.
Those are ordinary enterprise data-center conditions, chosen deliberately. Frontier compute that demands exotic facilities is compute most organizations cannot deploy.
Mission Control: the factory floor’s operating layer
NVIDIA Mission Control sits above the hardware. It covers cluster provisioning, workload orchestration, telemetry, automated recovery, and health diagnostics.
The pitch is operational rather than computational: enterprise IT gets the operating muscle of a hyperscaler without building one. The system watches, diagnoses, and recovers itself before a human would normally notice a problem.
That is what “factory” means in NVIDIA’s framing. Factories are defined by uptime and throughput, not peak benchmark numbers.
Mission Control turns a rack of fast hardware into infrastructure that can be run to a production standard.
Built for the data centers that exist
The chassis follows Open Compute Project standards. It is air-cooled, occupies 10 rack units, and can use either traditional AC power or an MGX-compatible design — the first time NVIDIA has offered that choice on a DGX system.
The flexibility is the quiet headline. Frontier-scale compute has historically implied hyperscale-only facilities; the B300 is engineered for data centers organizations already operate.
The building block, assembled
The design brief is clear: vast low-precision compute and memory bandwidth, eight GPUs joined as one, host and network systems built to keep them fed, and an operations layer that treats the system as a production line.
The DGX B300 is not merely a server with GPUs in it. It is NVIDIA’s standardized unit for the AI factory — intelligence manufacturing in ten rack units.