Which releases document running across several cards
Two of the 15 families document running across more than one card. Three state explicitly that their figure is for a single GPU, four name a card without a count, three give gigabytes with no card, and three state nothing. As of 2026-09-12.
| Item | What is published | Detail |
|---|---|---|
| A per-card multi-GPU figure | One family | 44.3GB across four or more H100 or H800 cards |
| A card count for a model | One family | Eight H100 or H800 cards for a 24B model |
| A single GPU stated explicitly | Three families | On a single GPU; on a single 80G GPU; single GPU memory usage |
| A card named without a count | Four families | An A100 80GB; RTX 30XX to 50XX; three chip architectures; a consumer 4090 |
| Gigabytes and no card | Three families | A figure with no machine shape attached |
| Nothing stated | Three families | No figure and no card |
Inclusion rule. One row per way of describing the machine a figure was measured on. A row is kept where one family uses that form, because multi-card documentation is rare. Order. From the most completely described multi-card statement to no statement.
1Sharding does not divide the requirement
One family states 60.3GB on one card and 44.3GB across four or more for the same output. A reader dividing by four would budget about fifteen gigabytes per card and be short by nearly thirty, because activations and buffers are replicated.
That is the single most useful thing a vendor can publish about multi-card running, and it appears once in this register. It is recorded as one vendor's measurement rather than as a scaling rule.
2A card count is a different kind of specification
Eight named data-centre cards specifies a node: the interconnect, the driver stack and the formats the run was exercised against. Read as a memory total it is a large number and a useless one.
The same release names one consumer card for its smaller model, so the count form works at both ends. What it does not give is a per-card figure somebody can match against hardware they own.
3Stating a single card is worth doing
Three families say explicitly that their figure is for one GPU, which removes an ambiguity a reader would otherwise carry. A figure with no machine shape could be a total across an unstated number of devices.
It is the cheapest form of precision in this column: four words, and the figure becomes comparable with every other single-card number here.
- Hardware60.3GB peak GPU memory at 768px on one GPU, and 44.3GB across four or moremeasured on H100 or H800 cards
- HardwareH100/H800 x 8 for the 24B model, RTX 4090 x 1 for the 4.5Bstated as a count of a named card
- HardwareApproximately 60GB VRAM on a single GPU, with at least one H100 recommendedthe vendor's own figure
- Hardware9.3G BF16 with cpu_offload, against 27.5G if CPU offload is not enabledsingle GPU memory usage on the same run
- Hardware78.55 GB peak GPU memory at 768px768px204f, 77.64 GB at 544px992px204f and 72.48 GB at 544px992px136fstated per configuration at batch size one
4Sources
Repository statements and hosted rates are linked from the model pages and the how read page, checked 2026-09-12. Hardware figures are set out on requirements. Nearby: Empty cells, Two kinds of open, The VRAM ladder.