Which repositories publish more than one memory figure
Eight of the 15 families publish more than one memory figure. What varies between the figures is different in each case: the offload setting, the frame size, the model size, the card count, the numeric format or the command. As of 2026-09-12.
| Item | What is published | Detail |
|---|---|---|
| The offload setting | One family | 9.3G with cpu_offload against 27.5G without |
| The frame size | One family | 60GB at 720px1280px against 45GB at 544px960px |
| The model size | Two families | About 14.7GB on a 1.3B against around 51.2GB on a 14B, both at 540P; and a three-step ladder from 4GB to 10GB |
| The card count | One family | 60.3GB on one GPU against 44.3GB across four or more |
| The numeric format | One family | At least 24GB against at least 12GB on an fp8 build |
| The command | One family | A consumer 4090 on one model against at least 80GB on others |
| The configuration | One family | 78.55 GB, 77.64 GB and 72.48 GB across three settings |
| One figure only | Three families | A single requirement or floor, with no second point |
| No memory figure at all | Four families | A timing on a named card, or nothing stated in this column |
Inclusion rule. One row per variable a vendor isolated. Families publishing a single figure keep a row, because a single point cannot be read as a gradient. Order. From the variable a reader controls most easily to the least.
1A pair is worth far more than a point
Two figures with one variable between them tell a reader which direction a number moves and by roughly how much. A single figure tells them one configuration and leaves every adjustment to guesswork.
The most useful pairs here hold the output constant and change the model, or hold the card constant and change the frame. Both forms let a reader extrapolate with some confidence.
2Not every pair is a comparison
Where two figures come from different tables at different resolutions on different model sizes, they describe two runs rather than a difference. Reading a ratio out of such a pair produces a number with no meaning.
The register keeps the configuration attached to each figure so that a reader can see which pairs are controlled. Several are, and they are the rows worth planning from.
3The card-count pair is the one most often assumed
One family publishes a single-card figure and a four-card figure for the same output, and the second is about a quarter lower rather than a quarter of the first. A reader assuming linear scaling would be short by a large margin.
Nobody else in the register publishes that pair, so it is the only measured answer here to what sharding buys. It is recorded as one vendor's measurement rather than as a rule.
- Hardware9.3G BF16 with cpu_offload, against 27.5G if CPU offload is not enabledsingle GPU memory usage on the same run
- HardwarePeak GPU memory of 60GB at 720px1280px and 129 frames, and 45GB at 544px960pxthe vendor's own figures
- HardwareAbout 14.7GB peak VRAM for a 540P video on the 1.3B model, and around 51.2GB on the 14Bthe same output size on two model sizes
- Hardware60.3GB peak GPU memory at 768px on one GPU, and 44.3GB across four or moremeasured on H100 or H800 cards
- HardwareFor 4.5B models, any machine with at least 24GB of GPU memory is sufficient, and a distilled fp8 configuration works on GPUs with at least 12GBa floor for the model and a lower one for a build
4Sources
Repository statements and hosted rates are linked from the model pages and the how read page, checked 2026-09-12. Hardware figures are set out on requirements. Nearby: Which variant, Consumer cards, Stated settings.