Stable Video Diffusion: what the repository publishes
Stable Video Diffusion publishes 2B params under a licence tagged stable-video-diffusion-community, which sends commercial use to a separate licence page. The card states 25 frames at 576x1024 and calls the results rather short, at 4sec or less. As of 2026-09-22.
| Field | What is stated | Detail |
|---|---|---|
| Licence | stable-video-diffusion-community | Commercial use directed to a separate licence page |
| Size | 2B params | The figure shown on the card itself |
| Hardware | About 180s on an A100 80GB card | A timing on a named card; no memory floor stated |
| Resolution | 576x1024 | What the model was trained to generate |
| Length | 25 frames, 4sec or less | Described as rather short, not as a ceiling |
Inclusion rule. The same five fields for every family here. A field the repository does not state is recorded as not stated rather than estimated from the parameter count. Order. Fixed field order, identical on every family page.
1A licence named after the release, which answers the commercial question by pointing elsewhere
The tag on the card is stable-video-diffusion-community. The card says the model is intended for both non-commercial and commercial usage, that the licence itself covers non-commercial or research purposes, and that commercial use should refer to a separate licence page. Three sentences, and the third is the one that decides anything.
This register records the name and the routing, not a reading of either. A licence that defers the commercial question to another document is a different object from one that settles it in the file, and collapsing both into the word open is the exact mistake this column exists to prevent.
2Length written as an admission rather than a ceiling
The card says the generated videos are rather short, at 4sec or less. That is not a specification and it is more honest than one: a ceiling implies the ceiling is usable, while this phrasing tells a reader the length is a limitation of the release.
Twenty-five frames at 576x1024 is what the model was trained to produce, given a context frame of the same size. Both numbers describe the training rather than switches somebody can turn, which is why this row does not read like a settings list.
3A timing on a named card, and no memory figure at all
The card states that the XT variant takes about 180 seconds on an A100 80GB card, and states no minimum memory anywhere. So the hardware cell here holds a duration and a card name, and the figure most readers arrive looking for is absent.
The combination is still worth something. Three minutes of an 80GB card for four seconds of output is a ratio, and a ratio is what any comparison against a per-second rate actually needs. What it cannot tell anybody is whether the model would load on a card they already own.
4The same field, across every family
| Family | Licence |
|---|---|
| Allegro | apache-2.0 |
| CogVideoX | Split by model size |
| Cosmos Predict 2.5 | NVIDIA Open Model License |
| FramePack | Apache 2.0 |
| HunyuanVideo | tencent-hunyuan-community |
| LongCat-Video | MIT License |
| LTX-Video | Apache 2.0 |
| MAGI-1 | Apache License 2.0 |
| Mochi 1 | Apache 2.0 |
| Open-Sora | Apache 2.0 |
| Pyramid Flow | stabilityai-ai-community |
| SkyReels-V2 | skywork-license |
| Stable Video Diffusion | stable-video-diffusion-community |
| Step-Video-T2V | MIT |
| Wan | Apache 2.0 |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
| Family | Size |
|---|---|
| Allegro | VAE 175M and DiT 2.8B |
| CogVideoX | 2B and 5B, carried in the model names |
| Cosmos Predict 2.5 | 2,059,174,912 parameters |
| FramePack | 13B models |
| HunyuanVideo | Over 13 billion parameters |
| LongCat-Video | 13.6B parameters |
| LTX-Video | A 13B model |
| MAGI-1 | 24B and 4.5B models |
| Mochi 1 | 10 billion parameters |
| Open-Sora | An 11B model |
| Pyramid Flow | No parameter count recorded |
| SkyReels-V2 | 1.3B, 5B and 14B variants |
| Stable Video Diffusion | 2B params |
| Step-Video-T2V | 30 billion parameters |
| Wan | A 5B text-to-video and image-to-video model |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
| Family | Hardware |
|---|---|
| Allegro | 9.3G BF16 with cpu_offload |
| CogVideoX | From 4GB in FP16 on the 2B |
| Cosmos Predict 2.5 | 32.54 GB of GPU VRAM |
| FramePack | 6GB minimum for 1800 frames at 30fps |
| HunyuanVideo | 60GB peak at 720px1280px, 129 frames |
| LongCat-Video | No memory figure recorded |
| LTX-Video | No minimum stated |
| MAGI-1 | H100 or H800 x 8 for the 24B |
| Mochi 1 | About 60GB VRAM on a single GPU |
| Open-Sora | 60.3GB peak at 768px on one GPU |
| Pyramid Flow | No memory figure recorded |
| SkyReels-V2 | About 14.7GB peak on the 1.3B at 540P |
| Stable Video Diffusion | About 180s on an A100 80GB card |
| Step-Video-T2V | 78.55 GB peak at 768px768px204f |
| Wan | A consumer 4090 on the 5B release |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
- Licencestable-video-diffusion-community, with commercial use directed to a separate licence pagethe tag on the model card
- Model size2B paramsthe figure shown on the card
- Resolution and framesTrained to generate 25 frames at resolution 576x1024, given a context frame of the same sizeas a property of the training
- LengthThe generated videos are rather short, at 4sec or lesswritten as a limitation rather than a ceiling
- HardwareSVD-XT takes about 180s on an A100 80GB carda timing on a named card with no memory floor stated
5Sources
Read from huggingface.co/stabilityai on 2026-09-22. Hardware figures across every family are collected on requirements; the licence column is set out on licence and what counts as a published weight is on how read. Previous family: Open-Sora. Next family: FramePack.