Step-Video-T2V: what the repository publishes
Step-Video-T2V publishes 30 billion parameters under MIT, the shortest licence text in this register attached to its largest published model. Peak GPU memory is stated per configuration: 78.55 GB at 768px768px with 204 frames, 72.48 GB at 544px992px with 136. As of 2026-09-22.
| Field | What is stated | Detail |
|---|---|---|
| Licence | MIT | The same tag on the repository and the model card |
| Size | 30 billion parameters | From the repository's own introduction |
| Hardware | 78.55 GB peak at 768px768px204f | 77.64 GB and 72.48 GB at two smaller settings |
| Resolution | 768px768px and 544px992px | Stated as the memory table's configurations |
| Length | Up to 204 frames | Frames, with no frame rate attached |
Inclusion rule. The same five fields for every family here. A field the repository does not state is recorded as not stated rather than estimated from the parameter count. Order. Fixed field order, identical on every family page.
1A four-paragraph licence on thirty billion parameters
MIT imposes attribution and stops. It is the shortest licence in this register and it sits on the largest model here, which is a clean demonstration that the licence column and the hardware column have nothing to say to each other. Permission does not predict cost.
Read practically, the rights question closes in a minute and the capacity question stays wide open. A permissive licence on a model that peaks near 79 gigabytes is permission to rent a card, not permission to avoid renting one.
2Three configurations, three figures, and the frame count is the dial that moves them
The repository publishes a table rather than a number: 78.55 GB at 768px768px with 204 frames, 77.64 GB at 544px992px with the same frame count, and 72.48 GB at 544px992px with 136 frames. Cutting 68 frames saves about five gigabytes. Changing the shape of the frame saves about one.
This is the rare entry where a reader can see which dial matters. Length moves this model's memory several times as much as resolution does, which inverts the assumption most capacity planning starts from, and it is only visible because three rows were published instead of one.
3What 204 frames is, and what the repository declines to convert it into
The stated ceiling is expressed in frames: up to 204 of them. No frame rate travels with it in the requirements table, so how many seconds that amounts to depends on a number the reader has to supply, and this register does not supply it for them.
The memory figures carry their own qualifiers: batch size one, without cfg distillation, with multi-GPU deployment discussed separately. Those clauses are the difference between a figure somebody can reproduce and a figure that merely sounds like a specification, so the row keeps them attached.
4The same field, across every family
| Family | Licence |
|---|---|
| Allegro | apache-2.0 |
| CogVideoX | Split by model size |
| Cosmos Predict 2.5 | NVIDIA Open Model License |
| FramePack | Apache 2.0 |
| HunyuanVideo | tencent-hunyuan-community |
| LongCat-Video | MIT License |
| LTX-Video | Apache 2.0 |
| MAGI-1 | Apache License 2.0 |
| Mochi 1 | Apache 2.0 |
| Open-Sora | Apache 2.0 |
| Pyramid Flow | stabilityai-ai-community |
| SkyReels-V2 | skywork-license |
| Stable Video Diffusion | stable-video-diffusion-community |
| Step-Video-T2V | MIT |
| Wan | Apache 2.0 |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
| Family | Size |
|---|---|
| Allegro | VAE 175M and DiT 2.8B |
| CogVideoX | 2B and 5B, carried in the model names |
| Cosmos Predict 2.5 | 2,059,174,912 parameters |
| FramePack | 13B models |
| HunyuanVideo | Over 13 billion parameters |
| LongCat-Video | 13.6B parameters |
| LTX-Video | A 13B model |
| MAGI-1 | 24B and 4.5B models |
| Mochi 1 | 10 billion parameters |
| Open-Sora | An 11B model |
| Pyramid Flow | No parameter count recorded |
| SkyReels-V2 | 1.3B, 5B and 14B variants |
| Stable Video Diffusion | 2B params |
| Step-Video-T2V | 30 billion parameters |
| Wan | A 5B text-to-video and image-to-video model |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
| Family | Hardware |
|---|---|
| Allegro | 9.3G BF16 with cpu_offload |
| CogVideoX | From 4GB in FP16 on the 2B |
| Cosmos Predict 2.5 | 32.54 GB of GPU VRAM |
| FramePack | 6GB minimum for 1800 frames at 30fps |
| HunyuanVideo | 60GB peak at 720px1280px, 129 frames |
| LongCat-Video | No memory figure recorded |
| LTX-Video | No minimum stated |
| MAGI-1 | H100 or H800 x 8 for the 24B |
| Mochi 1 | About 60GB VRAM on a single GPU |
| Open-Sora | 60.3GB peak at 768px on one GPU |
| Pyramid Flow | No memory figure recorded |
| SkyReels-V2 | About 14.7GB peak on the 1.3B at 540P |
| Stable Video Diffusion | About 180s on an A100 80GB card |
| Step-Video-T2V | 78.55 GB peak at 768px768px204f |
| Wan | A consumer 4090 on the 5B release |
Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.
- Licencemit, the same tag on the repository and on the model cardstated in the repository
- Model size30 billion parametersfrom the repository's own introduction
- Hardware78.55 GB peak GPU memory at 768px768px204f, 77.64 GB at 544px992px204f and 72.48 GB at 544px992px136fstated per configuration at batch size one
- LengthVideos up to 204 framesin frames, with no frame rate attached
5Sources
Read from github.com/stepfun-ai on 2026-09-22. Hardware figures across every family are collected on requirements; the licence column is set out on licence and what counts as a published weight is on how read. Previous family: Allegro. Next family: Open-Sora.