NodeValid

Published video-model weights, licences and hardware notes

Step-Video-T2V: what the repository publishes

Step-Video-T2V publishes 30 billion parameters under MIT, the shortest licence text in this register attached to its largest published model. Peak GPU memory is stated per configuration: 78.55 GB at 768px768px with 204 frames, 72.48 GB at 544px992px with 136. As of 2026-09-22.

Three configurations from one repositoryPeak GPU memory is published three times: 78.55 GB at 768px768px with 204 frames, 77.64 GB at 544px992px with the same frame count, and 72.48 GB at 544px992px with 136 frames. Cutting frames saves several gigabytes; reshaping the frame saves about one.Most memory asked for768px768px, 204 frames78.55 GB peak, the heaviest configuration published.544px992px, 204 frames77.64 GB peak. Reshaping the frame changes little.544px992px, 136 frames72.48 GB peak. Dropping 68 frames is what saves memory.Least, and the frame count is why
Fig. 1 Ordered by the memory each row asks for. The gap between the second and third rows is frames; the gap between the first and second is frame shape.
The five fields this register keeps, for Step-Video-T2V. Read from the vendor's own repository, with the licence confirmed on its model card on 2026-09-22.
FieldWhat is statedDetail
LicenceMITThe same tag on the repository and the model card
Size30 billion parametersFrom the repository's own introduction
Hardware78.55 GB peak at 768px768px204f77.64 GB and 72.48 GB at two smaller settings
Resolution768px768px and 544px992pxStated as the memory table's configurations
LengthUp to 204 framesFrames, with no frame rate attached

Inclusion rule. The same five fields for every family here. A field the repository does not state is recorded as not stated rather than estimated from the parameter count. Order. Fixed field order, identical on every family page.

1A four-paragraph licence on thirty billion parameters

MIT imposes attribution and stops. It is the shortest licence in this register and it sits on the largest model here, which is a clean demonstration that the licence column and the hardware column have nothing to say to each other. Permission does not predict cost.

Read practically, the rights question closes in a minute and the capacity question stays wide open. A permissive licence on a model that peaks near 79 gigabytes is permission to rent a card, not permission to avoid renting one.

2Three configurations, three figures, and the frame count is the dial that moves them

The repository publishes a table rather than a number: 78.55 GB at 768px768px with 204 frames, 77.64 GB at 544px992px with the same frame count, and 72.48 GB at 544px992px with 136 frames. Cutting 68 frames saves about five gigabytes. Changing the shape of the frame saves about one.

This is the rare entry where a reader can see which dial matters. Length moves this model's memory several times as much as resolution does, which inverts the assumption most capacity planning starts from, and it is only visible because three rows were published instead of one.

3What 204 frames is, and what the repository declines to convert it into

The stated ceiling is expressed in frames: up to 204 of them. No frame rate travels with it in the requirements table, so how many seconds that amounts to depends on a number the reader has to supply, and this register does not supply it for them.

The memory figures carry their own qualifiers: batch size one, without cfg distillation, with multi-GPU deployment discussed separately. Those clauses are the difference between a figure somebody can reproduce and a figure that merely sounds like a specification, so the row keeps them attached.

4The same field, across every family

Licence, family by familyOne field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.Licence, family by familyLicenceAllegroAllegro — Licence: apache-2.0CogVideoXCogVideoX — Licence: Split by model sizeCosmos Predict 2.5Cosmos Predict 2.5 — Licence: NVIDIA Open Model LicenseFramePackFramePack — Licence: Apache 2.0HunyuanVideoHunyuanVideo — Licence: tencent-hunyuan-communityLongCat-VideoLongCat-Video — Licence: MIT LicenseLTX-VideoLTX-Video — Licence: Apache 2.0MAGI-1MAGI-1 — Licence: Apache License 2.0Mochi 1Mochi 1 — Licence: Apache 2.0Open-SoraOpen-Sora — Licence: Apache 2.0Pyramid FlowPyramid Flow — Licence: stabilityai-ai-communitySkyReels-V2SkyReels-V2 — Licence: skywork-licenseStable Video DiffusionStable Video Diffusion — Licence: stable-video-diffusion-communityStep-Video-T2VStep-Video-T2V — Licence: MITWanWan — Licence: Apache 2.0
Fig. 2 One field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.
Every family on one field, with Step-Video-T2V marked. Recorded 2026-09-22.
FamilyLicence
Allegroapache-2.0
CogVideoXSplit by model size
Cosmos Predict 2.5NVIDIA Open Model License
FramePackApache 2.0
HunyuanVideotencent-hunyuan-community
LongCat-VideoMIT License
LTX-VideoApache 2.0
MAGI-1Apache License 2.0
Mochi 1Apache 2.0
Open-SoraApache 2.0
Pyramid Flowstabilityai-ai-community
SkyReels-V2skywork-license
Stable Video Diffusionstable-video-diffusion-community
Step-Video-T2VMIT
WanApache 2.0

Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.

Size, family by familyOne field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.Size, family by familySizeAllegroAllegro — Size: VAE 175M and DiT 2.8BCogVideoXCogVideoX — Size: 2B and 5B, carried in the model namesCosmos Predict 2.5Cosmos Predict 2.5 — Size: 2,059,174,912 parametersFramePackFramePack — Size: 13B modelsHunyuanVideoHunyuanVideo — Size: Over 13 billion parametersLongCat-VideoLongCat-Video — Size: 13.6B parametersLTX-VideoLTX-Video — Size: A 13B modelMAGI-1MAGI-1 — Size: 24B and 4.5B modelsMochi 1Mochi 1 — Size: 10 billion parametersOpen-SoraOpen-Sora — Size: An 11B modelPyramid FlowPyramid Flow — Size: No parameter count recordedSkyReels-V2SkyReels-V2 — Size: 1.3B, 5B and 14B variantsStable Video DiffusionStable Video Diffusion — Size: 2B paramsStep-Video-T2VStep-Video-T2V — Size: 30 billion parametersWanWan — Size: A 5B text-to-video and image-to-video model
Fig. 3 One field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.
Every family on one field, with Step-Video-T2V marked. Recorded 2026-09-22.
FamilySize
AllegroVAE 175M and DiT 2.8B
CogVideoX2B and 5B, carried in the model names
Cosmos Predict 2.52,059,174,912 parameters
FramePack13B models
HunyuanVideoOver 13 billion parameters
LongCat-Video13.6B parameters
LTX-VideoA 13B model
MAGI-124B and 4.5B models
Mochi 110 billion parameters
Open-SoraAn 11B model
Pyramid FlowNo parameter count recorded
SkyReels-V21.3B, 5B and 14B variants
Stable Video Diffusion2B params
Step-Video-T2V30 billion parameters
WanA 5B text-to-video and image-to-video model

Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.

Hardware, family by familyOne field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.Hardware, family by familyHardwareAllegroAllegro — Hardware: 9.3G BF16 with cpu_offloadCogVideoXCogVideoX — Hardware: From 4GB in FP16 on the 2BCosmos Predict 2.5Cosmos Predict 2.5 — Hardware: 32.54 GB of GPU VRAMFramePackFramePack — Hardware: 6GB minimum for 1800 frames at 30fpsHunyuanVideoHunyuanVideo — Hardware: 60GB peak at 720px1280px, 129 framesLongCat-VideoLongCat-Video — Hardware: No memory figure recordedLTX-VideoLTX-Video — Hardware: No minimum statedMAGI-1MAGI-1 — Hardware: H100 or H800 x 8 for the 24BMochi 1Mochi 1 — Hardware: About 60GB VRAM on a single GPUOpen-SoraOpen-Sora — Hardware: 60.3GB peak at 768px on one GPUPyramid FlowPyramid Flow — Hardware: No memory figure recordedSkyReels-V2SkyReels-V2 — Hardware: About 14.7GB peak on the 1.3B at 540PStable Video DiffusionStable Video Diffusion — Hardware: About 180s on an A100 80GB cardStep-Video-T2VStep-Video-T2V — Hardware: 78.55 GB peak at 768px768px204fWanWan — Hardware: A consumer 4090 on the 5B release
Fig. 4 One field at a time, every family beside it, Step-Video-T2V marked. Quoted from each repository and never derived from a parameter count. The exact phrasing sits in the table under this figure. Recorded 2026-09-22.
Every family on one field, with Step-Video-T2V marked. Recorded 2026-09-22.
FamilyHardware
Allegro9.3G BF16 with cpu_offload
CogVideoXFrom 4GB in FP16 on the 2B
Cosmos Predict 2.532.54 GB of GPU VRAM
FramePack6GB minimum for 1800 frames at 30fps
HunyuanVideo60GB peak at 720px1280px, 129 frames
LongCat-VideoNo memory figure recorded
LTX-VideoNo minimum stated
MAGI-1H100 or H800 x 8 for the 24B
Mochi 1About 60GB VRAM on a single GPU
Open-Sora60.3GB peak at 768px on one GPU
Pyramid FlowNo memory figure recorded
SkyReels-V2About 14.7GB peak on the 1.3B at 540P
Stable Video DiffusionAbout 180s on an A100 80GB card
Step-Video-T2V78.55 GB peak at 768px768px204f
WanA consumer 4090 on the 5B release

Inclusion rule. Families whose vendor has published weights openly. A family whose repository does not state this particular field keeps its row, marked hollow. Order. Alphabetical by family name.

5Sources

Read from github.com/stepfun-ai on 2026-09-22. Hardware figures across every family are collected on requirements; the licence column is set out on licence and what counts as a published weight is on how read. Previous family: Allegro. Next family: Open-Sora.