NodeValid

Published video-model weights, licences and hardware notes

Step-Video-T2V shows sizes only in its table

The two pixel sizes on these pages, 768px768px and 544px992px, are the configurations the memory table was measured at. They are the only output sizes stated, and they arrive attached to gigabytes rather than to a specification. As of 2026-09-22.

Step-Video-T2V in the resolution column, as publishedWhere a memory table is the only place output sizes appear, the resolution column is borrowing its values from the hardware column.Output size768px768px and544px992pxsizes the memoryfigureMemory it sizes78.55 GB peak at768px768px204fheld over thismany framesClip lengthUp to 204 framesSettings doing double dutyStated as the memory table's configurations
Fig. 1 Sizes that exist on the page because somebody measured memory at them.
Step-Video-T2V in the resolution column, with the columns that qualify it. Read from github.com/stepfun-ai on 2026-09-22.
ColumnWhat the pages state
Resolution, in the vendor's own words768px768px and 544px992px
The detail published beside itStated as the memory table's configurations
Hardware, from the same pages78.55 GB peak at 768px768px204f
Length, from the same pagesUp to 204 frames

Inclusion rule. The value this page is about, the detail the vendor attached to it, and the two columns that change how it should be read. A column the vendor left empty keeps its row and says so. Order. This column first, then its detail, then the two columns that qualify it.

1A square and a portrait frame

768px768px and 544px992px are different shapes, not two points on one scale. Publishing both means both were exercised, which is more than a single figure would establish.

What it does not establish is whether anything between or beyond them works, because neither was presented as a supported size.

2The memory spread across them is narrow

78.55 GB at the square setting with 204 frames, 77.64 GB at the portrait setting with the same frames, 72.48 GB at the portrait setting with 136. Changing the frame shape barely moves the requirement.

For a thirty billion parameter model that is expected: the weights dominate, and the frame is a small part of the total. It also means output size is not the lever for fitting this model onto a smaller card.

3Frames without a rate, again

The 204 frame ceiling has no frame rate attached anywhere in the repository, so the output sizes cannot be turned into a clip length in seconds.

The resolution and length columns here are both populated from the same table, which makes the entry coherent and leaves the specification a reader would want unwritten.

4Sources

Output sizes above appear on github.com/stepfun-ai, checked 2026-09-22. The full entry is at this family's page; what counts as a stated size is on resolution. Nearby: 256px and 768px, 576x1024 from training.