Diffusion transformer, named in a count
A diffusion transformer is the component that turns noise into a latent video over many steps. When a card labels part of its parameter count DiT, it is naming the architecture at the same time as the size. As of 2026-09-22.
| Question | Answered by the label |
|---|---|
| What kind of model is this | Yes |
| Which part is large | Yes |
| How many denoising steps | No |
| How long a clip it produces | No |
Inclusion rule. Questions a labelled count can and cannot answer. A question it cannot answer is kept, because the label is often over-read. Order. The two it answers first, then the two it does not.
1Labels are not marketing
DiT and VAE are specific terms with specific meanings. A reader who knows them knows roughly how the run behaves, which part quantising would help, and where the memory will go.
Compare a card that states a single round figure. Same order of magnitude, far less to go on, and no way to reason about the memory statement beside it.
2What the label still does not settle
Step counts, scheduler choices, attention implementation and training data are all outside it, and all of them affect both output and runtime.
The register records the labels the vendor used because they are part of the published value. It does not classify releases by architecture, because that would be a reading rather than a quotation.
3Where the label appears here
One card in the register labels part of its parameter count DiT, which is the only place in this column where the architecture is named inside the value rather than left to a paper.
The rest of the column holds counts without labels, so a reader who wants to know what kind of model a release is has to go elsewhere. The register quotes what was published rather than classifying releases itself.
A reading note, not an entry: no vendor value appears on this page. Where the register records this term for a particular release, it is on the size column and on that family's own page. Nearby terms: Peak memory, Memory floor.