How do AI video models work?

Every AI video starts as pure static. Frames stacked over time make a space-time cube. One second of Full HD at 24 fps is about 149 million numbers, millions of times more than the text you read in that second.

Read how it works: How do AI video models work? · The history of AI video models · More boxes on Glassbox

Chapters in this interactive model

  1. A video is a stack of images: Pictures over time make a block of numbers. And that block is huge.
  2. Diffusion: from noise to picture: Start with static. Guess the picture hiding in it. Remove a little noise. Repeat.
  3. Squeeze it, then cut it into patches: Compress the video into a small latent cube, chop it into space-time tokens, and let a transformer attend.
  4. Words steer the noise: A prompt nudges every denoising step. Guidance decides how hard.
  5. Keeping it consistent: Every frame can look fine and the video can still be wrong.
  6. Cost, trust and good use: Video is the most expensive thing AI makes, and the easiest to mistake for real.
Glassbox VideoModelClear

Drag to orbit · scroll or pinch to zoom