Chapter 1
A video is a stack of images
Pictures over time make a block of numbers. And that block is huge.
A video is not one moving thing. It is a quick run of still pictures, called frames. Show them fast enough and your brain sees smooth motion. Cinema settled on 24 frames per second (fps) when sound arrived in the late 1920s, and it still uses it (see FilmClear). TV and phones often use 25, 30 or 60.
Stack the frames one behind another and you get a block: two directions of space and one of time. This is the space-time cube, and it is exactly what an AI video model has to make. Cut the cube across and you get one frame. Cut it along time, one pixel row through every frame, and the ball's motion becomes a curved path.
Each frame is a grid of pixels, and each pixel is 3 numbers: red, green and blue. So one second of Full HD video at 24 fps is 1920 × 1080 × 3 × 24 ≈ 149 million numbers. Reading this page, you take in about 5 tokens of text a second. That gap, millions of times, is why making video is so much harder than making text (see LLMClear for tokens).
Real files are compressed, so a 1080p stream is closer to 5 megabits a second. But a model has to decide every pixel of every frame, and keep them all agreeing with each other.


