How do AI video models work?

Every AI video starts as pure static. Frames stacked over time make a space-time cube. One second of Full HD at 24 fps is about 149 million numbers, millions of times more than the text you read in that second.

Every AI video starts as pure static. Watch a real diffusion model pull a picture out of noise in your browser, steer it with a prompt, see why frames flicker when they are made one at a time, and test whether a hidden watermark survives an edit.

VideoModelClearOpened 22 Sept 202620 min to playFree · no sign-up

In 60 seconds

  1. A video is a huge stack of numbers

    Frames stacked over time make a space-time cube. One second of Full HD at 24 fps is about 149 million numbers, millions of times more than the text you read in that second. A video model has to decide all of them.

  2. Diffusion: noise in, picture out

    Models learn to guess the clean picture hidden under noise. To create, they start from pure static, guess, remove a little noise and repeat, 20 to 50 times. Our toy does it for real with an exact denoiser, which is why it can only copy its 288 pictures; real networks approximate and so can blend.

  3. Squeeze, then cut into patches

    An autoencoder squeezes video into a small latent cube, about 48 times smaller in open models. The cube is cut into space-time patches that act as tokens, and a transformer lets every patch attend to every other one. Attention cost grows with tokens squared.

  4. Words steer every step

    A text encoder turns the prompt into numbers the denoiser reads. Classifier-free guidance mixes a guess made with the prompt and one without: guess = without + w × (with − without). At w = 0 the prompt is ignored; higher obeys more but loses variety.

  5. Consistency is the hard part

    Make each frame alone and the object teleports and changes colour. Denoise the whole clip as one block and motion comes out smooth. Real models still slip on physics, hands and objects that appear from nowhere, because they learn patterns, not rules.

  6. Costly, and easy to mistake for real

    A 5-second clip took about 3.4 million joules in 2025 tests of open models, far more than a picture or a text reply. Deepfakes without consent cause real harm. Watermarks and content credentials help but can be weakened or stripped, and India's 2026 IT Rules require labels.

The history

From a row of cameras photographing a galloping horse in 1878 to models that turn a sentence into a talking film clip, and the rules that followed.

Read the full history
  1. 1878A galloping horse, frame by frame
  2. 2016The first tiny AI videos
  3. 2022Stable Diffusion is released openly
  4. 2025Pictures and sound in one go
  5. 2026India's deepfake rules become law

The full explanation

VideoModelClear, chapter by chapter

Chapter 1

A video is a stack of images

Pictures over time make a block of numbers. And that block is huge.

A video is not one moving thing. It is a quick run of still pictures, called frames. Show them fast enough and your brain sees smooth motion. Cinema settled on 24 frames per second (fps) when sound arrived in the late 1920s, and it still uses it (see FilmClear). TV and phones often use 25, 30 or 60.

Stack the frames one behind another and you get a block: two directions of space and one of time. This is the space-time cube, and it is exactly what an AI video model has to make. Cut the cube across and you get one frame. Cut it along time, one pixel row through every frame, and the ball's motion becomes a curved path.

Each frame is a grid of pixels, and each pixel is 3 numbers: red, green and blue. So one second of Full HD video at 24 fps is 1920 × 1080 × 3 × 24 ≈ 149 million numbers. Reading this page, you take in about 5 tokens of text a second. That gap, millions of times, is why making video is so much harder than making text (see LLMClear for tokens).

Real files are compressed, so a 1080p stream is closer to 5 megabits a second. But a model has to decide every pixel of every frame, and keep them all agreeing with each other.

Try “A stack of images” in the interactive model →

Chapter 2

Diffusion: from noise to picture

Start with static. Guess the picture hiding in it. Remove a little noise. Repeat.

Most AI video and image models today are diffusion models. The trick is surprising: they learn to remove noise. In training, you take a real picture, bury it under random static, and ask a neural network to guess the clean picture. Do that millions of times at every amount of noise, and the network gets very good at it (see NeuralNetClear for how networks learn).

To make something new, start from pure noise. Ask for a guess of the clean picture. The first guess is a blurry average, because it could be anything. Mix that guess back with a bit less noise, and ask again. Step by step, the picture sharpens out of the static. Real models take roughly 20 to 50 steps.

This one is real and runs right here. Its "memory" is 288 small pictures of shapes. Its guess is the mathematically best one: a weighted average of those 288 pictures, weighted by how well each fits the noise. The board shows which pictures it leans on at each step.

What is simplified: our denoiser is perfect for its 288 pictures, so it can only ever end on a copy of one of them. A real network is trained to approximate this guess, for billions of pictures. Its small, smooth errors are what let it mix ideas and make pictures nobody has seen. The same flaw in reverse is memorisation: real models sometimes reproduce training images too closely.

Try “Noise to picture” in the interactive model →

Chapter 3

Squeeze it, then cut it into patches

Compress the video into a small latent cube, chop it into space-time tokens, and let a transformer attend.

Denoising 149 million numbers a second directly would be far too slow. So video models first squeeze the video. An autoencoder has two halves: an encoder that turns each patch of pixels into a few numbers, and a decoder that turns them back. The small version is called the latent. All the diffusion happens in latent space, and the decoder paints the final pixels at the end.

Ours is real and was trained a moment ago on every frame in the toy clips. It turns each 4×4 block of pixels (48 numbers) into k numbers. Slide k down and watch the picture blur: the squeeze keeps the big shapes and throws away detail. Open video models reportedly shrink time by 4 and each side by 8, keeping 16 numbers per spot: about 48 times smaller.

Next, the latent cube is cut into space-time patches, a few cells wide and a few frames long. Each patch becomes one token, just like a word piece in a language model. OpenAI's Sora report (2024) described this idea as "spacetime patches".

A transformer then lets every token look at every other one with attention (see LLMClear). Click a cell: its token looks for tokens with similar content. The red ball's token finds the ball in other frames. That is how a model keeps things consistent over time. The catch: attention compares every pair, so 100,000 tokens means 10 billion pairs.

Try “Latents and patches” in the interactive model →

Chapter 4

Words steer the noise

A prompt nudges every denoising step. Guidance decides how hard.

So far the model makes something. To make what you ask for, the prompt has to take part in every step. First a text encoder, a language model like the ones in LLMClear, turns your words into a list of numbers. The denoiser reads those numbers alongside the noisy picture, usually through attention. Training pairs millions of videos with their captions, so the network learns which pictures go with which words.

That link is never perfect. So models use a trick called classifier-free guidance. At each step the network makes two guesses: one with the prompt and one without. The difference between them is "what the prompt adds". Guidance multiplies that difference by a number w and adds it back: guess = without + w × (with − without).

Our toy's text encoder is just a lookup for two words, and its grasp of them is deliberately weak, like a partly trained model. At w = 0 it ignores you. At w = 1 it follows the prompt only some of the time. Push w higher and it obeys more often. Real models usually run at about 5 to 7. Too high and pictures turn harsh and samey, with less variety.

The seed picks the starting noise. Same seed and same prompt give the same result every time. Same seed, different prompt: a different result, often with a similar layout, because the noise already leans towards certain places.

Try “Text to video” in the interactive model →

Chapter 5

Keeping it consistent

Every frame can look fine and the video can still be wrong.

A video model has a harder job than an image model. Every frame must look right, and every frame must also agree with the ones around it: the same ball, the same colour, moving smoothly. When that fails you see flicker, morphing (a thing slowly turns into something else) and things that pop in or out of existence.

Here are three real ways to make an 8-frame clip of the same prompt. Frame by frame: each frame is made alone, from its own noise. Each frame is fine, but the ball teleports and can even change colour. Same noise every frame: steady, but frozen, because the same noise always gives the same picture. Whole clip together: a video model denoises all 8 frames as one space-time block, so each frame's guess depends on the others. The motion comes out smooth.

That third way is what modern video models do, with attention across the whole space-time cube (chapter 3). Early text-to-video systems (2022) started from image models and added layers that look across time.

Why it is still hard. The model has no physics engine inside. It learned what videos usually look like, not the rules behind them. So it can get hands wrong, let objects pass through each other, or forget a detail after something walks in front of it. OpenAI itself listed failures like a cookie showing no bite mark after being bitten. Longer videos make it worse: more tokens, more to keep straight.

Try “Keeping it consistent” in the interactive model →

Chapter 6

Cost, trust and good use

Video is the most expensive thing AI makes, and the easiest to mistake for real.

Cost. Video is the heaviest thing these models make. In 2025, MIT Technology Review measured open models: one 5-second video took about 3.4 million joules (roughly 1 kWh), against about 2,300 to 4,400 J for one picture and roughly 100 to 7,000 J for a text reply. A Hugging Face study found the energy grows roughly with the square of the length: twice as long, about four times the energy. Big commercial models do not publish their figures, so treat all of these as estimates. Running models is its own topic (see InferenceClear, being built).

Trust. A realistic fake video of a real person is a deepfake. Making one without the person's consent can hurt them badly, and it is illegal in many places. In India, a viral deepfake of an actor in November 2023 led MeitY to issue an advisory to platforms that December. Amended IT Rules, reported as notified in February 2026, require AI-made audio and video to carry a prominent label, and platforms to remove unlawful content within 3 hours of an official order. The EU and China have labelling rules too.

Two tools help. A watermark hides a faint signal in the pixels that a detector can find, like Google DeepMind's SynthID. Content credentials (the C2PA standard, started in 2021 by Adobe, Microsoft, the BBC and others) attach a signed record of how a file was made. Neither is perfect: watermarks can be weakened, and credentials can be stripped. A video with no label is not proof it is real.

Use. In film, these tools help with storyboards and previs, rough moving sketches before a shoot (see StoryboardClear), concept ideas and some effects work (see VFXClear), with shots still edited and graded by people (see GradeClear). They differ from CGI rendering, which computes light in a 3D scene exactly as built (see CGIClear): a video model predicts pixels it thinks are likely. That makes it fast and flexible, but harder to control. People are still debating training on artists' work without permission, and the effect on jobs.

Try “Cost, trust and use” in the interactive model →

Test yourself

Frequently asked

What is a video, to a computer?

A run of still frames, each a grid of numbers. Each frame is a grid of pixels, and each pixel is three numbers for red, green and blue.

About how many numbers make one second of uncompressed 1080p video at 24 fps?

About 149 million. 1920 × 1080 pixels × 3 colours × 24 frames ≈ 149 million.

In the space-time cube, what does a slice along time (one row through every frame) show?

How things in that row move over time. Motion turns into a path. A still object would make a straight stripe.

What is a diffusion model trained to do?

Guess the clean picture hidden under noise. Training buries real pictures in noise and teaches the network to undo it.

Why is the first guess from pure noise a blurry average?

Pure noise fits every picture equally, so the best guess is their average. With no picture left in the noise, the safest guess is the average of everything it knows.

Our toy only ever makes copies of its 288 pictures. Why can real models make new ones?

Their networks approximate the guess smoothly, so they blend what they learned. A network cannot store every picture exactly, so it learns general patterns, which lets it combine them.

Why do video models work in a latent space?

It is many times smaller than the pixels, so denoising is affordable. Open models reportedly squeeze video about 48 times before any denoising happens.

What is a space-time patch?

A small block of the latent cube, a few cells wide and a few frames long, used as a token. Cutting the latent cube into patches turns video into tokens a transformer can read.

Why does attention get so expensive for long videos?

Each token compares itself with every other token, so the work grows with tokens squared. Twice the tokens means about four times the pairs to compare.

What does the text encoder do?

Turns the prompt into numbers the denoiser can read. The denoiser cannot read words, so a language model turns them into vectors first.

In classifier-free guidance, what happens at w = 0?

The prompt is ignored. guess = without + 0 × (with − without) = the guess without the prompt.

Same seed, different prompt. What do you get?

A different result from the same starting noise. The seed fixes the starting noise, but the prompt steers every step differently.

Why does making each frame separately cause flicker?

Nothing ties one frame to the next, so each frame makes its own choices. Each frame starts from its own noise and never sees its neighbours.

What keeps the "whole clip together" row smooth?

All frames are denoised as one block, so each frame depends on the others. Joint denoising lets every frame see every other frame at every step.

Why do video models still make physics mistakes?

They learned what videos look like, not the rules behind them. They learn patterns in pixels. Nothing inside checks gravity, solidity or counting fingers.

About how much more energy did a 5-second AI video take than one AI picture, in MIT Technology Review's 2025 tests?

More than 700 times. About 3.4 million joules against a few thousand for a picture. The reporters put it at more than 700 times.

A video has no AI label and no content credentials. What does that prove?

Nothing either way. Labels and credentials can be missing or stripped. Their absence proves nothing.

What makes a deepfake harmful even if it is labelled?

Using a real person's face or voice without their consent. Consent matters. A label does not undo the harm of putting words or actions on someone who never agreed.

Words worth knowing

Frame
One still picture in a video; cinema shows 24 every second.
Diffusion model
A generator that starts from random noise and removes it step by step to reveal a picture or video.
Denoiser
The trained network that guesses the clean picture from a noisy one.
Latent
A compressed version of the video, made by an autoencoder, that the model actually works on.
Space-time patch
A small block of the latent cube, a few cells wide and a few frames long, used as one token.
Attention
Each token scores every other token and takes a weighted mix of their information.
Classifier-free guidance
Pushing each step further towards the prompt by mixing guesses made with and without it.
Temporal consistency
Objects staying the same from frame to frame unless they should change.
Content credentials
A signed record attached to a file saying how it was made and edited (the C2PA standard).

Fork it. Teach with it.

This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/videomodelclear”.

git clone https://github.com/bdeeps/videomodelclear.git

Built with three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).

←→ previous / next box · / search

Would you like to see the full page, with the interactive model?