From a row of cameras photographing a galloping horse in 1878 to models that turn a sentence into a talking film clip, and the rules that followed.
Video has always been a stack of still pictures shown quickly. For decades computers could only play such stacks back. Then networks learned to invent images, diffusion learned to turn static into pictures, and transformers learned to handle patches of space and time together. By 2024 a sentence could become a realistic clip, and governments, India among them, began requiring labels on AI-made video.
Neural network that generated short videos from scratch
VGAN, Vondrick, Pirsiavash and Torralba, United States
2022
Diffusion model for video
Video Diffusion Models, Ho and colleagues, Google
2023
Widely described as the first publicly available text-to-video model
Runway Gen-2, United States
2021
Open standard for content credentials
C2PA, founded by Adobe, Arm, BBC, Intel, Microsoft and Truepic
2026
Indian rules requiring labels on AI-made media
IT Amendment Rules, MeitY, India
19 June 1878Moving pictures
1878 – 2012
Moving pictures
Photographers and film-makers learn that motion is just many still pictures shown quickly, and settle on 24 of them per second.
1878
19 June 1878
A galloping horse, frame by frame
Eadweard MuybridgePalo Alto, California, United States
Muybridge lined up a row of cameras along a racetrack, each fired by a thread as the horse ran past. The sequence proved that all four hooves leave the ground at once. Shown quickly one after another, the still photos looked like motion.
Why it mattered. It showed that movement can be stored as a stack of still pictures, the idea every video still uses.
The film industry, with the arrival of sound filmsUnited States
When films gained soundtracks in the late 1920s, projectors had to run at one steady speed so voices sounded right. The industry settled on 24 frames per second, which is still the standard for cinema.
Why it mattered. A single minute of film is 1,440 separate pictures, which is why video is so much bigger than one image.
Neural networks learn to invent images and the first tiny videos, and face-swapped "deepfakes" appear.
2014
June 2014
Two networks play a game
Ian Goodfellow and colleaguesUniversité de Montréal, Canada
In a generative adversarial network (GAN), one network makes fake images and another tries to spot them. Each gets better by competing with the other. A year earlier, Kingma and Welling had shown the variational autoencoder, another way to generate pictures.
Why it mattered. For the first time, networks could invent new, realistic-looking images instead of only sorting them.
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan and Surya GanguliStanford University, United States
Borrowing an idea from physics, the team slowly destroyed images by adding noise, then trained a network to run the process backwards. The results were blurry and the idea was mostly ignored for five years.
Why it mattered. It was the seed of diffusion, the method behind most video models today.
Carl Vondrick, Hamed Pirsiavash and Antonio TorralbaMIT, Cambridge, United States
Trained on over two million unlabelled videos, VGAN generated clips of 32 frames at 64 by 64 pixels, up to about a second long. It learned to draw a still background and a moving foreground separately. The paper appeared at NeurIPS 2016.
Why it mattered. It showed a network could learn how scenes move, not just how they look.
A user called "deepfakes" posted videos in which one person's face had been swapped onto another's using neural networks, and shared the code. The name stuck for any AI-faked video. Most early deepfakes targeted women without their consent.
Why it mattered. Faked video became a public worry years before text-to-video models existed.
Diffusion turns static into images, works in a compressed latent space, follows prompts and starts making short videos.
2020
June 2020
Diffusion finally works
Jonathan Ho, Ajay Jain and Pieter AbbeelUC Berkeley, United States
Denoising diffusion probabilistic models (DDPM) trained a network to predict the noise hidden in a picture, then removed it step by step, starting from pure static. The images matched the best GANs, and training was much more stable.
Why it mattered. Almost every image and video generator since is built on this recipe.
Adobe, Arm, BBC, Intel, Microsoft and TruepicUnited States and United Kingdom
Six organisations formed the Coalition for Content Provenance and Authenticity (C2PA) to write an open standard for "content credentials": signed records, attached to a file, of who made it and how it was edited. Many AI tools later used it to mark their output.
Why it mattered. It gave the world a shared way to say where a picture or video came from.
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn OmmerLMU Munich and Heidelberg University, Germany
Instead of denoising every pixel, latent diffusion first squeezes an image with an autoencoder into a much smaller grid of numbers, the "latent". Diffusion happens there, and a decoder turns the result back into pixels. This cut the cost dramatically.
Why it mattered. Video models compress in the same way, in both space and time, or they could never afford to run.
Classifier-free guidance trains one model both with and without the text prompt. At each step it pushes the picture further in the direction the prompt suggests. First shown at a NeurIPS workshop, it was posted on arXiv in July 2022.
Why it mattered. It is the "follow the prompt more closely" dial inside nearly every text-to-image and text-to-video model.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi and David FleetGoogle, United States
Video Diffusion Models denoised all the frames of a short clip together, so each frame could "see" its neighbours. This kept objects steadier from frame to frame than making each frame separately.
Why it mattered. Denoising frames together is the main trick for keeping video consistent.
Stability AI, CompVis at LMU Munich and RunwayMunich, Germany and London, United Kingdom
A latent diffusion model trained on billions of captioned images was released with its weights, so anyone with a gaming graphics card could run it at home. Thousands of tools and many later open video models were built on it.
Why it mattered. Image generation went from lab demo to something millions of people could use and study.
Uriel Singer, Adam Polyak, Devi Parikh, Yaniv Taigman and colleagues at Meta AITel Aviv, Israel and Menlo Park, United States
Meta's team took a text-to-image model and taught it motion from videos that had no captions at all. It produced short, low-resolution clips from a sentence.
Why it mattered. It showed video models could borrow what image models already knew about the world.
Jonathan Ho, William Chan, Chitwan Saharia and colleagues at GoogleMountain View, United States
A chain of seven diffusion models first made a small, jerky clip, then filled in extra frames and added detail. The paper reports 128 frames at 1280 by 768 pixels and 24 frames per second, about 5.3 seconds.
Why it mattered. It reached HD resolution by building a clip up in stages.
William Peebles and Saining XieUC Berkeley and New York University, United States
The diffusion transformer (DiT) cut the latent into small patches and processed them with a transformer, the same design behind chatbots. Bigger transformers gave better pictures, predictably.
Why it mattered. This design, scaled up to patches of space and time, became the backbone of Sora and many later video models.
From about one second in 2016 to minutes in 2024. Figures for company models are as reported by the companies.
2016 VGAN: 32 frames, up to about 1 second
2022 Imagen Video: 128 frames at 24 fps, about 5.3 seconds
2023 Runway Gen-2: about 4 seconds (reported)
2024 Sora: up to 1 minute (reported)
2024 Kling: up to 2 minutes (reported)
2023 – today
Video for everyone
Text-to-video tools reach the public, clips get longer, sharper and gain sound, and their cost becomes clear.
2023
20 March 2023
Text-to-video anyone can try
RunwayNew York, United States
Runway announced Gen-2, which made short clips of about four seconds from a text prompt or an image. It opened to the public in stages over the following months and was widely described as the first publicly available text-to-video model.
Why it mattered. Artists and students began experimenting with AI video themselves.
OpenAI showed Sora, which it reported could make videos up to one minute long. Its technical report described compressing video into a latent space and cutting it into "spacetime patches" for a diffusion transformer. It was not released to the public at first.
Why it mattered. It showed how far scaling a diffusion transformer on video could go, and set off a race.
Google announced Veo, which it reported could make 1080p videos beyond a minute long. Its videos were to carry SynthID watermarks. Veo 2 followed in December 2024.
Why it mattered. Big labs began shipping watermarking alongside video models.
Kuaishou, a short-video company, launched Kling for testing in China. It reported clips up to two minutes long at 1080p and 30 frames per second. Kling soon opened to users around the world.
Why it mattered. Strong video models were now coming from several countries, not one lab.
OpenAI; TencentSan Francisco, United States and Shenzhen, China
On 9 December OpenAI opened Sora to paying users in some countries, with clips up to 1080p and 20 seconds. The same month Tencent released HunyuanVideo, a 13-billion-parameter model with openly published weights, after Tsinghua University and Zhipu AI's open CogVideoX in August.
Why it mattered. Anyone could now study how a large video model is built, not just use one.
Google announced Veo 3, which generated speech, sound effects and music together with the video in a single model. Widely shared clips of talking characters made it hard to tell AI video from real footage at a glance.
Why it mattered. Realistic talking video made labels and provenance more urgent.
MIT Technology Review; Hugging Face researchersUnited States and Canada
Measurements on one open model found a 5-second clip used about 3.4 million joules, more than 700 times a high-quality image. A Hugging Face study found that doubling a video's length roughly quadruples the energy needed.
Why it mattered. Video is the most energy-hungry thing most people ask AI to make.
OpenAI released Sora 2, which added synchronised speech and sound, in a social app where people could put themselves into videos. It spread quickly, and so did complaints about copyright and look-alike videos of real people.
Why it mattered. AI video became something people shared in a feed, raising new questions about consent.
OpenAI announced it would close the Sora app, which shut on 26 April 2026, with the API ending in September. News reports pointed to very high running costs and problems with copyright and deepfakes.
Why it mattered. Making video is expensive, and even a famous product could not cover its costs.
Watermarks, content credentials and new laws in India, China and Europe try to make AI video easy to recognise.
2023
29 August 2023
An invisible watermark
Google DeepMindLondon, United Kingdom
SynthID hid a signal inside AI-made images from Google's Imagen model. People cannot see it, but a detector can find it, even after cropping or compression. In May 2024 Google said it would extend SynthID to video made by Veo.
Why it mattered. Watermarks inside the pixels can survive where a visible label is easily cropped off.
Ministry of Electronics and Information Technology (MeitY)New Delhi, India
In November 2023 a deepfake video of a well-known actor spread widely, and ministers called deepfakes a dangerous form of misinformation. In December MeitY advised all platforms to tell users clearly, at sign-up and when they upload, that impersonation and deceptive content are banned under the IT Rules.
Why it mattered. India began treating deepfakes as a problem platforms must actively deal with.
Ministry of Electronics and Information Technology (MeitY)New Delhi, India
A 1 March advisory asked platforms to seek government permission before offering untested AI models, and to mark AI-made content that could be used as a deepfake with a permanent unique identifier or metadata. After criticism, a 15 March advisory dropped the permission step but kept the call for labels.
Why it mattered. It was India's first push for AI content labels, a step towards later rules.
Cyberspace Administration of China and three other agenciesBeijing, China
China's labelling measures took effect, requiring visible labels on AI-made content such as video, plus hidden labels in the file's data. In Europe, the AI Act's Article 50 requires AI-made deepfakes to be disclosed, with those rules applying from August 2026.
Why it mattered. Several large markets moved from asking to requiring AI labels.
Ministry of Electronics and Information Technology (MeitY)New Delhi, India
MeitY published draft amendments to the IT Rules defining "synthetically generated information". The draft proposed that visible labels cover at least 10% of the screen, or play during the first 10% of an audio clip, and asked the public for comments.
Why it mattered. It was the first draft Indian law written specifically about AI-made media.
Ministry of Electronics and Information Technology (MeitY)New Delhi, India
The amended IT Rules, in force from 20 February 2026, require AI-made audio and video to carry a prominent label and metadata that platforms must not strip. The fixed 10% size was dropped. Platforms must remove unlawful content within 3 hours of a lawful order, and non-consensual intimate images, including deepfakes, within 2 hours.
Why it mattered. India now has some of the world's strictest labelling and takedown rules for AI video.
At 24 frames per second, one minute of film is 1,440 separate pictures.
One open video model used about 3.4 million joules to make a 5-second clip, more than 700 times the energy of a high-quality AI image, MIT Technology Review found in 2025.
A 2025 Hugging Face study found that doubling the length of an AI video roughly quadruples the energy it takes to make.
HunyuanVideo's autoencoder shrinks each frame 8 times in width and height and 4 times in time, so the model works on about 48 times fewer numbers than the raw video.
The word "deepfake" comes from the username of the Reddit account that posted face-swapped videos in 2017.
The people
Who figured it out
EM
Eadweard Muybridge
1830 – 1904 · Photographer · England and United States
Photographed a galloping horse with a row of cameras in 1878, showing motion as a sequence of stills.
IG
Ian Goodfellow
· Computer scientist · United States
Invented generative adversarial networks in 2014, which powered early image generators, video GANs and deepfakes.
JS
Jascha Sohl-Dickstein
· Physicist and computer scientist · United States
Led the 2015 paper that first used diffusion, adding and removing noise, to generate images.
JH
Jonathan Ho
· Computer scientist · United States
Lead author of DDPM (2020), classifier-free guidance and the first video diffusion papers.
RR
Robin Rombach
· Computer scientist · Germany
Lead author of latent diffusion, the compressed-space method behind Stable Diffusion and most video models.
SX
Saining Xie
· Computer scientist, New York University · China and United States
Co-created the diffusion transformer (DiT) with William Peebles, the design later scaled up for video.
DP
Devi Parikh
· Computer scientist, Meta AI · United States
A research director on Make-A-Video (2022), one of the first text-to-video diffusion systems.
SL
Sasha Luccioni
· AI and climate researcher, Hugging Face · Canada
Co-wrote the 2025 study showing that doubling an AI video's length roughly quadruples its energy use.