The history

The history of AI video models

From a row of cameras photographing a galloping horse in 1878 to models that turn a sentence into a talking film clip, and the rules that followed.

Video has always been a stack of still pictures shown quickly. For decades computers could only play such stacks back. Then networks learned to invent images, diffusion learned to turn static into pictures, and transformers learned to handle patches of space and time together. By 2024 a sentence could become a realistic clip, and governments, India among them, began requiring labels on AI-made video.

148
years
30
moments
8
people
11
places

1878

Motion captured as a sequence of photographs

Eadweard Muybridge, United States

2014

Generative adversarial network

Ian Goodfellow and colleagues, Canada

2016

Neural network that generated short videos from scratch

VGAN, Vondrick, Pirsiavash and Torralba, United States

2022

Diffusion model for video

Video Diffusion Models, Ho and colleagues, Google

2023

Widely described as the first publicly available text-to-video model

Runway Gen-2, United States

2021

Open standard for content credentials

C2PA, founded by Adobe, Arm, BBC, Intel, Microsoft and Truepic

2026

Indian rules requiring labels on AI-made media

IT Amendment Rules, MeitY, India

19 June 1878Moving pictures

1878 – 2012

Moving pictures

Photographers and film-makers learn that motion is just many still pictures shown quickly, and settle on 24 of them per second.

1878

19 June 1878

A galloping horse, frame by frame

Eadweard MuybridgePalo Alto, California, United States

Muybridge lined up a row of cameras along a racetrack, each fired by a thread as the horse ran past. The sequence proved that all four hooves leave the ground at once. Shown quickly one after another, the still photos looked like motion.

Why it mattered. It showed that movement can be stored as a stack of still pictures, the idea every video still uses.

1927

c. 1927

24 frames every second

The film industry, with the arrival of sound filmsUnited States

When films gained soundtracks in the late 1920s, projectors had to run at one steady speed so voices sounded right. The industry settled on 24 frames per second, which is still the standard for cinema.

Why it mattered. A single minute of film is 1,440 separate pictures, which is why video is so much bigger than one image.

2013 – 2019

Networks learn to imagine

Neural networks learn to invent images and the first tiny videos, and face-swapped "deepfakes" appear.

2014

June 2014

Two networks play a game

Ian Goodfellow and colleaguesUniversité de Montréal, Canada

In a generative adversarial network (GAN), one network makes fake images and another tries to spot them. Each gets better by competing with the other. A year earlier, Kingma and Welling had shown the variational autoencoder, another way to generate pictures.

Why it mattered. For the first time, networks could invent new, realistic-looking images instead of only sorting them.

2015

March 2015

Learning to undo noise

Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan and Surya GanguliStanford University, United States

Borrowing an idea from physics, the team slowly destroyed images by adding noise, then trained a network to run the process backwards. The results were blurry and the idea was mostly ignored for five years.

Why it mattered. It was the seed of diffusion, the method behind most video models today.

2016

September 2016

The first tiny AI videos

Carl Vondrick, Hamed Pirsiavash and Antonio TorralbaMIT, Cambridge, United States

Trained on over two million unlabelled videos, VGAN generated clips of 32 frames at 64 by 64 pixels, up to about a second long. It learned to draw a still background and a moving foreground separately. The paper appeared at NeurIPS 2016.

Why it mattered. It showed a network could learn how scenes move, not just how they look.

2017

late 2017

The word "deepfake"

An anonymous Reddit userOnline

A user called "deepfakes" posted videos in which one person's face had been swapped onto another's using neural networks, and shared the code. The name stuck for any AI-faked video. Most early deepfakes targeted women without their consent.

Why it mattered. Faked video became a public worry years before text-to-video models existed.

2020 – 2022

From noise to pictures

Diffusion turns static into images, works in a compressed latent space, follows prompts and starts making short videos.

2020

June 2020

Diffusion finally works

Jonathan Ho, Ajay Jain and Pieter AbbeelUC Berkeley, United States

Denoising diffusion probabilistic models (DDPM) trained a network to predict the noise hidden in a picture, then removed it step by step, starting from pure static. The images matched the best GANs, and training was much more stable.

Why it mattered. Almost every image and video generator since is built on this recipe.

2021

22 February 2021

A label that travels with the file

Adobe, Arm, BBC, Intel, Microsoft and TruepicUnited States and United Kingdom

Six organisations formed the Coalition for Content Provenance and Authenticity (C2PA) to write an open standard for "content credentials": signed records, attached to a file, of who made it and how it was edited. Many AI tools later used it to mark their output.

Why it mattered. It gave the world a shared way to say where a picture or video came from.

2021

December 2021

Do the work in a smaller space

Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn OmmerLMU Munich and Heidelberg University, Germany

Instead of denoising every pixel, latent diffusion first squeezes an image with an autoencoder into a much smaller grid of numbers, the "latent". Diffusion happens there, and a decoder turns the result back into pixels. This cut the cost dramatically.

Why it mattered. Video models compress in the same way, in both space and time, or they could never afford to run.

2021

December 2021

Turning up the prompt

Jonathan Ho and Tim SalimansGoogle, United States

Classifier-free guidance trains one model both with and without the text prompt. At each step it pushes the picture further in the direction the prompt suggests. First shown at a NeurIPS workshop, it was posted on arXiv in July 2022.

Why it mattered. It is the "follow the prompt more closely" dial inside nearly every text-to-image and text-to-video model.

2022

April 2022

Diffusion for video

Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi and David FleetGoogle, United States

Video Diffusion Models denoised all the frames of a short clip together, so each frame could "see" its neighbours. This kept objects steadier from frame to frame than making each frame separately.

Why it mattered. Denoising frames together is the main trick for keeping video consistent.

2022

22 August 2022

Stable Diffusion is released openly

Stability AI, CompVis at LMU Munich and RunwayMunich, Germany and London, United Kingdom

A latent diffusion model trained on billions of captioned images was released with its weights, so anyone with a gaming graphics card could run it at home. Thousands of tools and many later open video models were built on it.

Why it mattered. Image generation went from lab demo to something millions of people could use and study.

2022

29 September 2022

Make-A-Video

Uriel Singer, Adam Polyak, Devi Parikh, Yaniv Taigman and colleagues at Meta AITel Aviv, Israel and Menlo Park, United States

Meta's team took a text-to-image model and taught it motion from videos that had no captions at all. It produced short, low-resolution clips from a sentence.

Why it mattered. It showed video models could borrow what image models already knew about the world.

2022

5 October 2022

Imagen Video

Jonathan Ho, William Chan, Chitwan Saharia and colleagues at GoogleMountain View, United States

A chain of seven diffusion models first made a small, jerky clip, then filled in extra frames and added detail. The paper reports 128 frames at 1280 by 768 pixels and 24 frames per second, about 5.3 seconds.

Why it mattered. It reached HD resolution by building a clip up in stages.

2022

December 2022

Diffusion meets the transformer

William Peebles and Saining XieUC Berkeley and New York University, United States

The diffusion transformer (DiT) cut the latent into small patches and processed them with a transformer, the same design behind chatbots. Bigger transformers gave better pictures, predictably.

Why it mattered. This design, scaled up to patches of space and time, became the backbone of Sora and many later video models.

By the numbers

Longest clips reported by notable video models

From about one second in 2016 to minutes in 2024. Figures for company models are as reported by the companies.

1 hour1 min1 s 2020 2016: VGAN: 32 frames, up to about 1 second20162022: Imagen Video: 128 frames at 24 fps, about 5.3 seconds20222023: Runway Gen-2: about 4 seconds (reported)20232024: Sora: up to 1 minute (reported)20242024: Kling: up to 2 minutes (reported)
  1. 2016 VGAN: 32 frames, up to about 1 second
  2. 2022 Imagen Video: 128 frames at 24 fps, about 5.3 seconds
  3. 2023 Runway Gen-2: about 4 seconds (reported)
  4. 2024 Sora: up to 1 minute (reported)
  5. 2024 Kling: up to 2 minutes (reported)

2023 – today

Video for everyone

Text-to-video tools reach the public, clips get longer, sharper and gain sound, and their cost becomes clear.

2023

20 March 2023

Text-to-video anyone can try

RunwayNew York, United States

Runway announced Gen-2, which made short clips of about four seconds from a text prompt or an image. It opened to the public in stages over the following months and was widely described as the first publicly available text-to-video model.

Why it mattered. Artists and students began experimenting with AI video themselves.

2024

15 February 2024

Sora: up to a minute of video

OpenAISan Francisco, United States

OpenAI showed Sora, which it reported could make videos up to one minute long. Its technical report described compressing video into a latent space and cutting it into "spacetime patches" for a diffusion transformer. It was not released to the public at first.

Why it mattered. It showed how far scaling a diffusion transformer on video could go, and set off a race.

2024

14 May 2024

Veo, with watermarks built in

Google DeepMindMountain View, United States

Google announced Veo, which it reported could make 1080p videos beyond a minute long. Its videos were to carry SynthID watermarks. Veo 2 followed in December 2024.

Why it mattered. Big labs began shipping watermarking alongside video models.

2024

6 June 2024

Kling from China

KuaishouBeijing, China

Kuaishou, a short-video company, launched Kling for testing in China. It reported clips up to two minutes long at 1080p and 30 frames per second. Kling soon opened to users around the world.

Why it mattered. Strong video models were now coming from several countries, not one lab.

2024

December 2024

Sora opens, and open models catch up

OpenAI; TencentSan Francisco, United States and Shenzhen, China

On 9 December OpenAI opened Sora to paying users in some countries, with clips up to 1080p and 20 seconds. The same month Tencent released HunyuanVideo, a 13-billion-parameter model with openly published weights, after Tsinghua University and Zhipu AI's open CogVideoX in August.

Why it mattered. Anyone could now study how a large video model is built, not just use one.

2025

20 May 2025

Pictures and sound in one go

Google DeepMindMountain View, United States

Google announced Veo 3, which generated speech, sound effects and music together with the video in a single model. Widely shared clips of talking characters made it hard to tell AI video from real footage at a glance.

Why it mattered. Realistic talking video made labels and provenance more urgent.

2025

May and September 2025

Counting the energy of a video

MIT Technology Review; Hugging Face researchersUnited States and Canada

Measurements on one open model found a 5-second clip used about 3.4 million joules, more than 700 times a high-quality image. A Hugging Face study found that doubling a video's length roughly quadruples the energy needed.

Why it mattered. Video is the most energy-hungry thing most people ask AI to make.

2025

30 September 2025

Sora 2 and an app

OpenAISan Francisco, United States

OpenAI released Sora 2, which added synchronised speech and sound, in a social app where people could put themselves into videos. It spread quickly, and so did complaints about copyright and look-alike videos of real people.

Why it mattered. AI video became something people shared in a feed, raising new questions about consent.

2026

24 March 2026

Sora is switched off

OpenAISan Francisco, United States

OpenAI announced it would close the Sora app, which shut on 26 April 2026, with the API ending in September. News reports pointed to very high running costs and problems with copyright and deepfakes.

Why it mattered. Making video is expensive, and even a famous product could not cover its costs.

By the numbers

Pixels in one generated clip

Width × height × frames. It grew more than 50,000-fold in eight years, which is why video models must compress before they generate.

100,0001,000,00010,000,000100,000,0001,000,000,00010,000,000,000 2020 2016: VGAN: 64 × 64 pixels × 32 frames20162022: Imagen Video: 1280 × 768 × 128 frames20222024: Kling: 1920 × 1080 × 3,600 frames (2 minutes at 30 fps, reported maximum)2024
  1. 2016 VGAN: 64 × 64 pixels × 32 frames
  2. 2022 Imagen Video: 1280 × 768 × 128 frames
  3. 2024 Kling: 1920 × 1080 × 3,600 frames (2 minutes at 30 fps, reported maximum)

2023 – today

Labels and laws

Watermarks, content credentials and new laws in India, China and Europe try to make AI video easy to recognise.

2023

29 August 2023

An invisible watermark

Google DeepMindLondon, United Kingdom

SynthID hid a signal inside AI-made images from Google's Imagen model. People cannot see it, but a detector can find it, even after cropping or compression. In May 2024 Google said it would extend SynthID to video made by Veo.

Why it mattered. Watermarks inside the pixels can survive where a visible label is easily cropped off.

2023

26 December 2023

India warns platforms about deepfakes

Ministry of Electronics and Information Technology (MeitY)New Delhi, India

In November 2023 a deepfake video of a well-known actor spread widely, and ministers called deepfakes a dangerous form of misinformation. In December MeitY advised all platforms to tell users clearly, at sign-up and when they upload, that impersonation and deceptive content are banned under the IT Rules.

Why it mattered. India began treating deepfakes as a problem platforms must actively deal with.

2024

1 and 15 March 2024

Labels for AI-made content in India

Ministry of Electronics and Information Technology (MeitY)New Delhi, India

A 1 March advisory asked platforms to seek government permission before offering untested AI models, and to mark AI-made content that could be used as a deepfake with a permanent unique identifier or metadata. After criticism, a 15 March advisory dropped the permission step but kept the call for labels.

Why it mattered. It was India's first push for AI content labels, a step towards later rules.

2025

1 September 2025

China requires AI labels

Cyberspace Administration of China and three other agenciesBeijing, China

China's labelling measures took effect, requiring visible labels on AI-made content such as video, plus hidden labels in the file's data. In Europe, the AI Act's Article 50 requires AI-made deepfakes to be disclosed, with those rules applying from August 2026.

Why it mattered. Several large markets moved from asking to requiring AI labels.

2025

22 October 2025

India drafts rules for synthetic media

Ministry of Electronics and Information Technology (MeitY)New Delhi, India

MeitY published draft amendments to the IT Rules defining "synthetically generated information". The draft proposed that visible labels cover at least 10% of the screen, or play during the first 10% of an audio clip, and asked the public for comments.

Why it mattered. It was the first draft Indian law written specifically about AI-made media.

2026

10 February 2026

India's deepfake rules become law

Ministry of Electronics and Information Technology (MeitY)New Delhi, India

The amended IT Rules, in force from 20 February 2026, require AI-made audio and video to carry a prominent label and metadata that platforms must not strip. The fixed 10% size was dropped. Platforms must remove unlawful content within 3 hours of a lawful order, and non-consensual intimate images, including deepfakes, within 2 hours.

Why it mattered. India now has some of the world's strictest labelling and takedown rules for AI video.

Did you know?

At 24 frames per second, one minute of film is 1,440 separate pictures.

One open video model used about 3.4 million joules to make a 5-second clip, more than 700 times the energy of a high-quality AI image, MIT Technology Review found in 2025.

A 2025 Hugging Face study found that doubling the length of an AI video roughly quadruples the energy it takes to make.

HunyuanVideo's autoencoder shrinks each frame 8 times in width and height and 4 times in time, so the model works on about 48 times fewer numbers than the raw video.

The word "deepfake" comes from the username of the Reddit account that posted face-swapped videos in 2017.

The people

Who figured it out

Eadweard Muybridge

1830 – 1904 · Photographer · England and United States

Photographed a galloping horse with a row of cameras in 1878, showing motion as a sequence of stills.

Ian Goodfellow

· Computer scientist · United States

Invented generative adversarial networks in 2014, which powered early image generators, video GANs and deepfakes.

Jascha Sohl-Dickstein

· Physicist and computer scientist · United States

Led the 2015 paper that first used diffusion, adding and removing noise, to generate images.

Jonathan Ho

· Computer scientist · United States

Lead author of DDPM (2020), classifier-free guidance and the first video diffusion papers.

Robin Rombach

· Computer scientist · Germany

Lead author of latent diffusion, the compressed-space method behind Stable Diffusion and most video models.

Saining Xie

· Computer scientist, New York University · China and United States

Co-created the diffusion transformer (DiT) with William Peebles, the design later scaled up for video.

Devi Parikh

· Computer scientist, Meta AI · United States

A research director on Make-A-Video (2022), one of the first text-to-video diffusion systems.

Sasha Luccioni

· AI and climate researcher, Hugging Face · Canada

Co-wrote the 2025 study showing that doubling an AI video's length roughly quadruples its energy use.

Where it happened

11 places, one idea

Sources

Where this comes from

Dates marked “c.” are approximate, and historians sometimes disagree about who was first. If you spot a mistake, tell us.

  1. Eadweard Muybridge Wikipedia
  2. The Horse in Motion Wikipedia
  3. Frame rate Wikipedia
  4. Auto-Encoding Variational Bayes (Kingma & Welling, 2013) arXiv
  5. Generative Adversarial Networks (Goodfellow et al., 2014) arXiv
  6. Deep Unsupervised Learning using Nonequilibrium Thermodynamics (Sohl-Dickstein et al., 2015) arXiv
  7. Generating Videos with Scene Dynamics (Vondrick, Pirsiavash & Torralba, 2016) arXiv / NeurIPS
  8. Deepfake Wikipedia
  9. Denoising Diffusion Probabilistic Models (Ho, Jain & Abbeel, 2020) arXiv
  10. C2PA founding press release Coalition for Content Provenance and Authenticity
  11. Technology and media entities join forces to create standards group aimed at building trust in online content Microsoft
  12. High-Resolution Image Synthesis with Latent Diffusion Models (Rombach et al., 2021) arXiv
  13. Classifier-Free Diffusion Guidance (Ho & Salimans, 2022) arXiv
  14. Video Diffusion Models (Ho et al., 2022) arXiv
  15. Stable Diffusion public release Stability AI
  16. Make-A-Video: Text-to-Video Generation without Text-Video Data (Singer et al., 2022) arXiv
  17. Imagen Video: High Definition Video Generation with Diffusion Models (Ho et al., 2022) arXiv
  18. Scalable Diffusion Models with Transformers (Peebles & Xie, 2022) arXiv
  19. Gen-2: Generate novel videos with text, images or video clips Runway
  20. Runway Gen-2 is the First Publicly Available Text-to-Video Generator PetaPixel
  21. Identifying AI-generated images with SynthID Google DeepMind
  22. MoS Rajeev Chandrasekhar on deepfake video featuring an actor Business Today
  23. MeitY issues advisory to all intermediaries to comply with existing IT rules Press Information Bureau, Government of India
  24. MeitY issues advisory against AI-Deepfakes on social media and other platforms SCC Online
  25. MeitY issues advisory on Misinformation and Deepfake mandating unique metadata SCC Online
  26. MeitY Revises AI Advisory, Does Away with Government Permission Requirement AZB & Partners
  27. Video generation models as world simulators OpenAI
  28. Sora (text-to-video model) Wikipedia
  29. Google I/O 2024: Introducing Veo and Imagen 3 Google
  30. Watermarking AI-generated text and video with SynthID Google DeepMind
  31. Kuaishou Unveils Proprietary Video Generation Model 'Kling' Kuaishou Technology
  32. What to know about this new Chinese text-to-video AI model MIT Technology Review
  33. Movie Gen: A Cast of Media Foundation Models Meta AI
  34. Sora is here OpenAI
  35. Veo (text-to-video model) Wikipedia
  36. Sora 2 is here OpenAI
  37. Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems EU Artificial Intelligence Act (Future of Life Institute)
  38. Measures for Labeling of AI-Generated Synthetic Content China Law Translate
  39. Explanatory note: draft amendments on synthetically generated information (22 October 2025) Ministry of Electronics and Information Technology, Government of India
  40. Unpacking the Draft IT Rules on Synthetic Information MediaNama
  41. We did the math on AI's energy footprint MIT Technology Review
  42. Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models (Delavande, Pierrard & Luccioni, 2025) arXiv
  43. Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Amendment Rules, 2026 Ministry of Electronics and Information Technology, Government of India
  44. Government sets 3-hour deadline for social media platforms to take down AI-generated deepfake content All India Radio News
  45. MeitY notifies the IT Amendment Rules 2026 Khaitan & Co
  46. OpenAI's Sora app is shutting down TechCrunch
  47. HunyuanVideo: A Systematic Framework For Large Video Generative Models (Tencent, 2024) arXiv
  48. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer (Yang et al., 2024) arXiv
  49. What to know about the Sora discontinuation OpenAI Help Center
  50. Devi Parikh Wikipedia

That's the history. Now see how it works.