The history

The history of AI inference

From motor-driven knobs in 1960 to data centres answering billions of questions a day: how we learned to run trained models fast, small and cheap.

For most of AI's history, the hard part was teaching a model. Running it seemed easy. But once models became useful, they were asked questions millions, then billions, of times. Engineers built GPUs, TPUs and phone chips made of multiply-add units, learned to shrink models to 8 and 4 bits, and invented new ways to serve chatbots, while the electricity bill became a public question.

65
years
30
moments
8
people
7
places

1960

Neural network weights as hardware knobs

Frank Rosenblatt, Mark I Perceptron, United States

1980

Expert system in daily business use

XCON at Digital Equipment Corporation, United States

2009

Large neural networks on GPUs

Raina, Madhavan and Ng, Stanford

2015 (revealed 2016)

Data-centre chip built for neural-network inference

Google TPU, United States

2017

Neural engines in phones

Huawei Kirin 970 and Apple A11 Bionic

2023

4-bit LLM running on a laptop, open source

Georgi Gerganov, llama.cpp, Bulgaria

23 June 1960Rules and first machines

1958 – 1999

Rules and first machines

Hardware perceptrons, expert systems that run fixed rules, and the first networks reading cheques for banks.

1960

23 June 1960

A neural network built in hardware

Frank Rosenblatt and teamCornell Aeronautical Laboratory, Buffalo, United States

The Mark I Perceptron was shown to the public. Its weights were motor-driven knobs (potentiometers), and a camera of 400 light cells fed it pictures. Once set, the knobs turned inputs into an answer with no computer program at all.

Why it mattered. It showed that a network's answer is just inputs times weights, added up: something you can build in hardware.

1965

Dendral: rules that reason like a chemist

Edward Feigenbaum, Joshua Lederberg, Bruce Buchanan and Carl DjerassiStanford University, United States

Dendral used rules written by chemists to guess which molecule matched a mass-spectrometry reading. It is often called the first expert system.

Why it mattered. It started the idea of putting expert knowledge into a program that answers new questions.

1974

c. 1972–1976

MYCIN suggests antibiotics

Edward Shortliffe and colleaguesStanford University, United States

MYCIN held about 600 if-then rules about blood infections. For each patient it chained through them to suggest likely bacteria and a treatment, with a confidence score. In tests its advice compared well with specialists, but it was never used on real patients.

Why it mattered. Every question meant running the same fixed rules again: an early form of inference.

1980

XCON: expert rules at work in a factory

John McDermott (Carnegie Mellon University) for Digital Equipment CorporationSalem, New Hampshire, United States

XCON (first called R1) used thousands of rules to check and complete orders for VAX computers. By 1986 it had handled 80,000 orders with 95 to 98% accuracy, and was estimated to save DEC about $25 million a year.

Why it mattered. One of the first AI systems answering real questions, every day, for a business.

1998

1990s

A neural network reads cheques

Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Bell Labs and NCR, United States

Convolutional networks for reading handwritten digits were built into cheque-reading machines. By the late 1990s such systems read several million cheques a day, according to the team's 1998 paper.

Why it mattered. A trained network answering questions at industrial scale, long before chatbots.

1989 – 2019

Smaller, faster, on the phone

Pruning, 8-bit numbers, distillation and compression make models small enough for phones.

1989

November 1989

Optimal Brain Damage: pruning a network

Yann LeCun, John Denker and Sara SollaAT&T Bell Laboratories, Holmdel, United States

The team showed that many weights in a trained network can be removed with little loss, if you choose them by how much each one matters.

Why it mattered. Pruning became one of the main ways to make models smaller and cheaper to run.

2011

December 2011

8-bit numbers are enough

Vincent Vanhoucke, Andrew Senior and Mark MaoGoogle, Mountain View, United States

For speech recognition, the team ran a neural network with 8-bit integers on ordinary CPUs and got about a 3 times speed-up over a well-tuned floating-point version, with no real loss in accuracy.

Why it mattered. An early, clear sign that inference does not need full-precision numbers.

2015

March 2015

Distillation: a small student copies a big teacher

Geoffrey Hinton, Oriol Vinyals and Jeff DeanGoogle, Mountain View, United States

The paper showed how to train a small model to copy the full answer probabilities of a big one, keeping much of its skill in a model that is far cheaper to run.

Why it mattered. Distillation is now a standard way to make models small enough for phones and busy servers.

2015

October 2015

Deep Compression: 35 times smaller

Song Han, Huizi Mao and William DallyStanford University, United States

Pruning, quantisation and clever coding together shrank the AlexNet network from about 240 MB to about 6.9 MB with no loss of accuracy.

Why it mattered. It showed big networks could fit in a phone app, and won a best paper award at ICLR 2016.

2017

April 2017

Your keyboard learns without sending your words

Brendan McMahan, Daniel Ramage and colleaguesGoogle, Seattle and Mountain View, United States

Google tested federated learning in its Gboard keyboard: phones improve a small shared prediction model locally and send only updates, not what you typed. The keyboard's next-word model itself runs on the phone.

Why it mattered. A clear example of useful, private, on-device inference.

2017

December 2017

INT8 inference goes into production tools

Benoit Jacob and colleaguesGoogle, United States

The team described how to run networks using only 8-bit integer arithmetic, and how to train them so accuracy holds up. The method went into TensorFlow Lite, used to run models on phones.

Why it mattered. INT8 became the everyday standard for running AI on phones and many servers.

2019

March 2019

Speech recognition that fits in a phone

Google speech teamGoogle, United States

Google rolled out an all-neural speech recogniser that runs entirely on Pixel phones. Quantisation shrank it from 450 MB to 80 MB and made it about four times faster.

Why it mattered. Voice typing could work offline and privately, with no round trip to a server.

2009 – 2014

GPUs wake up deep learning

Graphics chips, built for games, turn out to be perfect for the multiply-adds of neural networks.

2009

June 2009

GPUs make deep learning practical

Rajat Raina, Anand Madhavan and Andrew NgStanford University, United States

The team ran large neural networks on graphics chips (GPUs) meant for video games, and reported training up to about 70 times faster than on a dual-core CPU. A job of weeks became about a day.

Why it mattered. The same parallel multiply-adds later made GPUs the main engine for AI inference too.

2012

September 2012

AlexNet wins on two gaming GPUs

Alex Krizhevsky, Ilya Sutskever and Geoffrey HintonUniversity of Toronto, Canada

A deep network trained on two NVIDIA GTX 580 graphics cards won the ImageNet image-recognition contest by a wide margin.

Why it mattered. It set off the deep learning boom, and with it the demand for chips that run networks fast.

2015 – 2019

Chips and tools built for inference

TPUs, inference engines, a shared model format and neural engines in phones.

2016

May 2016 (announced); in use since 2015

TPU: a chip built for inference

Norman Jouppi and a Google teamGoogle, Mountain View, United States

Google revealed its Tensor Processing Unit, already running in its data centres since 2015. Its heart was a grid of 65,536 8-bit multiply-add units, reaching 92 trillion operations a second. Google reported 15 to 30 times the speed of the CPUs and GPUs it then used, for its inference jobs.

Why it mattered. It proved that a chip made mostly of multiply-add units, using small numbers, is the efficient way to run trained networks.

2016

13 September 2016 (renamed)

TensorRT tunes models for GPUs

NVIDIASanta Clara, United States

NVIDIA's GPU Inference Engine, renamed TensorRT in September 2016, took a trained network and rebuilt it to run fast: merging layers, choosing the best maths routines and using lower precision such as INT8.

Why it mattered. Inference became its own engineering job, with its own tools, separate from training.

2017

September 2017

ONNX: one file format for trained models

Microsoft and Facebook, later joined by AWS and othersRedmond and Menlo Park, United States

ONNX let a model trained in one framework be saved and run in another. In December 2018 Microsoft open-sourced ONNX Runtime, an engine for running such models on many kinds of hardware.

Why it mattered. Training tools and inference engines could now be mixed and matched.

2017

12 September 2017

A neural engine in a phone

Apple (A11 Bionic); Huawei had announced its Kirin 970 with an NPU days earlierCupertino, United States; Shenzhen, China

The iPhone X's A11 chip included a two-core Neural Engine that Apple said could do 600 billion operations a second, used for Face ID and Animoji. Huawei's Kirin 970, shown at IFA Berlin that month, also had a neural processing unit.

Why it mattered. Phones got their own inference hardware, so AI could run without the cloud.

By the numbers

Memory bandwidth of NVIDIA's flagship data-centre GPUs

Bandwidth grew about 30-fold in 12 years. For LLM decoding, this number, more than raw maths speed, sets the tokens per second.

1001,00010,000 20152020 2012: Tesla K20X: 250 GB/s20122016: Tesla P100: 732 GB/s (first with HBM2)20162017: Tesla V100: 900 GB/s20172020: A100 40GB: 1,555 GB/s20202021: A100 80GB: 2,039 GB/s20212022: H100 SXM: 3,350 GB/s20222024: H200: 4,800 GB/s20242024: B200: 8,000 GB/s
  1. 2012 Tesla K20X: 250 GB/s
  2. 2016 Tesla P100: 732 GB/s (first with HBM2)
  3. 2017 Tesla V100: 900 GB/s
  4. 2020 A100 40GB: 1,555 GB/s
  5. 2021 A100 80GB: 2,039 GB/s
  6. 2022 H100 SXM: 3,350 GB/s
  7. 2024 H200: 4,800 GB/s
  8. 2024 B200: 8,000 GB/s

2022 – 2025

Serving language models to the world

4-bit LLMs on laptops, new ways to batch and cache, and a public reckoning with energy.

2022

May 2022

FlashAttention: fewer trips to memory

Tri Dao, Daniel Fu, Stefano Ermon, Atri Rudra and Christopher RéStanford University, United States

FlashAttention computed attention in small tiles that stay in the GPU's fast on-chip memory, instead of writing big tables out to slower memory. The answers are exactly the same, only faster.

Why it mattered. It showed that for AI chips, moving data costs more than doing the maths.

2022

July 2022

Continuous batching for chatbots

Gyeong-In Yu, Joo Seong Jeong and colleaguesSeoul National University and FriendliAI, South Korea

Orca let a server add new requests to a running batch after every token, instead of waiting for the whole batch to finish. Short and long answers no longer held each other up.

Why it mattered. Almost every LLM server now batches this way, sharing each read of the weights across many users.

2022

August–October 2022

Big language models squeezed to 8 and 4 bits

Tim Dettmers and colleagues (LLM.int8()); Elias Frantar and colleagues (GPTQ)University of Washington, United States; IST Austria

LLM.int8() ran models with 175 billion parameters in 8-bit with no loss in quality. Weeks later GPTQ showed weights could go down to 3 or 4 bits with small losses.

Why it mattered. Huge models could suddenly fit on far fewer GPUs, and smaller ones on laptops.

2022

4 July 2022

Bhashini: AI for India's languages

Government of India (Ministry of Electronics and IT)Gandhinagar, India

The Prime Minister launched Digital India Bhashini at Digital India Week 2022, a platform for translation and speech tools in Indian languages, open to apps and services.

Why it mattered. Serving AI in many languages, cheaply and at scale, became national infrastructure.

2023

March 2023

A chatbot model on a laptop

Georgi GerganovSofia, Bulgaria

Gerganov released llama.cpp, plain C/C++ code that ran Meta's LLaMA models on an ordinary MacBook using 4-bit weights. Thousands of people began running LLMs on their own computers.

Why it mattered. It made on-device LLM inference practical for anyone, and pushed quantisation into everyday use.

2023

June 2023 (released); October 2023 (paper)

vLLM and PagedAttention

Woosuk Kwon, Zhuohan Li, Ion Stoica and colleaguesUC Berkeley, United States

The team found that LLM servers wasted 60 to 80% of their KV-cache memory. PagedAttention stores the cache in small pages, like an operating system handles memory, so many more users fit on each GPU. They reported 2 to 4 times the throughput of earlier systems.

Why it mattered. Their open-source vLLM became one of the most widely used ways to serve LLMs.

2023

November 2022 (preprint); July 2023 (ICML)

Speculative decoding: guess ahead, check in one go

Yaniv Leviathan, Matan Kalman and Yossi Matias (Google); independently Charlie Chen and colleagues (DeepMind)Google Research, Israel; DeepMind, London

A small, fast model drafts several tokens; the big model checks them all in a single pass and keeps the ones it agrees with. The output is exactly what the big model alone would give, typically 2 to 3 times faster.

Why it mattered. It beats the one-token-at-a-time memory wall without changing the answers.

2023

6 September 2023

Hello! UPI: pay by speaking

NPCI, with AI4Bharat at IIT Madras and BhashiniMumbai, India

At the Global Fintech Fest, NPCI launched voice-based UPI payments in Hindi and English, with language models co-developed with AI4Bharat at IIT Madras under Bhashini. UPI itself handled over 20 billion transactions in August 2025.

Why it mattered. A glimpse of AI inference at the scale of a whole country's payments.

2023

November 2023

Measuring the energy of each answer

Sasha Luccioni, Yacine Jernite and Emma StrubellHugging Face and Carnegie Mellon University

The study measured the energy of 1,000 inferences across many tasks and models. Generating images used far more energy than classifying text, and big general-purpose models used far more than small task-specific ones.

Why it mattered. It put inference, not just training, at the centre of the AI energy debate.

2023

December 2023

An LLM built into a phone

GoogleMountain View, United States

Google put Gemini Nano, a small language model, onto the Pixel 8 Pro to run features like summarising recordings and smart replies on the device. Apple announced its own on-device model, about 3 billion parameters, in June 2024.

Why it mattered. Language models joined the other small models already running inside phones.

2025

10 April 2025

The IEA counts AI's electricity

International Energy AgencyParis, France

The IEA's Energy and AI report estimated that data centres used about 415 TWh in 2024, around 1.5% of the world's electricity, and could reach about 945 TWh by 2030, with AI the main driver.

Why it mattered. It gave governments a shared, careful baseline for the energy cost of AI.

2025

21 August 2025

A company reports energy per prompt

GoogleMountain View, United States

Google reported that a median text prompt to its Gemini app used about 0.24 Wh of energy and 0.26 mL of water, and said the figure had fallen 33-fold in a year. Epoch AI had earlier estimated about 0.3 Wh for a typical GPT-4o query.

Why it mattered. Per-answer figures finally became public, though they are company numbers that are hard to check.

Did you know?

For a language model, each new token takes about 2 floating-point operations per parameter: an 8-billion-parameter model needs about 16 billion for every token.

Google's first TPU had 65,536 multiply-add units working in lockstep, each on 8-bit numbers.

LLM servers were once found to waste 60 to 80% of the memory set aside for the KV cache; paging it like an operating system fixed most of that.

Google reported that the energy for a median Gemini text prompt fell 33-fold in the year to mid-2025.

Data centres used about 1.5% of the world's electricity in 2024, according to the IEA.

The people

Who figured it out

Frank Rosenblatt

1928 – 1971 · Psychologist and engineer · United States

Built the Mark I Perceptron, a neural network whose weights were motor-driven knobs.

Edward Shortliffe

born 1947 · Physician and computer scientist · Canada / United States

Created MYCIN, a rule-based system that suggested antibiotics.

Norman Jouppi

— · Computer architect · United States

Led the design of Google's first Tensor Processing Unit for inference.

Song Han

— · Computer scientist · China / United States

Showed with Deep Compression that networks can shrink 35-fold with pruning and quantisation.

Sara Hooker

— · Computer scientist · United States / Canada

Argued in 'The Hardware Lottery' (2020) that the ideas that win in AI are shaped by the chips available.

Tri Dao

— · Computer scientist · Vietnam / United States

Co-created FlashAttention, which speeds up attention by cutting trips to memory.

Georgi Gerganov

— · Software engineer · Bulgaria

Wrote llama.cpp, which put quantised LLMs on ordinary laptops and phones.

Sasha Luccioni

— · AI researcher · Canada

Measured the energy and carbon of AI inference across tasks and models.

Where it happened

7 places, one idea

Sources

Where this comes from

Dates marked “c.” are approximate, and historians sometimes disagree about who was first. If you spot a mistake, tell us.

  1. Mark I Perceptron Wikipedia
  2. Professor's perceptron paved the way for AI – 60 years too soon Cornell Chronicle (Cornell University)
  3. Dendral Wikipedia
  4. Mycin Wikipedia
  5. Rule-Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project (Buchanan & Shortliffe, 1984) Stanford University / Addison-Wesley
  6. Xcon Wikipedia
  7. Optimal Brain Damage (LeCun, Denker & Solla, NIPS 1989) NeurIPS Proceedings
  8. Gradient-Based Learning Applied to Document Recognition (LeCun et al., 1998) Proceedings of the IEEE
  9. LeNet Wikipedia
  10. Large-scale deep unsupervised learning using graphics processors (Raina, Madhavan & Ng, ICML 2009) ACM Digital Library
  11. Improving the speed of neural networks on CPUs (Vanhoucke, Senior & Mao, 2011) Google Research
  12. ImageNet Classification with Deep Convolutional Neural Networks (Krizhevsky, Sutskever & Hinton, 2012) NeurIPS Proceedings
  13. AlexNet Wikipedia
  14. Distilling the Knowledge in a Neural Network (Hinton, Vinyals & Dean, 2015) arXiv
  15. Deep Compression (Han, Mao & Dally, 2015) arXiv
  16. In-Datacenter Performance Analysis of a Tensor Processing Unit (Jouppi et al., ISCA 2017) arXiv
  17. Tensor Processing Unit Wikipedia
  18. Production Deep Learning with NVIDIA GPU Inference Engine NVIDIA Technical Blog
  19. TensorRT Wikipedia
  20. ONNX V1 released Engineering at Meta
  21. Open Neural Network Exchange Wikipedia
  22. ONNX Runtime is now open source Microsoft Azure Blog
  23. Apple A11 Wikipedia
  24. Apple unveils A11 Bionic neural engine chip in iPhone X CNBC
  25. HiSilicon (Kirin 970) Wikipedia
  26. Federated Learning: Collaborative Machine Learning without Centralized Training Data Google Research Blog
  27. Federated Learning for Mobile Keyboard Prediction (Hard et al., 2018) arXiv
  28. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (Jacob et al., 2017) arXiv
  29. An All-Neural On-Device Speech Recognizer Google Research Blog
  30. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al., 2022) arXiv
  31. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar et al., 2022) arXiv
  32. FlashAttention (Dao et al., 2022) arXiv
  33. Orca: A Distributed Serving System for Transformer-Based Generative Models (Yu et al., OSDI 2022) USENIX
  34. Fast Inference from Transformers via Speculative Decoding (Leviathan, Kalman & Matias, ICML 2023) PMLR
  35. Accelerating Large Language Model Decoding with Speculative Sampling (Chen et al., 2023) arXiv
  36. llama.cpp GitHub (ggml-org)
  37. llama.cpp Wikipedia
  38. Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., SOSP 2023) arXiv
  39. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention vLLM Blog
  40. PM inaugurates Digital India Week 2022 in Gandhinagar Press Information Bureau, Government of India
  41. NPCI launches new products; users can now make voice UPI payments YourStory
  42. UPI transactions cross 20 billion mark for the first time in August All India Radio News
  43. Gemini (language model) Wikipedia
  44. Power Hungry Processing: Watts Driving the Cost of AI Deployment? (Luccioni, Jernite & Strubell) arXiv
  45. Energy and AI: Executive summary International Energy Agency
  46. Measuring the environmental impact of AI inference Google Cloud Blog
  47. Measuring the environmental impact of delivering AI at Google Scale arXiv
  48. How much energy does ChatGPT use? Epoch AI
  49. NVIDIA H100 Tensor Core GPU NVIDIA
  50. NVIDIA H200 Tensor Core GPU NVIDIA
  51. NVIDIA A100 Tensor Core GPU NVIDIA
  52. Nvidia Tesla / data-centre GPU specifications Wikipedia
  53. NVIDIA GB200 NVL72 NVIDIA
  54. Scaling Laws for Neural Language Models (Kaplan et al., 2020) arXiv
  55. The Hardware Lottery (Sara Hooker, 2020) arXiv
  56. Introducing Apple's On-Device and Server Foundation Models Apple Machine Learning Research

That's the history. Now see how it works.