From motor-driven knobs in 1960 to data centres answering billions of questions a day: how we learned to run trained models fast, small and cheap.
For most of AI's history, the hard part was teaching a model. Running it seemed easy. But once models became useful, they were asked questions millions, then billions, of times. Engineers built GPUs, TPUs and phone chips made of multiply-add units, learned to shrink models to 8 and 4 bits, and invented new ways to serve chatbots, while the electricity bill became a public question.
Frank Rosenblatt, Mark I Perceptron, United States
1980
Expert system in daily business use
XCON at Digital Equipment Corporation, United States
2009
Large neural networks on GPUs
Raina, Madhavan and Ng, Stanford
2015 (revealed 2016)
Data-centre chip built for neural-network inference
Google TPU, United States
2017
Neural engines in phones
Huawei Kirin 970 and Apple A11 Bionic
2023
4-bit LLM running on a laptop, open source
Georgi Gerganov, llama.cpp, Bulgaria
23 June 1960Rules and first machines
1958 – 1999
Rules and first machines
Hardware perceptrons, expert systems that run fixed rules, and the first networks reading cheques for banks.
1960
23 June 1960
A neural network built in hardware
Frank Rosenblatt and teamCornell Aeronautical Laboratory, Buffalo, United States
The Mark I Perceptron was shown to the public. Its weights were motor-driven knobs (potentiometers), and a camera of 400 light cells fed it pictures. Once set, the knobs turned inputs into an answer with no computer program at all.
Why it mattered. It showed that a network's answer is just inputs times weights, added up: something you can build in hardware.
Edward Shortliffe and colleaguesStanford University, United States
MYCIN held about 600 if-then rules about blood infections. For each patient it chained through them to suggest likely bacteria and a treatment, with a confidence score. In tests its advice compared well with specialists, but it was never used on real patients.
Why it mattered. Every question meant running the same fixed rules again: an early form of inference.
John McDermott (Carnegie Mellon University) for Digital Equipment CorporationSalem, New Hampshire, United States
XCON (first called R1) used thousands of rules to check and complete orders for VAX computers. By 1986 it had handled 80,000 orders with 95 to 98% accuracy, and was estimated to save DEC about $25 million a year.
Why it mattered. One of the first AI systems answering real questions, every day, for a business.
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Bell Labs and NCR, United States
Convolutional networks for reading handwritten digits were built into cheque-reading machines. By the late 1990s such systems read several million cheques a day, according to the team's 1998 paper.
Why it mattered. A trained network answering questions at industrial scale, long before chatbots.
Vincent Vanhoucke, Andrew Senior and Mark MaoGoogle, Mountain View, United States
For speech recognition, the team ran a neural network with 8-bit integers on ordinary CPUs and got about a 3 times speed-up over a well-tuned floating-point version, with no real loss in accuracy.
Why it mattered. An early, clear sign that inference does not need full-precision numbers.
Distillation: a small student copies a big teacher
Geoffrey Hinton, Oriol Vinyals and Jeff DeanGoogle, Mountain View, United States
The paper showed how to train a small model to copy the full answer probabilities of a big one, keeping much of its skill in a model that is far cheaper to run.
Why it mattered. Distillation is now a standard way to make models small enough for phones and busy servers.
Brendan McMahan, Daniel Ramage and colleaguesGoogle, Seattle and Mountain View, United States
Google tested federated learning in its Gboard keyboard: phones improve a small shared prediction model locally and send only updates, not what you typed. The keyboard's next-word model itself runs on the phone.
Why it mattered. A clear example of useful, private, on-device inference.
The team described how to run networks using only 8-bit integer arithmetic, and how to train them so accuracy holds up. The method went into TensorFlow Lite, used to run models on phones.
Why it mattered. INT8 became the everyday standard for running AI on phones and many servers.
Google rolled out an all-neural speech recogniser that runs entirely on Pixel phones. Quantisation shrank it from 450 MB to 80 MB and made it about four times faster.
Why it mattered. Voice typing could work offline and privately, with no round trip to a server.
Graphics chips, built for games, turn out to be perfect for the multiply-adds of neural networks.
2009
June 2009
GPUs make deep learning practical
Rajat Raina, Anand Madhavan and Andrew NgStanford University, United States
The team ran large neural networks on graphics chips (GPUs) meant for video games, and reported training up to about 70 times faster than on a dual-core CPU. A job of weeks became about a day.
Why it mattered. The same parallel multiply-adds later made GPUs the main engine for AI inference too.
TPUs, inference engines, a shared model format and neural engines in phones.
2016
May 2016 (announced); in use since 2015
TPU: a chip built for inference
Norman Jouppi and a Google teamGoogle, Mountain View, United States
Google revealed its Tensor Processing Unit, already running in its data centres since 2015. Its heart was a grid of 65,536 8-bit multiply-add units, reaching 92 trillion operations a second. Google reported 15 to 30 times the speed of the CPUs and GPUs it then used, for its inference jobs.
Why it mattered. It proved that a chip made mostly of multiply-add units, using small numbers, is the efficient way to run trained networks.
NVIDIA's GPU Inference Engine, renamed TensorRT in September 2016, took a trained network and rebuilt it to run fast: merging layers, choosing the best maths routines and using lower precision such as INT8.
Why it mattered. Inference became its own engineering job, with its own tools, separate from training.
Microsoft and Facebook, later joined by AWS and othersRedmond and Menlo Park, United States
ONNX let a model trained in one framework be saved and run in another. In December 2018 Microsoft open-sourced ONNX Runtime, an engine for running such models on many kinds of hardware.
Why it mattered. Training tools and inference engines could now be mixed and matched.
Apple (A11 Bionic); Huawei had announced its Kirin 970 with an NPU days earlierCupertino, United States; Shenzhen, China
The iPhone X's A11 chip included a two-core Neural Engine that Apple said could do 600 billion operations a second, used for Face ID and Animoji. Huawei's Kirin 970, shown at IFA Berlin that month, also had a neural processing unit.
Why it mattered. Phones got their own inference hardware, so AI could run without the cloud.
Memory bandwidth of NVIDIA's flagship data-centre GPUs
Bandwidth grew about 30-fold in 12 years. For LLM decoding, this number, more than raw maths speed, sets the tokens per second.
2012 Tesla K20X: 250 GB/s
2016 Tesla P100: 732 GB/s (first with HBM2)
2017 Tesla V100: 900 GB/s
2020 A100 40GB: 1,555 GB/s
2021 A100 80GB: 2,039 GB/s
2022 H100 SXM: 3,350 GB/s
2024 H200: 4,800 GB/s
2024 B200: 8,000 GB/s
2022 – 2025
Serving language models to the world
4-bit LLMs on laptops, new ways to batch and cache, and a public reckoning with energy.
2022
May 2022
FlashAttention: fewer trips to memory
Tri Dao, Daniel Fu, Stefano Ermon, Atri Rudra and Christopher RéStanford University, United States
FlashAttention computed attention in small tiles that stay in the GPU's fast on-chip memory, instead of writing big tables out to slower memory. The answers are exactly the same, only faster.
Why it mattered. It showed that for AI chips, moving data costs more than doing the maths.
Gyeong-In Yu, Joo Seong Jeong and colleaguesSeoul National University and FriendliAI, South Korea
Orca let a server add new requests to a running batch after every token, instead of waiting for the whole batch to finish. Short and long answers no longer held each other up.
Why it mattered. Almost every LLM server now batches this way, sharing each read of the weights across many users.
Tim Dettmers and colleagues (LLM.int8()); Elias Frantar and colleagues (GPTQ)University of Washington, United States; IST Austria
LLM.int8() ran models with 175 billion parameters in 8-bit with no loss in quality. Weeks later GPTQ showed weights could go down to 3 or 4 bits with small losses.
Why it mattered. Huge models could suddenly fit on far fewer GPUs, and smaller ones on laptops.
Government of India (Ministry of Electronics and IT)Gandhinagar, India
The Prime Minister launched Digital India Bhashini at Digital India Week 2022, a platform for translation and speech tools in Indian languages, open to apps and services.
Why it mattered. Serving AI in many languages, cheaply and at scale, became national infrastructure.
Gerganov released llama.cpp, plain C/C++ code that ran Meta's LLaMA models on an ordinary MacBook using 4-bit weights. Thousands of people began running LLMs on their own computers.
Why it mattered. It made on-device LLM inference practical for anyone, and pushed quantisation into everyday use.
Woosuk Kwon, Zhuohan Li, Ion Stoica and colleaguesUC Berkeley, United States
The team found that LLM servers wasted 60 to 80% of their KV-cache memory. PagedAttention stores the cache in small pages, like an operating system handles memory, so many more users fit on each GPU. They reported 2 to 4 times the throughput of earlier systems.
Why it mattered. Their open-source vLLM became one of the most widely used ways to serve LLMs.
Speculative decoding: guess ahead, check in one go
Yaniv Leviathan, Matan Kalman and Yossi Matias (Google); independently Charlie Chen and colleagues (DeepMind)Google Research, Israel; DeepMind, London
A small, fast model drafts several tokens; the big model checks them all in a single pass and keeps the ones it agrees with. The output is exactly what the big model alone would give, typically 2 to 3 times faster.
Why it mattered. It beats the one-token-at-a-time memory wall without changing the answers.
NPCI, with AI4Bharat at IIT Madras and BhashiniMumbai, India
At the Global Fintech Fest, NPCI launched voice-based UPI payments in Hindi and English, with language models co-developed with AI4Bharat at IIT Madras under Bhashini. UPI itself handled over 20 billion transactions in August 2025.
Why it mattered. A glimpse of AI inference at the scale of a whole country's payments.
Sasha Luccioni, Yacine Jernite and Emma StrubellHugging Face and Carnegie Mellon University
The study measured the energy of 1,000 inferences across many tasks and models. Generating images used far more energy than classifying text, and big general-purpose models used far more than small task-specific ones.
Why it mattered. It put inference, not just training, at the centre of the AI energy debate.
Google put Gemini Nano, a small language model, onto the Pixel 8 Pro to run features like summarising recordings and smart replies on the device. Apple announced its own on-device model, about 3 billion parameters, in June 2024.
Why it mattered. Language models joined the other small models already running inside phones.
The IEA's Energy and AI report estimated that data centres used about 415 TWh in 2024, around 1.5% of the world's electricity, and could reach about 945 TWh by 2030, with AI the main driver.
Why it mattered. It gave governments a shared, careful baseline for the energy cost of AI.
Google reported that a median text prompt to its Gemini app used about 0.24 Wh of energy and 0.26 mL of water, and said the figure had fallen 33-fold in a year. Epoch AI had earlier estimated about 0.3 Wh for a typical GPT-4o query.
Why it mattered. Per-answer figures finally became public, though they are company numbers that are hard to check.
For a language model, each new token takes about 2 floating-point operations per parameter: an 8-billion-parameter model needs about 16 billion for every token.
Google's first TPU had 65,536 multiply-add units working in lockstep, each on 8-bit numbers.
LLM servers were once found to waste 60 to 80% of the memory set aside for the KV cache; paging it like an operating system fixed most of that.
Google reported that the energy for a median Gemini text prompt fell 33-fold in the year to mid-2025.
Data centres used about 1.5% of the world's electricity in 2024, according to the IEA.
The people
Who figured it out
FR
Frank Rosenblatt
1928 – 1971 · Psychologist and engineer · United States
Built the Mark I Perceptron, a neural network whose weights were motor-driven knobs.
ES
Edward Shortliffe
born 1947 · Physician and computer scientist · Canada / United States
Created MYCIN, a rule-based system that suggested antibiotics.
NJ
Norman Jouppi
— · Computer architect · United States
Led the design of Google's first Tensor Processing Unit for inference.
SH
Song Han
— · Computer scientist · China / United States
Showed with Deep Compression that networks can shrink 35-fold with pruning and quantisation.
SH
Sara Hooker
— · Computer scientist · United States / Canada
Argued in 'The Hardware Lottery' (2020) that the ideas that win in AI are shaped by the chips available.
TD
Tri Dao
— · Computer scientist · Vietnam / United States
Co-created FlashAttention, which speeds up attention by cutting trips to memory.
GG
Georgi Gerganov
— · Software engineer · Bulgaria
Wrote llama.cpp, which put quantised LLMs on ordinary laptops and phones.
SL
Sasha Luccioni
— · AI researcher · Canada
Measured the energy and carbon of AI inference across tasks and models.