About 80 years from a neuron drawn on paper to networks that read, see and talk.
In 1943 two scientists sketched a simple maths "neuron", loosely inspired by brain cells. Learning machines caused excitement in the 1950s, then doubt and a long winter, because nobody could train networks with many layers. Backpropagation, big datasets like ImageNet and fast graphics chips finally made deep networks work, and now they sit inside phones, banks and hospitals.
Deep (multilayer) networks trained, often called the first
Alexey Ivakhnenko and Valentin Lapa, Kyiv, Ukraine
1979 to 1980
Convolution-style network for images
Kunihiko Fukushima's Neocognitron, Japan
2012
Deep network wins ImageNet
AlexNet, University of Toronto, Canada
2016
Program beats a top Go professional in a match
AlphaGo, Google DeepMind, in Seoul
2024
Nobel Prize for neural network work
John Hopfield and Geoffrey Hinton, Physics
1943The artificial neuron
1943 – 1969
The artificial neuron
Scientists sketch simple maths neurons and build the first machines that learn from examples. The hype runs ahead of the results.
1943
A neuron made of maths
Warren McCulloch and Walter PittsChicago, United States
A brain scientist and a young self-taught logician wrote a paper describing a very simple "neuron" on paper. It adds up its inputs and fires only if the total crosses a threshold. They showed that networks of these units could, in principle, work out logic.
Why it mattered. It was the first artificial neuron, a loose and simplified sketch inspired by real brain cells, not a copy of them.
Psychologist Donald Hebb published The Organization of Behavior. He suggested that when one brain cell keeps helping another to fire, the link between them grows stronger. That is one idea of how learning could live in connections.
Why it mattered. It gave neural networks their big idea about learning: change the strength of the connections.
Frank RosenblattCornell Aeronautical Laboratory, Buffalo, United States
Psychologist Frank Rosenblatt wrote a report describing the perceptron, a machine that could learn to sort patterns by adjusting its weights after each mistake. He soon tested the idea as a program on an IBM 704 computer.
Why it mattered. For the first time, an artificial neuron came with a rule for learning from examples.
Frank Rosenblatt and the US Office of Naval ResearchWashington, D.C., United States
At a US Navy press event, a perceptron running on an IBM 704 learned to tell cards marked on the left from cards marked on the right after about 50 tries. The next day The New York Times said the Navy expected such machines to one day walk, talk, see and write. Rosenblatt's full paper came out in Psychological Review later that year.
Why it mattered. It was a real step, but the hype ran far ahead of what one layer of neurons could do.
Frank Rosenblatt and teamCornell Aeronautical Laboratory, Buffalo, United States
The Mark I Perceptron was a large machine built for image recognition. Its camera was a grid of 20 by 20 light sensors, just 400 pixels, and its weights were dials turned by small electric motors. It was shown in public in June 1960.
Why it mattered. It was one of the first machines built to learn to recognise simple pictures.
Bernard Widrow and Ted HoffStanford University, United States
Bernard Widrow and his student Ted Hoff built ADALINE, a single artificial neuron that nudged its weights to shrink its error bit by bit. Their learning rule, now called the delta rule or least mean squares, is still used.
Why it mattered. A version of it was later used to cancel echoes on phone lines, often called one of the first real-world uses of a neural network.
Alexey Ivakhnenko and Valentin LapaKyiv, Ukrainian SSR (now Ukraine)
At the Institute of Cybernetics in Kyiv, Alexey Ivakhnenko and Valentin Lapa trained networks with several layers, adding and pruning one layer at a time. Their method was called the Group Method of Data Handling. By 1971 Ivakhnenko reported a network eight layers deep.
Why it mattered. It is often called the first working deep learning, decades before the word existed.
Funding dries up after sharp criticism, but a few researchers in Finland, the US and Japan quietly invent the pieces that will matter later.
1969
Perceptrons, the book
Marvin Minsky and Seymour PapertMIT, Cambridge, United States
Two MIT researchers published a careful mathematical book showing what a single-layer perceptron cannot do. One famous example: it cannot learn XOR, "one or the other but not both". Multilayer networks could, but nobody yet knew a good way to train them.
Why it mattered. The book helped push funding and interest away from neural networks for more than a decade.
In his master's thesis, Finnish student Seppo Linnainmaa described reverse-mode automatic differentiation. It works backwards through a long chain of sums to find how much each step affected the final answer.
Why it mattered. It is the maths at the heart of backpropagation, although he was not thinking about neural networks.
The British government asked mathematician James Lighthill to judge AI research. His report said AI had not lived up to its grand promises, and UK funding was cut sharply.
Why it mattered. Together with the doubts after Perceptrons, it marks what people now call the first AI winter.
In his PhD thesis, Paul Werbos proposed using this backwards way of computing blame to train networks with many layers. Few people noticed at the time.
Why it mattered. The key idea for training deep networks existed, but it waited twelve years to catch on.
Kunihiko Fukushima built a layered network for recognising handwritten shapes. Early layers spotted small strokes, and later layers combined them into bigger patterns, an idea inspired by studies of the cat's visual cortex.
Why it mattered. Its design is the ancestor of the convolutional networks that read images today.
Physicist John Hopfield described a network that stores patterns like memories. Give it a damaged or partial pattern and it settles, step by step, into the closest stored one, like a ball rolling into a valley.
Why it mattered. It brought physicists into the field and helped wake neural networks from their winter.
Backpropagation makes multilayer networks trainable. Networks read handwriting and cheques, but computers are still too slow for big ones.
1986
9 October 1986
Backpropagation catches on
David Rumelhart, Geoffrey Hinton and Ronald WilliamsSan Diego and Pittsburgh, United States
A short paper in Nature showed that backpropagation could train networks with hidden layers, and that those layers learned useful features by themselves. They did not invent the method, which had earlier roots, but they made it popular.
Why it mattered. Multilayer networks could finally be trained, and the XOR problem stopped being a wall.
Yann LeCun and colleaguesAT&T Bell Labs, Holmdel, United States
Yann LeCun's team trained a convolutional network with backpropagation to read handwritten digits on US mail. The network shared the same small filters across the whole image, so it needed far fewer weights.
Why it mattered. It showed that learned networks could beat hand-written rules at reading real, messy handwriting.
George CybenkoUniversity of Illinois, United States
Mathematician George Cybenko proved that a network with one hidden layer, if it is wide enough, can get as close as you like to any smooth pattern. Others proved similar results the same year.
Why it mattered. It explained why networks are so flexible, and also why they can memorise noise if you are not careful.
Sepp Hochreiter and Jürgen SchmidhuberMunich, Germany, and Lugano, Switzerland
Sepp Hochreiter and Jürgen Schmidhuber designed the LSTM, a network with memory cells and gates that decide what to keep and what to forget. It could learn from long sequences, like sentences or speech.
Why it mattered. LSTMs later powered speech recognition and translation on phones for years.
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Labs, United States
A long paper described LeNet-5, a seven-layer convolutional network with about 60,000 weights, and a system for reading handwritten bank cheques. By around 2000, systems like it were reported to read roughly 10 percent or more of cheques in the United States.
Why it mattered. It was one of the first neural networks doing everyday work at a large scale.
Many-layer networks, huge labelled datasets and graphics chips come together. Networks start to beat older methods at seeing, hearing and playing games.
2006
July 2006
Deep networks, one layer at a time
Geoffrey Hinton, Simon Osindero and Yee-Whye TehUniversity of Toronto, Canada
Hinton and colleagues showed a way to train a deep network by first teaching it one layer at a time, then fine-tuning the whole thing. They called it a deep belief net.
Why it mattered. It made "deep learning" a respectable phrase again and drew researchers back to many-layer networks.
Fei-Fei Li, Jia Deng and colleaguesPrinceton University, United States
Fei-Fei Li's team presented ImageNet, a huge collection of photos sorted into thousands of categories, labelled with help from online workers. From 2010 it hosted a yearly contest, the ImageNet Large Scale Visual Recognition Challenge.
Why it mattered. Big, varied data is one of the best cures for overfitting, and ImageNet gave networks enough of it.
Alex Krizhevsky, Ilya Sutskever and Geoffrey HintonUniversity of Toronto, Canada
A deep convolutional network with about 60 million weights, trained on two gaming graphics cards, won the ImageNet contest. Its top-5 error was about 15.3 percent, against about 26.2 percent for the best entry from another team. It used dropout, randomly switching off neurons during training, to fight overfitting.
Why it mattered. This result convinced much of the field, and soon the tech industry, that deep learning worked.
AlphaGo, which used neural networks to judge positions and pick moves, beat Lee Sedol, one of the world's best Go players, by 4 games to 1. Its move 37 in game 2 surprised experts. Lee won game 4 with a brilliant move of his own.
Why it mattered. Go has more possible positions than atoms in the universe, so this showed learned intuition could beat brute force.
Ashish Vaswani and seven co-authorsGoogle Brain and Google Research, Mountain View, United States
Eight researchers at Google described the transformer, a network that lets every word look at every other word through "attention". It trained fast on many chips at once. Our box LLMClear explains how it works.
Why it mattered. Transformers became the engine of today's chatbots and large language models.
How often the winning program's five best guesses all missed the right label. Deep networks arrived in 2012, and the error fell fast.
2010 NEC and UIUC, not a deep network: 28.2%
2011 XRCE, not a deep network: 25.8%
2012 AlexNet (SuperVision): 15.3%
2013 Clarifai: 11.2%
2014 GoogLeNet: 6.7%
2015 ResNet (Microsoft Research Asia): 3.57%
2016 Trimps-Soushen: about 2.99%
2017 SENet (team WMW): about 2.25%
2018 – today
Everywhere, and debated
Neural networks spread into science, government and daily life, including India's languages and payments. People argue about fairness, energy and jobs.
2018
4 June 2018
#AIforAll: India's AI strategy
NITI AayogNew Delhi, India
India's policy think tank NITI Aayog published a discussion paper, the National Strategy for Artificial Intelligence. It picked five areas where AI could help most: health, farming, education, smart cities and transport.
Why it mattered. It set out India's first national plan for using AI for everyone, not just for big companies.
Geoffrey Hinton, Yann LeCun and Yoshua BengioNew York, United States (ACM)
The ACM gave its 2018 Turing Award, often called the Nobel Prize of computing, to three researchers who kept faith in neural networks during the lean years.
Why it mattered. It marked how far neural networks had travelled, from fringe idea to the centre of computer science.
John Jumper, Demis Hassabis and the DeepMind teamLondon, United Kingdom
At the CASP14 contest, DeepMind's AlphaFold 2 predicted the 3D shapes of proteins with accuracy close to lab experiments for many targets. Protein shapes decide how medicines and diseases work.
Why it mattered. It showed neural networks helping with a real science problem, and later won a share of the 2024 Nobel Prize in Chemistry.
Ministry of Electronics and IT, Government of IndiaGandhinagar, India
The Prime Minister launched Digital India Bhashini during Digital India Week. It aims to build AI tools for translation and speech in Indian languages, using datasets that citizens can help create through BhashaDaan.
Why it mattered. Neural networks learn from data, so a country with many languages needs data in all of them.
Union Cabinet, Government of IndiaNew Delhi, India
The Union Cabinet approved the IndiaAI Mission with an outlay of about Rs 10,372 crore over five years. Its plans include shared GPU computing for researchers and startups, datasets, skills and "safe and trusted AI".
Why it mattered. Training neural networks needs huge computing power, and the mission aims to make it available within India.
John Hopfield and Geoffrey HintonStockholm, Sweden
The Royal Swedish Academy of Sciences gave the Nobel Prize in Physics to John Hopfield and Geoffrey Hinton "for foundational discoveries and inventions that enable machine learning with artificial neural networks". Both had borrowed tools from physics to build their networks.
Why it mattered. It was the first Nobel Prize given for work on artificial neural networks.
Reported sizes of some famous networks, from tens of thousands of weights to hundreds of billions.
1998 LeNet-5: about 60,000 (reported)
2012 AlexNet: about 60 million (reported)
2019 GPT-2: about 1.5 billion (reported)
2020 GPT-3: about 175 billion (reported)
Did you know?
An artificial neuron does only three things: multiply each input by a weight, add them up, and squash the total.
In July 1958 The New York Times reported that the Navy expected the perceptron to one day walk, talk, see and write. It could tell left from right.
The Mark I Perceptron's camera had just 400 light sensors, 20 by 20. A cheap phone camera today has millions.
AlexNet, the network that started the deep learning boom in 2012, was trained on two graphics cards made for video games.
Geoffrey Hinton's Nobel Prize is in physics, not computer science: there is no Nobel for computing, and the prize honoured ideas borrowed from physics.
The people
Who figured it out
WM
Warren McCulloch
1898 – 1969 · Neurophysiologist · United States
Co-wrote the 1943 paper that described the first artificial neuron.
WP
Walter Pitts
1923 – 1969 · Logician · United States
A largely self-taught teenager who became McCulloch's co-author and turned neurons into logic.
DH
Donald Hebb
1904 – 1985 · Psychologist · Canada
Suggested that learning happens as connections between brain cells grow stronger with use.
FR
Frank Rosenblatt
1928 – 1971 · Psychologist and engineer · United States
Invented the perceptron and built the Mark I, one of the first machines that learned from examples.
AI
Alexey Ivakhnenko
1913 – 2007 · Mathematician and engineer · Ukraine
Trained multilayer networks in Kyiv in the 1960s, often called the first deep learning.
KF
Kunihiko Fukushima
1936 – · Computer scientist · Japan
Built the Neocognitron, the layered image network that inspired convolutional networks.
JH
John Hopfield
1933 – · Physicist · United States
Designed a network that stores memories like valleys in a landscape, and shared the 2024 Nobel Prize in Physics.
GH
Geoffrey Hinton
1947 – · Computer scientist and psychologist · United Kingdom and Canada
Helped popularise backpropagation, revived deep learning in 2006, and shared both the Turing Award and a Nobel Prize.
YL
Yann LeCun
1960 – · Computer scientist · France
Built convolutional networks that read zip codes and bank cheques at Bell Labs.
YB
Yoshua Bengio
1964 – · Computer scientist · France and Canada
A leader of deep learning research in Montreal who shared the 2018 Turing Award.
FL
Fei-Fei Li
1976 – · Computer scientist · China and United States
Led the creation of ImageNet, the giant labelled photo collection that powered the deep learning boom.