The history

The history of neural networks

About 80 years from a neuron drawn on paper to networks that read, see and talk.

In 1943 two scientists sketched a simple maths "neuron", loosely inspired by brain cells. Learning machines caused excitement in the 1950s, then doubt and a long winter, because nobody could train networks with many layers. Backpropagation, big datasets like ImageNet and fast graphics chips finally made deep networks work, and now they sit inside phones, banks and hospitals.

81
years
29
moments
11
people
10
places

1943

Artificial neuron described

Warren McCulloch and Walter Pitts, United States

1958

Learning machine shown to the press

Frank Rosenblatt's perceptron, United States

1965 to 1971

Deep (multilayer) networks trained, often called the first

Alexey Ivakhnenko and Valentin Lapa, Kyiv, Ukraine

1979 to 1980

Convolution-style network for images

Kunihiko Fukushima's Neocognitron, Japan

2012

Deep network wins ImageNet

AlexNet, University of Toronto, Canada

2016

Program beats a top Go professional in a match

AlphaGo, Google DeepMind, in Seoul

2024

Nobel Prize for neural network work

John Hopfield and Geoffrey Hinton, Physics

1943The artificial neuron

1943 – 1969

The artificial neuron

Scientists sketch simple maths neurons and build the first machines that learn from examples. The hype runs ahead of the results.

1943

A neuron made of maths

Warren McCulloch and Walter PittsChicago, United States

A brain scientist and a young self-taught logician wrote a paper describing a very simple "neuron" on paper. It adds up its inputs and fires only if the total crosses a threshold. They showed that networks of these units could, in principle, work out logic.

Why it mattered. It was the first artificial neuron, a loose and simplified sketch inspired by real brain cells, not a copy of them.

1949

Cells that fire together wire together

Donald HebbMontreal, Canada

Psychologist Donald Hebb published The Organization of Behavior. He suggested that when one brain cell keeps helping another to fire, the link between them grows stronger. That is one idea of how learning could live in connections.

Why it mattered. It gave neural networks their big idea about learning: change the strength of the connections.

1957

The perceptron is proposed

Frank RosenblattCornell Aeronautical Laboratory, Buffalo, United States

Psychologist Frank Rosenblatt wrote a report describing the perceptron, a machine that could learn to sort patterns by adjusting its weights after each mistake. He soon tested the idea as a program on an IBM 704 computer.

Why it mattered. For the first time, an artificial neuron came with a rule for learning from examples.

1958

7 to 8 July 1958

A learning machine makes headlines

Frank Rosenblatt and the US Office of Naval ResearchWashington, D.C., United States

At a US Navy press event, a perceptron running on an IBM 704 learned to tell cards marked on the left from cards marked on the right after about 50 tries. The next day The New York Times said the Navy expected such machines to one day walk, talk, see and write. Rosenblatt's full paper came out in Psychological Review later that year.

Why it mattered. It was a real step, but the hype ran far ahead of what one layer of neurons could do.

1960

23 June 1960

The Mark I Perceptron gets an eye

Frank Rosenblatt and teamCornell Aeronautical Laboratory, Buffalo, United States

The Mark I Perceptron was a large machine built for image recognition. Its camera was a grid of 20 by 20 light sensors, just 400 pixels, and its weights were dials turned by small electric motors. It was shown in public in June 1960.

Why it mattered. It was one of the first machines built to learn to recognise simple pictures.

1960

ADALINE learns from its error

Bernard Widrow and Ted HoffStanford University, United States

Bernard Widrow and his student Ted Hoff built ADALINE, a single artificial neuron that nudged its weights to shrink its error bit by bit. Their learning rule, now called the delta rule or least mean squares, is still used.

Why it mattered. A version of it was later used to cancel echoes on phone lines, often called one of the first real-world uses of a neural network.

1965

1965 to 1967

Deep networks in Kyiv

Alexey Ivakhnenko and Valentin LapaKyiv, Ukrainian SSR (now Ukraine)

At the Institute of Cybernetics in Kyiv, Alexey Ivakhnenko and Valentin Lapa trained networks with several layers, adding and pruning one layer at a time. Their method was called the Group Method of Data Handling. By 1971 Ivakhnenko reported a network eight layers deep.

Why it mattered. It is often called the first working deep learning, decades before the word existed.

1969 – 1986

Winter and quiet progress

Funding dries up after sharp criticism, but a few researchers in Finland, the US and Japan quietly invent the pieces that will matter later.

1969

Perceptrons, the book

Marvin Minsky and Seymour PapertMIT, Cambridge, United States

Two MIT researchers published a careful mathematical book showing what a single-layer perceptron cannot do. One famous example: it cannot learn XOR, "one or the other but not both". Multilayer networks could, but nobody yet knew a good way to train them.

Why it mattered. The book helped push funding and interest away from neural networks for more than a decade.

1970

A quiet trick for computing blame

Seppo LinnainmaaUniversity of Helsinki, Finland

In his master's thesis, Finnish student Seppo Linnainmaa described reverse-mode automatic differentiation. It works backwards through a long chain of sums to find how much each step affected the final answer.

Why it mattered. It is the maths at the heart of backpropagation, although he was not thinking about neural networks.

1973

The Lighthill report

James LighthillUnited Kingdom

The British government asked mathematician James Lighthill to judge AI research. His report said AI had not lived up to its grand promises, and UK funding was cut sharply.

Why it mattered. Together with the doubts after Perceptrons, it marks what people now call the first AI winter.

1974

Backprop for neural networks, on paper

Paul WerbosHarvard University, United States

In his PhD thesis, Paul Werbos proposed using this backwards way of computing blame to train networks with many layers. Few people noticed at the time.

Why it mattered. The key idea for training deep networks existed, but it waited twelve years to catch on.

1980

1979 (Japanese), 1980 (English)

The Neocognitron

Kunihiko FukushimaNHK laboratories, Tokyo, Japan

Kunihiko Fukushima built a layered network for recognising handwritten shapes. Early layers spotted small strokes, and later layers combined them into bigger patterns, an idea inspired by studies of the cat's visual cortex.

Why it mattered. Its design is the ancestor of the convolutional networks that read images today.

1982

April 1982

A network that remembers

John HopfieldCaltech, Pasadena, United States

Physicist John Hopfield described a network that stores patterns like memories. Give it a damaged or partial pattern and it settles, step by step, into the closest stored one, like a ball rolling into a valley.

Why it mattered. It brought physicists into the field and helped wake neural networks from their winter.

1986 – 2006

Learning to learn: backprop

Backpropagation makes multilayer networks trainable. Networks read handwriting and cheques, but computers are still too slow for big ones.

1986

9 October 1986

Backpropagation catches on

David Rumelhart, Geoffrey Hinton and Ronald WilliamsSan Diego and Pittsburgh, United States

A short paper in Nature showed that backpropagation could train networks with hidden layers, and that those layers learned useful features by themselves. They did not invent the method, which had earlier roots, but they made it popular.

Why it mattered. Multilayer networks could finally be trained, and the XOR problem stopped being a wall.

1989

Reading handwritten zip codes

Yann LeCun and colleaguesAT&T Bell Labs, Holmdel, United States

Yann LeCun's team trained a convolutional network with backpropagation to read handwritten digits on US mail. The network shared the same small filters across the whole image, so it needed far fewer weights.

Why it mattered. It showed that learned networks could beat hand-written rules at reading real, messy handwriting.

1989

One hidden layer can fit almost anything

George CybenkoUniversity of Illinois, United States

Mathematician George Cybenko proved that a network with one hidden layer, if it is wide enough, can get as close as you like to any smooth pattern. Others proved similar results the same year.

Why it mattered. It explained why networks are so flexible, and also why they can memorise noise if you are not careful.

1997

November 1997

Long short-term memory

Sepp Hochreiter and Jürgen SchmidhuberMunich, Germany, and Lugano, Switzerland

Sepp Hochreiter and Jürgen Schmidhuber designed the LSTM, a network with memory cells and gates that decide what to keep and what to forget. It could learn from long sequences, like sentences or speech.

Why it mattered. LSTMs later powered speech recognition and translation on phones for years.

1998

November 1998

LeNet-5 reads cheques

Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Labs, United States

A long paper described LeNet-5, a seven-layer convolutional network with about 60,000 weights, and a system for reading handwritten bank cheques. By around 2000, systems like it were reported to read roughly 10 percent or more of cheques in the United States.

Why it mattered. It was one of the first neural networks doing everyday work at a large scale.

2006 – 2018

Deep learning

Many-layer networks, huge labelled datasets and graphics chips come together. Networks start to beat older methods at seeing, hearing and playing games.

2006

July 2006

Deep networks, one layer at a time

Geoffrey Hinton, Simon Osindero and Yee-Whye TehUniversity of Toronto, Canada

Hinton and colleagues showed a way to train a deep network by first teaching it one layer at a time, then fine-tuning the whole thing. They called it a deep belief net.

Why it mattered. It made "deep learning" a respectable phrase again and drew researchers back to many-layer networks.

2009

June 2009

ImageNet: millions of labelled pictures

Fei-Fei Li, Jia Deng and colleaguesPrinceton University, United States

Fei-Fei Li's team presented ImageNet, a huge collection of photos sorted into thousands of categories, labelled with help from online workers. From 2010 it hosted a yearly contest, the ImageNet Large Scale Visual Recognition Challenge.

Why it mattered. Big, varied data is one of the best cures for overfitting, and ImageNet gave networks enough of it.

2012

30 September 2012 (results)

AlexNet wins ImageNet

Alex Krizhevsky, Ilya Sutskever and Geoffrey HintonUniversity of Toronto, Canada

A deep convolutional network with about 60 million weights, trained on two gaming graphics cards, won the ImageNet contest. Its top-5 error was about 15.3 percent, against about 26.2 percent for the best entry from another team. It used dropout, randomly switching off neurons during training, to fight overfitting.

Why it mattered. This result convinced much of the field, and soon the tech industry, that deep learning worked.

2016

9 to 15 March 2016

AlphaGo beats Lee Sedol

Google DeepMindSeoul, South Korea

AlphaGo, which used neural networks to judge positions and pick moves, beat Lee Sedol, one of the world's best Go players, by 4 games to 1. Its move 37 in game 2 surprised experts. Lee won game 4 with a brilliant move of his own.

Why it mattered. Go has more possible positions than atoms in the universe, so this showed learned intuition could beat brute force.

2017

June 2017

Attention is all you need

Ashish Vaswani and seven co-authorsGoogle Brain and Google Research, Mountain View, United States

Eight researchers at Google described the transformer, a network that lets every word look at every other word through "attention". It trained fast on many chips at once. Our box LLMClear explains how it works.

Why it mattered. Transformers became the engine of today's chatbots and large language models.

By the numbers

Winning error in the ImageNet contest

How often the winning program's five best guesses all missed the right label. Deep networks arrived in 2012, and the error fell fast.

0 % top-5 error10 % top-5 error20 % top-5 error30 % top-5 error 20102015 2010: NEC and UIUC, not a deep network: 28.2%20102011: XRCE, not a deep network: 25.8%20112012: AlexNet (SuperVision): 15.3%20122013: Clarifai: 11.2%20132014: GoogLeNet: 6.7%20142015: ResNet (Microsoft Research Asia): 3.57%20152016: Trimps-Soushen: about 2.99%20162017: SENet (team WMW): about 2.25%2017
  1. 2010 NEC and UIUC, not a deep network: 28.2%
  2. 2011 XRCE, not a deep network: 25.8%
  3. 2012 AlexNet (SuperVision): 15.3%
  4. 2013 Clarifai: 11.2%
  5. 2014 GoogLeNet: 6.7%
  6. 2015 ResNet (Microsoft Research Asia): 3.57%
  7. 2016 Trimps-Soushen: about 2.99%
  8. 2017 SENet (team WMW): about 2.25%

2018 – today

Everywhere, and debated

Neural networks spread into science, government and daily life, including India's languages and payments. People argue about fairness, energy and jobs.

2018

4 June 2018

#AIforAll: India's AI strategy

NITI AayogNew Delhi, India

India's policy think tank NITI Aayog published a discussion paper, the National Strategy for Artificial Intelligence. It picked five areas where AI could help most: health, farming, education, smart cities and transport.

Why it mattered. It set out India's first national plan for using AI for everyone, not just for big companies.

2019

27 March 2019

A Turing Award for deep learning

Geoffrey Hinton, Yann LeCun and Yoshua BengioNew York, United States (ACM)

The ACM gave its 2018 Turing Award, often called the Nobel Prize of computing, to three researchers who kept faith in neural networks during the lean years.

Why it mattered. It marked how far neural networks had travelled, from fringe idea to the centre of computer science.

2020

30 November 2020

AlphaFold 2 predicts protein shapes

John Jumper, Demis Hassabis and the DeepMind teamLondon, United Kingdom

At the CASP14 contest, DeepMind's AlphaFold 2 predicted the 3D shapes of proteins with accuracy close to lab experiments for many targets. Protein shapes decide how medicines and diseases work.

Why it mattered. It showed neural networks helping with a real science problem, and later won a share of the 2024 Nobel Prize in Chemistry.

2022

4 July 2022

Bhashini: AI for Indian languages

Ministry of Electronics and IT, Government of IndiaGandhinagar, India

The Prime Minister launched Digital India Bhashini during Digital India Week. It aims to build AI tools for translation and speech in Indian languages, using datasets that citizens can help create through BhashaDaan.

Why it mattered. Neural networks learn from data, so a country with many languages needs data in all of them.

2024

7 March 2024

The IndiaAI Mission

Union Cabinet, Government of IndiaNew Delhi, India

The Union Cabinet approved the IndiaAI Mission with an outlay of about Rs 10,372 crore over five years. Its plans include shared GPU computing for researchers and startups, datasets, skills and "safe and trusted AI".

Why it mattered. Training neural networks needs huge computing power, and the mission aims to make it available within India.

2024

8 October 2024

A Nobel Prize in Physics for neural networks

John Hopfield and Geoffrey HintonStockholm, Sweden

The Royal Swedish Academy of Sciences gave the Nobel Prize in Physics to John Hopfield and Geoffrey Hinton "for foundational discoveries and inventions that enable machine learning with artificial neural networks". Both had borrowed tools from physics to build their networks.

Why it mattered. It was the first Nobel Prize given for work on artificial neural networks.

By the numbers

How many weights a famous network has

Reported sizes of some famous networks, from tens of thousands of weights to hundreds of billions.

10,000100,0001,000,00010,000,000100,000,0001,000,000,00010,000,000,000100,000,000,0001,000,000,000,000 20002005201020152020 1998: LeNet-5: about 60,000 (reported)19982012: AlexNet: about 60 million (reported)20122019: GPT-2: about 1.5 billion (reported)20192020: GPT-3: about 175 billion (reported)
  1. 1998 LeNet-5: about 60,000 (reported)
  2. 2012 AlexNet: about 60 million (reported)
  3. 2019 GPT-2: about 1.5 billion (reported)
  4. 2020 GPT-3: about 175 billion (reported)

Did you know?

An artificial neuron does only three things: multiply each input by a weight, add them up, and squash the total.

In July 1958 The New York Times reported that the Navy expected the perceptron to one day walk, talk, see and write. It could tell left from right.

The Mark I Perceptron's camera had just 400 light sensors, 20 by 20. A cheap phone camera today has millions.

AlexNet, the network that started the deep learning boom in 2012, was trained on two graphics cards made for video games.

Geoffrey Hinton's Nobel Prize is in physics, not computer science: there is no Nobel for computing, and the prize honoured ideas borrowed from physics.

The people

Who figured it out

Warren McCulloch

1898 – 1969 · Neurophysiologist · United States

Co-wrote the 1943 paper that described the first artificial neuron.

Walter Pitts

1923 – 1969 · Logician · United States

A largely self-taught teenager who became McCulloch's co-author and turned neurons into logic.

Donald Hebb

1904 – 1985 · Psychologist · Canada

Suggested that learning happens as connections between brain cells grow stronger with use.

Frank Rosenblatt

1928 – 1971 · Psychologist and engineer · United States

Invented the perceptron and built the Mark I, one of the first machines that learned from examples.

Alexey Ivakhnenko

1913 – 2007 · Mathematician and engineer · Ukraine

Trained multilayer networks in Kyiv in the 1960s, often called the first deep learning.

Kunihiko Fukushima

1936 – · Computer scientist · Japan

Built the Neocognitron, the layered image network that inspired convolutional networks.

John Hopfield

1933 – · Physicist · United States

Designed a network that stores memories like valleys in a landscape, and shared the 2024 Nobel Prize in Physics.

Geoffrey Hinton

1947 – · Computer scientist and psychologist · United Kingdom and Canada

Helped popularise backpropagation, revived deep learning in 2006, and shared both the Turing Award and a Nobel Prize.

Yann LeCun

1960 – · Computer scientist · France

Built convolutional networks that read zip codes and bank cheques at Bell Labs.

Yoshua Bengio

1964 – · Computer scientist · France and Canada

A leader of deep learning research in Montreal who shared the 2018 Turing Award.

Fei-Fei Li

1976 – · Computer scientist · China and United States

Led the creation of ImageNet, the giant labelled photo collection that powered the deep learning boom.

Where it happened

10 places, one idea

Sources

Where this comes from

Dates marked “c.” are approximate, and historians sometimes disagree about who was first. If you spot a mistake, tell us.

  1. History of artificial neural networks Wikipedia
  2. A logical calculus of the ideas immanent in nervous activity (McCulloch and Pitts, 1943) Bulletin of Mathematical Biophysics, Springer
  3. Artificial neuron Wikipedia
  4. Hebbian theory Wikipedia
  5. Perceptron Wikipedia
  6. Professor's perceptron paved the way for AI, 60 years too soon Cornell Chronicle, Cornell University
  7. New Navy device learns by doing (8 July 1958) The New York Times archive
  8. Mark I Perceptron Wikipedia
  9. ADALINE Wikipedia
  10. Group method of data handling Wikipedia
  11. Perceptrons (book) Wikipedia
  12. AI winter Wikipedia
  13. Lighthill report Wikipedia
  14. Backpropagation Wikipedia
  15. Seppo Linnainmaa Wikipedia
  16. Paul Werbos Wikipedia
  17. Neocognitron Wikipedia
  18. Hopfield network Wikipedia
  19. Neural networks and physical systems with emergent collective computational abilities (Hopfield, 1982) PNAS
  20. Learning representations by back-propagating errors (Rumelhart, Hinton and Williams, 1986) Nature
  21. LeNet Wikipedia
  22. Gradient-based learning applied to document recognition (LeCun et al., 1998) Proceedings of the IEEE
  23. Yann LeCun Wikipedia
  24. Universal approximation theorem Wikipedia
  25. Long short-term memory (Hochreiter and Schmidhuber, 1997) Neural Computation, MIT Press
  26. Long short-term memory Wikipedia
  27. A fast learning algorithm for deep belief nets (Hinton, Osindero and Teh, 2006) Neural Computation, MIT Press
  28. Deep belief network Wikipedia
  29. ImageNet Wikipedia
  30. ImageNet: A large-scale hierarchical image database (Deng et al., CVPR 2009) IEEE
  31. AlexNet Wikipedia
  32. ImageNet classification with deep convolutional neural networks (Krizhevsky, Sutskever and Hinton, 2012) NeurIPS Proceedings
  33. ILSVRC 2012 results ImageNet (image-net.org)
  34. ImageNet Large Scale Visual Recognition Challenge (Russakovsky et al., 2015) International Journal of Computer Vision / arXiv
  35. ILSVRC 2013 results ImageNet (image-net.org)
  36. ILSVRC 2014 results ImageNet (image-net.org)
  37. Deep residual learning for image recognition (He et al., 2015) arXiv
  38. ILSVRC 2016 results ImageNet (image-net.org)
  39. Squeeze-and-excitation networks (Hu et al., 2017) arXiv
  40. AlphaGo Google DeepMind
  41. AlphaGo versus Lee Sedol Wikipedia
  42. Attention is all you need (Vaswani et al., 2017) arXiv
  43. Fathers of the deep learning revolution receive ACM A.M. Turing Award ACM
  44. AlphaFold: a solution to a 50-year-old grand challenge in biology Google DeepMind
  45. AlphaFold Wikipedia
  46. The Nobel Prize in Physics 2024: press release nobelprize.org
  47. The Nobel Prize in Physics 2024: popular science background nobelprize.org
  48. National Strategy for Artificial Intelligence #AIforAll (discussion paper, June 2018) NITI Aayog
  49. PM inaugurates Digital India Week 2022 in Gandhinagar Press Information Bureau (PIB)
  50. Bhashini Wikipedia
  51. Cabinet approves ambitious IndiaAI Mission to strengthen the AI innovation ecosystem Press Information Bureau (PIB)
  52. Better language models and their implications (GPT-2) OpenAI
  53. Language models are few-shot learners (Brown et al., 2020) arXiv
  54. Geoffrey Hinton Wikipedia
  55. Fei-Fei Li Wikipedia
  56. Alexey Ivakhnenko Wikipedia
  57. Frank Rosenblatt Wikipedia

That's the history. Now see how it works.