The history

The history of image recognition

About 70 years from the first digital photo to machines that look and talk.

In 1957 a photo of a baby became the first grid of pixels in a computer. Researchers thought seeing would be easy to program, and learned over decades that it was one of the hardest problems in computing. Hand-made edge and corner finders gave way to networks that learn their own filters from millions of pictures, and now vision is joining language, bringing new uses and new questions about fairness and privacy.

67
years
30
moments
9
people
10
places

1957

First digital photo scanned into a computer

Russell Kirsch and team, National Bureau of Standards, United States

1979 to 1980

Layered, convolution-style vision network

Kunihiko Fukushima's Neocognitron, Japan

1989

Convolutional network trained with backpropagation

Yann LeCun and colleagues, AT&T Bell Labs, United States

2001

Real-time face detection

Paul Viola and Michael Jones, United States

2012

Deep network wins ImageNet

AlexNet, University of Toronto, Canada

2015

Network over 100 layers wins ImageNet

ResNet, Microsoft Research Asia, China

1957Pixels and first ideas

1957 – 1979

Pixels and first ideas

Photos become numbers, brain scientists find edge-detecting cells, and early programs try to understand simple blocks. Seeing turns out to be hard.

1957

The first digital photo

Russell Kirsch and colleaguesNational Bureau of Standards, Washington DC, United States

Russell Kirsch's team built a drum scanner for SEAC, one of the first programmable computers in the US. The first picture they fed in was a photo of his three-month-old son, Walden, turned into a grid of 176 × 176 pixels.

Why it mattered. For the first time a photo became numbers inside a computer, the starting point of all image processing.

1959

Cells that answer to edges

David Hubel and Torsten WieselJohns Hopkins University, Baltimore, United States

Recording from single brain cells in the visual cortex of cats, Hubel and Wiesel found cells that fired only for a line or edge at a particular angle and place. Later work showed more complex cells that combine them.

Why it mattered. The idea of simple edge detectors feeding more complex ones inspired the layered design of convolutional networks, loosely.

1963

The blocks world

Lawrence RobertsMIT, Cambridge, United States

In his PhD thesis, Lawrence Roberts wrote a program that found the edges of simple blocks in a photo and worked out their 3D shapes. It only worked for plain, clean polyhedra, not messy real scenes.

Why it mattered. Often called the first computer vision thesis, it set the goal of turning pixels into an understanding of 3D objects.

1966

7 July 1966

Vision as a summer project

Seymour Papert and studentsMIT Artificial Intelligence Group, Cambridge, United States

A memo from MIT planned a summer project for students to build a large part of a visual system, including telling objects from the background and naming them. It turned out to be far harder than anyone expected.

Why it mattered. It became a famous lesson: seeing feels effortless to people, but getting a computer to do it took decades.

1969

A chip that captures light

Willard Boyle and George SmithBell Labs, Murray Hill, United States

Boyle and Smith invented the charge-coupled device (CCD), a chip that stores light as packets of electric charge and reads them out pixel by pixel. It became the eye of digital cameras and telescopes.

Why it mattered. Cheap image sensors made digital pictures, and so computer vision, part of everyday life. They shared the 2009 Nobel Prize in Physics.

1979 – 2008

Hand-made features

People design edge, corner and feature detectors by hand, while the first convolutional networks quietly learn to read digits and cheques.

1980

1979 (Japanese), 1980 (English)

The Neocognitron

Kunihiko FukushimaNHK Broadcasting Science Research Laboratories, Tokyo, Japan

Inspired by Hubel and Wiesel, Kunihiko Fukushima built a network of layers: simple cells that find small features and complex cells that pool them, so a shape could be recognised even when shifted.

Why it mattered. It is the direct ancestor of the convolutional network, with filters and pooling, the design this box trains.

1981

A Nobel Prize for how we see

David Hubel and Torsten WieselStockholm, Sweden

Hubel and Wiesel shared the Nobel Prize in Physiology or Medicine for their discoveries about how the visual system processes information, with Roger Sperry, who was honoured for separate work on the brain's two halves.

Why it mattered. Their map of the visual cortex is still the textbook starting point, and EyeClear and BrainClear show the living version.

1982

Marr's Vision

David MarrMIT, Cambridge, United States

David Marr's book, published after his early death in 1980, described vision as a series of steps: from edges to a sketch of surfaces to 3D objects. It shaped how researchers thought about vision for years.

Why it mattered. It gave computer vision a clear plan of stages, many of which learned networks later handled on their own.

1986

November 1986

The Canny edge detector

John CannyMIT, Cambridge, United States

John Canny worked out, with careful maths, what a good edge detector should do: find real edges, place them exactly, and mark each only once. His method is still built into image software today.

Why it mattered. Hand-designed edge finding was the first step in most vision systems for decades.

1988

Finding corners

Chris Harris and Mike StephensPlessey Research, Roke Manor, United Kingdom

Harris and Stephens described a corner detector that looks for places where the picture changes in two directions at once. Corners are easy to find again in another photo, so they help match and track things.

Why it mattered. Chapter 2's corner filter uses their formula.

1989

A network reads zip codes

Yann LeCun and colleaguesAT&T Bell Labs, Holmdel, United States

Yann LeCun's team trained a convolutional network with backpropagation to read handwritten digits on US mail. The same small filters were shared across the whole picture, so it needed far fewer weights.

Why it mattered. It showed that a network could learn its own filters from examples, instead of people designing them.

1994

The QR code

Masahiro Hara and team at Denso WaveDenso Wave, Japan

Engineers at Denso Wave, a Toyota group company, designed the QR code to track car parts. Its three big corner squares can be found quickly from any angle, and built-in error correction fixes damaged parts.

Why it mattered. Decades later, QR codes carry UPI payments in India, and reading them still needs no neural network.

1998

November 1998

LeNet-5 reads cheques

Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Labs, United States

A long paper described LeNet-5 and a system that read the amounts on handwritten bank cheques. NCR put such readers to work from the mid-1990s, and by around 2000 they were reported to read roughly 10% of cheques in the US.

Why it mattered. It was one of the first times image recognition did everyday work at large scale.

1999

SIFT: features that survive turning and zooming

David LoweUniversity of British Columbia, Vancouver, Canada

David Lowe's SIFT method found distinctive points in a picture and described each so it could be matched in another photo, even after turning, zooming or a change of light.

Why it mattered. For a decade, hand-made features like SIFT powered panorama stitching, object recognition and robot navigation.

2001

Real-time face detection

Paul Viola and Michael JonesMitsubishi Electric Research Labs and Compaq, Cambridge, United States

Viola and Jones found faces with thousands of simple light-and-dark box features, chosen by a learning method and checked in a fast cascade. It ran at 15 frames a second on a 700 MHz Pentium III.

Why it mattered. It put face-finding boxes in digital cameras, and its sliding-window idea is the one chapter 4 uses.

2005

Finding people by their edges

Navneet Dalal and Bill TriggsINRIA, Grenoble, France

Dalal and Triggs described people by histograms of edge directions (HOG) in small cells, and slid a detector over the picture. It became a standard way to find pedestrians.

Why it mattered. It was among the best hand-made detectors just before learned features took over.

2009 – 2019

Deep learning learns to see

Huge labelled datasets and graphics chips let deep networks learn their own features. Accuracy leaps, and so do concerns about bias.

2009

June 2009

ImageNet

Fei-Fei Li, Jia Deng and colleaguesPrinceton University, United States

Fei-Fei Li's team presented ImageNet, millions of photos sorted into thousands of categories, labelled with help from online workers. From 2010 it ran a yearly contest with 1,000 categories.

Why it mattered. It gave learning systems enough varied examples to learn from, and a fair test to compare them.

2012

September 2012

AlexNet

Alex Krizhevsky, Ilya Sutskever and Geoffrey HintonUniversity of Toronto, Canada

A deep convolutional network trained on two gaming graphics cards won the ImageNet contest with a top-5 error of about 15.3%, far ahead of about 26.2% for the next team. It used augmentation and dropout to avoid memorising.

Why it mattered. This result convinced much of the field to switch from hand-made features to learned ones.

2014

June 2014

DeepFace

Yaniv Taigman and colleaguesFacebook AI Research, Menlo Park, United States

A deep network trained on about 4 million face photos from Facebook users matched pairs of faces with 97.35% accuracy on a standard test, close to the 97.53% people score.

Why it mattered. Face recognition jumped in accuracy, and so did questions about who gets recognised, and with whose photos.

2015

December 2015

ResNet: 152 layers

Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian SunMicrosoft Research Asia, Beijing, China

Very deep networks used to train badly. ResNet added shortcut connections that let each block learn only a small change, which made networks 152 layers deep trainable. It won ImageNet 2015 with a 3.57% top-5 error.

Why it mattered. Residual connections are now in almost every large network, including language models.

2015

May 2015

U-Net draws masks

Olaf Ronneberger, Philipp Fischer and Thomas BroxUniversity of Freiburg, Germany

U-Net shrinks a picture through convolution layers and then grows it back, labelling every pixel. It was built to outline cells in microscope images from only a few labelled examples.

Why it mattered. It became a standard tool for segmentation, especially in medicine.

2016

June 2016

YOLO: you only look once

Joseph Redmon, Santosh Divvala, Ross Girshick and Ali FarhadiUniversity of Washington and Allen Institute for AI, Seattle, United States

Instead of sliding a window, YOLO splits the picture into a grid and predicts all boxes and classes in one pass of one network. It ran at about 45 frames a second.

Why it mattered. Fast single-pass detectors made real-time detection practical in cars, drones, phones and cameras.

2018

February 2018

Gender Shades

Joy Buolamwini and Timnit GebruMIT Media Lab, Cambridge, United States

Testing three commercial face-analysis systems on 1,270 faces of parliament members from six countries, they found gender errors of at most 0.8% for lighter-skinned men but up to 34.7% for darker-skinned women.

Why it mattered. It showed that unbalanced data can make a system fail most for the people least represented, and pushed companies to improve.

2019

June 2019

Eye screening tested in India

Varun Gulshan and colleagues, with Aravind and Sankara NethralayaMadurai and Chennai, India

A deep-learning system for diabetic retinopathy was tested on 3,049 patients at two Indian eye hospitals. It spotted cases needing a doctor about as well as, or better than, trained graders and eye specialists.

Why it mattered. Screening aids like this could help reach millions of people with diabetes where eye doctors are scarce.

2019

December 2019

NIST measures bias in face recognition

Patrick Grother, Mei Ngan and Kayee HanaokaNIST, Gaithersburg, United States

The US standards agency tested 189 face recognition algorithms. False matches were often 10 to more than 100 times more common for some demographic groups than others, though the size of the gap depended strongly on the algorithm.

Why it mattered. It gave governments and buyers hard numbers on how accurate and how fair different systems are.

By the numbers

Winning error in the ImageNet contest

How often the winning program's five best guesses all missed the right label. Deep networks arrived in 2012, and the error fell fast.

0 % top-5 error10 % top-5 error20 % top-5 error30 % top-5 error 20102015 2010: NEC and UIUC, hand-made features: 28.2%20102011: XRCE, hand-made features: 25.8%20112012: AlexNet: 15.3%20122013: Clarifai: 11.2%20132014: GoogLeNet: 6.7%20142015: ResNet: 3.57%20152016: Trimps-Soushen: about 2.99%20162017: SENet: about 2.25%2017
  1. 2010 NEC and UIUC, hand-made features: 28.2%
  2. 2011 XRCE, hand-made features: 25.8%
  3. 2012 AlexNet: 15.3%
  4. 2013 Clarifai: 11.2%
  5. 2014 GoogLeNet: 6.7%
  6. 2015 ResNet: 3.57%
  7. 2016 Trimps-Soushen: about 2.99%
  8. 2017 SENet: about 2.25%

2020 – today

Vision meets language

Transformers and caption-trained models link pictures with words, while countries debate face recognition and privacy.

2020

October 2020

Vision Transformer

Alexey Dosovitskiy and colleaguesGoogle Research, Berlin, Germany

The paper 'An Image is Worth 16x16 Words' cut pictures into patches and fed them to a transformer, the design behind language models. With enough training data, it matched or beat the best CNNs.

Why it mattered. Vision and language models began to share one design.

2021

January 2021

CLIP: pictures meet words

Alec Radford and colleaguesOpenAI, San Francisco, United States

CLIP learned from about 400 million pictures with captions from the web, matching each picture with its text. It could then recognise things it was never given labels for, just from a written description.

Why it mattered. It led to vision-language models that caption photos and answer questions about them.

2022

1 December 2022

DigiYatra face boarding

Ministry of Civil Aviation and the Digi Yatra FoundationDelhi, Bengaluru and Varanasi airports, India

India launched DigiYatra, which lets passengers who sign up walk through airport gates by face instead of showing papers. The government says it is voluntary and privacy-focused; digital-rights groups have raised concerns about consent, data sharing and transparency.

Why it mattered. It shows both the convenience of face recognition and why clear rules on consent and data matter.

2023

April 2023

Segment Anything

Alexander Kirillov and colleaguesMeta AI, Menlo Park, United States

Meta released a model that draws a mask around almost any object you point at, trained on 11 million images with more than 1.1 billion masks.

Why it mattered. Segmentation went from a narrow tool to something anyone could use on any picture.

2024

The EU's AI Act

European UnionBrussels, Belgium

The EU's AI Act became law in 2024. Among other rules, it bans real-time face recognition by police in public places except in narrow cases, such as searching for a victim of kidnapping.

Why it mattered. It is one of the first big laws to draw lines around what computer vision may be used for.

By the numbers

How many pictures famous vision datasets hold

Datasets grew from tens of thousands of labelled digits to billions of pictures with captions scraped from the web.

1,00010,000100,0001,000,00010,000,000100,000,0001,000,000,00010,000,000,000 20002005201020152020 1998: MNIST handwritten digits: 70,00019982004: Caltech-101: 9,146 pictures in 101 categories20042009: ImageNet at launch: about 3.2 million (later over 14 million)20092014: COCO, with boxes and masks: about 328,00020142021: CLIP: about 400 million picture–caption pairs (reported)20212022: LAION-5B: 5.85 billion picture–caption pairs2023: Segment Anything: 11 million pictures, 1.1 billion masks2023
  1. 1998 MNIST handwritten digits: 70,000
  2. 2004 Caltech-101: 9,146 pictures in 101 categories
  3. 2009 ImageNet at launch: about 3.2 million (later over 14 million)
  4. 2014 COCO, with boxes and masks: about 328,000
  5. 2021 CLIP: about 400 million picture–caption pairs (reported)
  6. 2022 LAION-5B: 5.85 billion picture–caption pairs
  7. 2023 Segment Anything: 11 million pictures, 1.1 billion masks

Did you know?

The first digital photo, in 1957, was 176 × 176 pixels, about 31,000. A cheap phone camera today has more than 10 million.

In 1966 MIT planned to build a large part of a vision system as a summer project for students. It took the whole field more than 40 years.

Scanning a UPI QR code needs no neural network: its three corner squares always show stripes in the ratio 1:1:3:1:1, from any angle.

A change of just a few steps out of 255 to every pixel, chosen with the network's own maths, can make it name the wrong object.

ImageNet's millions of labels were added by thousands of online workers, who checked the pictures one by one.

The people

Who figured it out

Russell Kirsch

1929 – 2020 · Computer scientist · United States

Led the team that scanned the first digital photo, and is often credited with the idea of the pixel.

David Hubel

1926 – 2013 · Neuroscientist · Canada and United States

With Torsten Wiesel, found the edge-detecting cells of the visual cortex.

Torsten Wiesel

born 1924 · Neuroscientist · Sweden

Shared the 1981 Nobel Prize with Hubel for decoding how the brain's visual system works.

Kunihiko Fukushima

born 1936 · Computer scientist · Japan

Invented the Neocognitron, the forerunner of convolutional networks.

Yann LeCun

born 1960 · Computer scientist · France

Made convolutional networks learn with backpropagation and read real handwriting.

Fei-Fei Li

born 1976 · Computer scientist · China and United States

Led ImageNet, the dataset and contest that powered the deep learning boom in vision.

Kaiming He

born c. 1984 · Computer scientist · China

Lead author of ResNet, whose shortcut connections made very deep networks trainable.

Joy Buolamwini

born 1990 · Computer scientist and activist · Canada, Ghana and United States

Showed with Gender Shades how face analysis failed darker-skinned women most, and founded the Algorithmic Justice League.

Timnit Gebru

born c. 1983 · Computer scientist · Ethiopia and United States

Co-author of Gender Shades and a leading voice on bias and accountability in AI.

Where it happened

10 places, one idea

Sources

Where this comes from

Dates marked “c.” are approximate, and historians sometimes disagree about who was first. If you spot a mistake, tell us.

  1. First Digital Image NIST
  2. Russell Kirsch Wikipedia
  3. Receptive fields of single neurones in the cat's striate cortex (Hubel and Wiesel, 1959) The Journal of Physiology
  4. David H. Hubel Wikipedia
  5. Lawrence Roberts publishes "Machine Perception of Three-Dimensional Solids" History of Information
  6. Larry Roberts (computer scientist) Wikipedia
  7. The Summer Vision Project (AI Memo 100, 1966) MIT CSAIL
  8. History of computer vision Wikipedia
  9. The Nobel Prize in Physics 2009 NobelPrize.org
  10. Charge-coupled device Wikipedia
  11. Neocognitron (Fukushima, 1980) Biological Cybernetics, Springer
  12. Neocognitron Wikipedia
  13. The Nobel Prize in Physiology or Medicine 1981 NobelPrize.org
  14. Torsten Wiesel Wikipedia
  15. David Marr (neuroscientist) Wikipedia
  16. Vision (Marr) MIT Press
  17. Canny edge detector Wikipedia
  18. A Computational Approach to Edge Detection (Canny, 1986) IEEE Transactions on Pattern Analysis and Machine Intelligence
  19. Harris corner detector Wikipedia
  20. A Combined Corner and Edge Detector (Harris and Stephens, 1988) Alvey Vision Conference, BMVA
  21. LeNet Wikipedia
  22. Yann LeCun Wikipedia
  23. QR code Wikipedia
  24. History of QR Code Denso Wave
  25. Gradient-based learning applied to document recognition (LeCun et al., 1998) Proceedings of the IEEE
  26. Scale-invariant feature transform Wikipedia
  27. Object recognition from local scale-invariant features (Lowe, 1999) IEEE ICCV
  28. Viola–Jones object detection framework Wikipedia
  29. Rapid object detection using a boosted cascade of simple features (Viola and Jones, 2001) IEEE CVPR
  30. Histogram of oriented gradients Wikipedia
  31. Histograms of oriented gradients for human detection (Dalal and Triggs, 2005) IEEE CVPR
  32. ImageNet: A large-scale hierarchical image database (Deng et al., 2009) IEEE CVPR
  33. ImageNet Wikipedia
  34. ImageNet Classification with Deep Convolutional Neural Networks (Krizhevsky et al., 2012) NeurIPS
  35. AlexNet Wikipedia
  36. ILSVRC 2012 results ImageNet
  37. DeepFace: Closing the Gap to Human-Level Performance in Face Verification (CVPR 2014) CVF Open Access
  38. DeepFace Wikipedia
  39. Deep Residual Learning for Image Recognition (He et al., 2015) arXiv
  40. Residual neural network Wikipedia
  41. U-Net: Convolutional Networks for Biomedical Image Segmentation arXiv
  42. U-Net Wikipedia
  43. You Only Look Once: Unified, Real-Time Object Detection arXiv
  44. You Only Look Once Wikipedia
  45. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification Proceedings of Machine Learning Research 81
  46. Joy Buolamwini Wikipedia
  47. Performance of a Deep-Learning Algorithm vs Manual Grading for Detecting Diabetic Retinopathy in India JAMA Ophthalmology (PMC)
  48. Face Recognition Vendor Test Part 3: Demographic Effects (NISTIR 8280) NIST
  49. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale arXiv
  50. Vision transformer Wikipedia
  51. Learning Transferable Visual Models From Natural Language Supervision arXiv
  52. Contrastive Language-Image Pre-training Wikipedia
  53. DigiYatra app launch: facial recognition installed at three airports Business Today
  54. Digi Yatra Wikipedia
  55. DigiYatra Internet Freedom Foundation
  56. Segment Anything Meta AI
  57. SA-1B Dataset Meta AI
  58. Regulation (EU) 2024/1689 (Artificial Intelligence Act) EUR-Lex
  59. Artificial Intelligence Act Wikipedia
  60. ImageNet Large Scale Visual Recognition Challenge (Russakovsky et al., 2015) International Journal of Computer Vision (arXiv)
  61. Microsoft COCO: Common Objects in Context (Lin et al., 2014) arXiv
  62. LAION-5B: An open large-scale dataset for training next generation image-text models arXiv
  63. MNIST database Wikipedia
  64. Caltech 101 Wikipedia

That's the history. Now see how it works.