About 70 years from the first digital photo to machines that look and talk.
In 1957 a photo of a baby became the first grid of pixels in a computer. Researchers thought seeing would be easy to program, and learned over decades that it was one of the hardest problems in computing. Hand-made edge and corner finders gave way to networks that learn their own filters from millions of pictures, and now vision is joining language, bringing new uses and new questions about fairness and privacy.
Russell Kirsch and team, National Bureau of Standards, United States
1979 to 1980
Layered, convolution-style vision network
Kunihiko Fukushima's Neocognitron, Japan
1989
Convolutional network trained with backpropagation
Yann LeCun and colleagues, AT&T Bell Labs, United States
2001
Real-time face detection
Paul Viola and Michael Jones, United States
2012
Deep network wins ImageNet
AlexNet, University of Toronto, Canada
2015
Network over 100 layers wins ImageNet
ResNet, Microsoft Research Asia, China
1957Pixels and first ideas
1957 – 1979
Pixels and first ideas
Photos become numbers, brain scientists find edge-detecting cells, and early programs try to understand simple blocks. Seeing turns out to be hard.
1957
The first digital photo
Russell Kirsch and colleaguesNational Bureau of Standards, Washington DC, United States
Russell Kirsch's team built a drum scanner for SEAC, one of the first programmable computers in the US. The first picture they fed in was a photo of his three-month-old son, Walden, turned into a grid of 176 × 176 pixels.
Why it mattered. For the first time a photo became numbers inside a computer, the starting point of all image processing.
David Hubel and Torsten WieselJohns Hopkins University, Baltimore, United States
Recording from single brain cells in the visual cortex of cats, Hubel and Wiesel found cells that fired only for a line or edge at a particular angle and place. Later work showed more complex cells that combine them.
Why it mattered. The idea of simple edge detectors feeding more complex ones inspired the layered design of convolutional networks, loosely.
In his PhD thesis, Lawrence Roberts wrote a program that found the edges of simple blocks in a photo and worked out their 3D shapes. It only worked for plain, clean polyhedra, not messy real scenes.
Why it mattered. Often called the first computer vision thesis, it set the goal of turning pixels into an understanding of 3D objects.
Seymour Papert and studentsMIT Artificial Intelligence Group, Cambridge, United States
A memo from MIT planned a summer project for students to build a large part of a visual system, including telling objects from the background and naming them. It turned out to be far harder than anyone expected.
Why it mattered. It became a famous lesson: seeing feels effortless to people, but getting a computer to do it took decades.
Willard Boyle and George SmithBell Labs, Murray Hill, United States
Boyle and Smith invented the charge-coupled device (CCD), a chip that stores light as packets of electric charge and reads them out pixel by pixel. It became the eye of digital cameras and telescopes.
Why it mattered. Cheap image sensors made digital pictures, and so computer vision, part of everyday life. They shared the 2009 Nobel Prize in Physics.
People design edge, corner and feature detectors by hand, while the first convolutional networks quietly learn to read digits and cheques.
1980
1979 (Japanese), 1980 (English)
The Neocognitron
Kunihiko FukushimaNHK Broadcasting Science Research Laboratories, Tokyo, Japan
Inspired by Hubel and Wiesel, Kunihiko Fukushima built a network of layers: simple cells that find small features and complex cells that pool them, so a shape could be recognised even when shifted.
Why it mattered. It is the direct ancestor of the convolutional network, with filters and pooling, the design this box trains.
Hubel and Wiesel shared the Nobel Prize in Physiology or Medicine for their discoveries about how the visual system processes information, with Roger Sperry, who was honoured for separate work on the brain's two halves.
Why it mattered. Their map of the visual cortex is still the textbook starting point, and EyeClear and BrainClear show the living version.
David Marr's book, published after his early death in 1980, described vision as a series of steps: from edges to a sketch of surfaces to 3D objects. It shaped how researchers thought about vision for years.
Why it mattered. It gave computer vision a clear plan of stages, many of which learned networks later handled on their own.
John Canny worked out, with careful maths, what a good edge detector should do: find real edges, place them exactly, and mark each only once. His method is still built into image software today.
Why it mattered. Hand-designed edge finding was the first step in most vision systems for decades.
Chris Harris and Mike StephensPlessey Research, Roke Manor, United Kingdom
Harris and Stephens described a corner detector that looks for places where the picture changes in two directions at once. Corners are easy to find again in another photo, so they help match and track things.
Why it mattered. Chapter 2's corner filter uses their formula.
Yann LeCun and colleaguesAT&T Bell Labs, Holmdel, United States
Yann LeCun's team trained a convolutional network with backpropagation to read handwritten digits on US mail. The same small filters were shared across the whole picture, so it needed far fewer weights.
Why it mattered. It showed that a network could learn its own filters from examples, instead of people designing them.
Masahiro Hara and team at Denso WaveDenso Wave, Japan
Engineers at Denso Wave, a Toyota group company, designed the QR code to track car parts. Its three big corner squares can be found quickly from any angle, and built-in error correction fixes damaged parts.
Why it mattered. Decades later, QR codes carry UPI payments in India, and reading them still needs no neural network.
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick HaffnerAT&T Labs, United States
A long paper described LeNet-5 and a system that read the amounts on handwritten bank cheques. NCR put such readers to work from the mid-1990s, and by around 2000 they were reported to read roughly 10% of cheques in the US.
Why it mattered. It was one of the first times image recognition did everyday work at large scale.
David LoweUniversity of British Columbia, Vancouver, Canada
David Lowe's SIFT method found distinctive points in a picture and described each so it could be matched in another photo, even after turning, zooming or a change of light.
Why it mattered. For a decade, hand-made features like SIFT powered panorama stitching, object recognition and robot navigation.
Paul Viola and Michael JonesMitsubishi Electric Research Labs and Compaq, Cambridge, United States
Viola and Jones found faces with thousands of simple light-and-dark box features, chosen by a learning method and checked in a fast cascade. It ran at 15 frames a second on a 700 MHz Pentium III.
Why it mattered. It put face-finding boxes in digital cameras, and its sliding-window idea is the one chapter 4 uses.
Navneet Dalal and Bill TriggsINRIA, Grenoble, France
Dalal and Triggs described people by histograms of edge directions (HOG) in small cells, and slid a detector over the picture. It became a standard way to find pedestrians.
Why it mattered. It was among the best hand-made detectors just before learned features took over.
Huge labelled datasets and graphics chips let deep networks learn their own features. Accuracy leaps, and so do concerns about bias.
2009
June 2009
ImageNet
Fei-Fei Li, Jia Deng and colleaguesPrinceton University, United States
Fei-Fei Li's team presented ImageNet, millions of photos sorted into thousands of categories, labelled with help from online workers. From 2010 it ran a yearly contest with 1,000 categories.
Why it mattered. It gave learning systems enough varied examples to learn from, and a fair test to compare them.
Alex Krizhevsky, Ilya Sutskever and Geoffrey HintonUniversity of Toronto, Canada
A deep convolutional network trained on two gaming graphics cards won the ImageNet contest with a top-5 error of about 15.3%, far ahead of about 26.2% for the next team. It used augmentation and dropout to avoid memorising.
Why it mattered. This result convinced much of the field to switch from hand-made features to learned ones.
Yaniv Taigman and colleaguesFacebook AI Research, Menlo Park, United States
A deep network trained on about 4 million face photos from Facebook users matched pairs of faces with 97.35% accuracy on a standard test, close to the 97.53% people score.
Why it mattered. Face recognition jumped in accuracy, and so did questions about who gets recognised, and with whose photos.
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian SunMicrosoft Research Asia, Beijing, China
Very deep networks used to train badly. ResNet added shortcut connections that let each block learn only a small change, which made networks 152 layers deep trainable. It won ImageNet 2015 with a 3.57% top-5 error.
Why it mattered. Residual connections are now in almost every large network, including language models.
Olaf Ronneberger, Philipp Fischer and Thomas BroxUniversity of Freiburg, Germany
U-Net shrinks a picture through convolution layers and then grows it back, labelling every pixel. It was built to outline cells in microscope images from only a few labelled examples.
Why it mattered. It became a standard tool for segmentation, especially in medicine.
Joseph Redmon, Santosh Divvala, Ross Girshick and Ali FarhadiUniversity of Washington and Allen Institute for AI, Seattle, United States
Instead of sliding a window, YOLO splits the picture into a grid and predicts all boxes and classes in one pass of one network. It ran at about 45 frames a second.
Why it mattered. Fast single-pass detectors made real-time detection practical in cars, drones, phones and cameras.
Joy Buolamwini and Timnit GebruMIT Media Lab, Cambridge, United States
Testing three commercial face-analysis systems on 1,270 faces of parliament members from six countries, they found gender errors of at most 0.8% for lighter-skinned men but up to 34.7% for darker-skinned women.
Why it mattered. It showed that unbalanced data can make a system fail most for the people least represented, and pushed companies to improve.
Varun Gulshan and colleagues, with Aravind and Sankara NethralayaMadurai and Chennai, India
A deep-learning system for diabetic retinopathy was tested on 3,049 patients at two Indian eye hospitals. It spotted cases needing a doctor about as well as, or better than, trained graders and eye specialists.
Why it mattered. Screening aids like this could help reach millions of people with diabetes where eye doctors are scarce.
Patrick Grother, Mei Ngan and Kayee HanaokaNIST, Gaithersburg, United States
The US standards agency tested 189 face recognition algorithms. False matches were often 10 to more than 100 times more common for some demographic groups than others, though the size of the gap depended strongly on the algorithm.
Why it mattered. It gave governments and buyers hard numbers on how accurate and how fair different systems are.
How often the winning program's five best guesses all missed the right label. Deep networks arrived in 2012, and the error fell fast.
2010 NEC and UIUC, hand-made features: 28.2%
2011 XRCE, hand-made features: 25.8%
2012 AlexNet: 15.3%
2013 Clarifai: 11.2%
2014 GoogLeNet: 6.7%
2015 ResNet: 3.57%
2016 Trimps-Soushen: about 2.99%
2017 SENet: about 2.25%
2020 – today
Vision meets language
Transformers and caption-trained models link pictures with words, while countries debate face recognition and privacy.
2020
October 2020
Vision Transformer
Alexey Dosovitskiy and colleaguesGoogle Research, Berlin, Germany
The paper 'An Image is Worth 16x16 Words' cut pictures into patches and fed them to a transformer, the design behind language models. With enough training data, it matched or beat the best CNNs.
Why it mattered. Vision and language models began to share one design.
Alec Radford and colleaguesOpenAI, San Francisco, United States
CLIP learned from about 400 million pictures with captions from the web, matching each picture with its text. It could then recognise things it was never given labels for, just from a written description.
Why it mattered. It led to vision-language models that caption photos and answer questions about them.
Ministry of Civil Aviation and the Digi Yatra FoundationDelhi, Bengaluru and Varanasi airports, India
India launched DigiYatra, which lets passengers who sign up walk through airport gates by face instead of showing papers. The government says it is voluntary and privacy-focused; digital-rights groups have raised concerns about consent, data sharing and transparency.
Why it mattered. It shows both the convenience of face recognition and why clear rules on consent and data matter.
The EU's AI Act became law in 2024. Among other rules, it bans real-time face recognition by police in public places except in narrow cases, such as searching for a victim of kidnapping.
Why it mattered. It is one of the first big laws to draw lines around what computer vision may be used for.