How does image recognition work?

To a computer, a mango is 2,304 numbers. Here a tiny camera renders one, hand-made filters find its edges, and a real neural network learns to tell mangoes from bananas, chillies and coins right in your browser. Then it finds them in a scene, draws boxes and masks, and gets fooled by invisible noise.

VisionClearOpened 21 Sept 202620 min to playFree · no sign-up

In 60 seconds

  1. A picture is only numbers

    A camera sensor turns light into a grid of pixels, each with a red, green and blue value from 0 to 255. A 32 × 24 picture is 2,304 numbers. Turn the object, move the camera or change the light and most of those numbers change, though it is the same mango. Recognising means finding what stays the same.

  2. Filters find edges, layers find parts

    A 3 × 3 grid of weights slid over the picture, a convolution, lights up where brightness jumps: edges. Edges two ways at once are corners; busy patterns are texture. Max pooling keeps the biggest number in each 2 × 2 block. A CNN stacks learned filters, so early layers find edges and colours and later layers combine them into parts.

  3. Learning from labelled pictures

    A small CNN with 21,081 weights trains in your browser on 1,200 made-up pictures of a mango, banana, chilli, coin or nothing, and is tested on 300 wild ones it never saw. Data augmentation, randomly turning, flipping and re-lighting each picture, stops it from memorising and lifts test accuracy. A confusion matrix shows what it mixes up.

  4. Where, not just what

    Slide the recogniser over 897 windows of a scene and keep the confident ones. Non-max suppression keeps the best box and deletes duplicates that overlap it, measured by IoU, overlap divided by union. A second small network labels every pixel, turning boxes into masks. Detectors like YOLO do it in one pass, fast enough for video.

  5. It fails in strange ways

    A tiny change to every pixel, worked out from the network's own gradient, can flip its answer. A network trained with every fruit on its own background learns the background instead. Dim light outside its training range lowers accuracy. The 2018 Gender Shades study found face-analysis errors of up to 34.7% for darker-skinned women against at most 0.8% for lighter-skinned men.

  6. Everywhere, with limits

    Image recognition searches photos, helps screen for diabetic eye disease in India, suggests crop diseases, reads number plates and checks factory parts. Reading a UPI QR code mostly uses classic geometry with no learning. Vision transformers and vision-language models are next, but they can be confidently wrong, and face recognition raises real privacy questions.

The history

About 70 years from the first digital photo to machines that look and talk.

Read the full history
  1. 1957The first digital photo
  2. 1966Vision as a summer project
  3. 1989A network reads zip codes
  4. 2009ImageNet
  5. 2015ResNet: 152 layers
  6. 2018Gender Shades

The full explanation

VisionClear, chapter by chapter

Chapter 1

To a computer, a picture is just numbers

A camera turns a mango into a grid of red, green and blue values. That grid is all the computer gets.

You look at this mango and just see a mango. A computer never does. A camera's sensor (CameraClear shows how) measures light in a grid of tiny squares called pixels, and each pixel becomes three numbers: how much red, green and blue light arrived, each from 0 to 255.

So this 32 × 24 picture is 2,304 numbers in a row. A phone photo of 12 megapixels is about 36 million. Somewhere in those numbers is a mango, but no single number says "mango".

The picture on the right is made live by a tiny renderer in this page: for every pixel it shoots rays from the camera into the 3D scene on the left, finds what they hit, and works out the light. Click any pixel to read its numbers.

Now the hard part. Turn the mango, move the camera or dim the lamp, and most of the numbers change. A different angle, size or light gives a completely different grid, for the same mango. Recognising things means finding what stays the same while all the numbers move. Your eyes and brain do this without effort (EyeClear and BrainClear show how). Computers had to learn it.

Try “Pixels are numbers” in the interactive model →

Chapter 2

Filters find edges, then shapes

A small grid of weights slides over the picture. Stack layers of them and the network sees parts, then objects.

If a picture is just numbers, how do you find a mango in them? Start small. Look at a 3 × 3 patch, multiply each pixel by a weight, and add up. This little grid of weights is a filter (or kernel), and sliding it across the whole picture is called convolution. NeuralNetClear shows it on a digit; here it runs on a colour scene.

The right weights make the sum big only at an edge, where brightness jumps. Combine an up-down and a side-to-side edge filter and you get outlines. Where edges run two ways at once, you have a corner, like the tip of a chilli. Other filters answer to texture: busy, fine patterns. Each answer, spot by spot, makes a new picture called a feature map.

Pooling then shrinks each map: keep only the biggest number in every 2 × 2 block. Four numbers become one, so the next layer looks at a wider area, and a small shift of the object hardly changes the answer.

A convolutional neural network (CNN) stacks these steps, but nobody sets its weights by hand: it learns them from examples. Open "Inside the trained CNN" to see the network this box trains in your browser. Its first layer learns edge and colour filters on its own. Its second layer combines them into bigger parts. Big networks such as ResNet repeat this 50 or more times, climbing from edges to textures to parts to whole objects. The idea was loosely inspired by cells in the cat's visual cortex (see BrainClear), but a CNN is maths, not a brain.

Try “Filters and features” in the interactive model →

Chapter 3

Train a recogniser, live

Show a network 1,200 labelled pictures, again and again, and it learns to tell a mango from a chilli.

Now let's teach a network to recognise things. When you opened this box, the page drew 1,200 made-up pictures, 240 each of a mango, a banana, a chilli, a coin and nothing, each 20 × 20 pixels with its own colour, size, angle and background. Each one comes with its right answer, a label.

Training works as in NeuralNetClear: show a batch of 16 pictures, measure how wrong the answers are, and nudge all 21,081 weights a little to be less wrong (backpropagation). One pass through all 1,200 is an epoch. This network gets 8.

The real test is pictures it has never seen. Every half pass, it is tested on 300 "wild" pictures: turned any way, in dim or bright light, with colour casts and more noise. The confusion matrix shows what it mixes up: each row is the true answer, each column its guess.

Data augmentation is a cheap trick with a big effect. Each time a training picture is used, it is randomly turned, flipped, zoomed, shifted, re-lit and made noisy. The network never sees exactly the same picture twice, so it can't just memorise. Turn augmentation off and train again: training accuracy shoots towards 100%, but on wild pictures it does clearly worse. It learned the pictures, not the fruit.

Try “Train a recogniser” in the interactive model →

Chapter 4

Where is it? Boxes, then masks

Slide the recogniser across a whole scene, keep the confident boxes, remove the duplicates, then colour every pixel.

A recogniser says what is in a picture. A detector must also say where, for every object, even when there are several. The simplest way: slide the recogniser across the scene like a magnifying glass. This one checks 897 windows of 20 × 20 pixels, each 2 pixels from the last, and each window guesses a class, a confidence and a tight box.

Keep only windows above a confidence threshold. Here each box floats above the picture at a height equal to its confidence, and the glass sheet is the threshold. Lower it and you find more objects, but also more junk.

One mango sets off many overlapping windows. Non-max suppression (NMS) cleans up: take the most confident box, delete every box that overlaps it too much, and repeat. Overlap is measured by IoU, intersection over union: the shared area divided by the total area covered. 1 means identical, 0 means no overlap. A detection usually counts as right if its IoU with the true box is at least 0.5.

Boxes are rough. Segmentation labels every single pixel instead: a second small network here scores each pixel "object or not", and each detection paints its own object's pixels. Modern detectors such as YOLO look at the whole picture once instead of sliding, which is how they run at video speed in cars and phones. This little detector trained in your browser in seconds, so it misses some objects and raises some false alarms: the readout counts both honestly.

Try “Detection and masks” in the interactive model →

Chapter 5

How image recognition fails

Invisible noise, lazy shortcuts, bad light and unbalanced data. Knowing the failures is part of knowing the tool.

A network that scores 90% in a test can still fail in strange ways, because it never "understands" a mango. It found number patterns that happened to work. Here are four ways that goes wrong, three of them run live on this box's own networks.

Adversarial noise. Using the network's own maths, you can work out a tiny change for every pixel that pushes the answer the wrong way. Each number moves by only a few steps out of 255, too little for you to notice, yet the label flips. Random noise of the same size usually does nothing. Researchers have shown the same with stickers on real stop signs.

Shortcuts. A second network here was trained where every banana sat on green and every mango on blue. It scored almost 100%. Swap the backgrounds and it follows the colour behind the fruit, because that was the easiest pattern. The saliency map shows which pixels its answer depends on. Real models have done this too: one reading chest X-rays was found to pick up clues about which hospital took the picture, which was linked to how sick patients there tended to be.

Light and data bias. A network only knows the world its training pictures showed. Dim the light far below anything it trained on and accuracy falls. The same thing happens with people. The 2018 Gender Shades study tested three commercial face-analysis systems: they got the gender of lighter-skinned men wrong at most 0.8% of the time, but of darker-skinned women up to 34.7%. Much of the gap came from unbalanced training data. The companies improved their systems after the study, and a large 2019 test by the US standards agency NIST found error gaps between groups in many (not all) algorithms.

Privacy and surveillance. Face recognition can find missing children, unlock phones and speed up airport queues, like India's DigiYatra. It can also track people without their consent, and a false match can wrongly accuse someone. Countries are drawing different lines: the EU's AI Act bans most live face recognition by police in public places, while others are expanding its use. Where to draw the line is a debate for everyone, not only engineers.

Try “When it fails” in the interactive model →

Chapter 6

Where computers see, and what comes next

From photo search to eye screening, farms, roads, factories and UPI payments. Plus the models that look and talk.

The same ideas, pixels in, filters, layers, an answer out, run in many places around you. Pick one.

Phone photo search. Type "mango" in a gallery app and a network like this box's finds the pictures, without anyone having labelled them. The search here is live: this box's own network looks through 16 new pictures.

Eye screening. People with diabetes can slowly lose sight from diabetic retinopathy. In a 2019 study of 3,049 patients at Aravind Eye Hospital and Sankara Nethralaya in India, a network spotted cases that needed a doctor about as well as, or better than, human graders. It is a screening aid: doctors still decide.

Crops. Phone apps can suggest a leaf disease from a photo, and millions of farmers in India have reportedly tried them. But be careful: one well-known network scored 99% on tidy lab photos and only about 31% on photos taken elsewhere. Chapter 5's lesson again.

Roads and factories. Traffic cameras count vehicles and read number plates for fines. Factory cameras check every tablet, circuit board or coin for flaws, far faster than people.

UPI QR codes. Surprise: scanning a QR code mostly uses no learning. The scanner hunts for the three corner squares, whose stripes always come in the ratio 1:1:3:1:1, then straightens the grid and fixes errors with clever maths. Classic vision is still everywhere.

What's next. Vision transformers cut a picture into patches and treat them like words (see LLMClear). Vision-language models learn from hundreds of millions of pictures with captions, so you can ask a question about a photo and get an answer in words. They can also describe things wrongly and confidently, so check what matters. Video models (VideoModelClear) and running models fast on phones (InferenceClear) are the next steps.

Try “Where it's used” in the interactive model →

Test yourself

Frequently asked

What does a computer actually receive from a colour camera?

A grid of numbers: red, green and blue for every pixel. The sensor measures light in each pixel. The picture is just those numbers.

How many numbers are in a 32 × 24 colour picture?

2,304. 32 × 24 = 768 pixels, and each has 3 numbers: 768 × 3 = 2,304.

Why is recognising a mango hard for a computer?

Turning it, moving it or changing the light changes most of the numbers. The same mango can give completely different grids. The computer must find what stays the same.

What is a feature map?

The picture of a filter's answers at every spot. Slide a filter over the picture and write down its sum at every spot: that new picture is a feature map.

What does 2 × 2 max pooling do to a 48 × 32 map?

Makes it 24 × 16, keeping the biggest of each 2 × 2 block. Every 2 × 2 block becomes one number, its maximum, so width and height both halve.

Who chooses the filters in a CNN?

They are learned from examples during training. Training adjusts the filter weights so the network's answers get better. Edge filters appear on their own.

Why is the network tested on pictures it never trained on?

To check it learned the objects, not just memorised its examples. Only new pictures show whether it will work in the real world.

What does data augmentation do?

Randomly turns, flips, re-lights and adds noise to training pictures. The network never sees the same picture twice, so it must learn what stays the same.

In a confusion matrix, what do the numbers off the diagonal show?

Mistakes: pictures of one thing guessed as another. The diagonal is right answers. Everything else is a mix-up, like a banana called a chilli.

Two boxes share an overlap of 20 square pixels, and together cover 80 square pixels. What is their IoU?

0.25. IoU = intersection ÷ union = 20 ÷ 80 = 0.25.

What does non-max suppression remove?

Boxes that overlap a more confident box too much. It keeps the best box and deletes its near-duplicates, then repeats with the next best.

You lower the confidence threshold a lot. What usually happens?

More objects found, but more false alarms too. A lower bar lets in weak true detections and weak junk alike: a trade-off between misses and false alarms.

What makes adversarial noise different from random noise?

It is worked out from the network's own gradient to push the answer the wrong way. Each tiny change is chosen to hurt the answer most. Random changes of the same size mostly cancel out.

A model trained with every banana on green calls a banana on blue a "mango". Why?

It learned the background colour as a shortcut. In its training data the background alone gave the right answer, so it learned that instead of the shape.

What did the Gender Shades study find?

Error rates were far higher for darker-skinned women than for lighter-skinned men. Up to 34.7% error for darker-skinned women against at most 0.8% for lighter-skinned men, across three commercial systems.

In the 2019 Indian eye-screening study, what role did the network play?

A screening aid that flags who should see a doctor. It matched or beat human graders at spotting referable cases, but doctors still decide.

How does a phone mainly find a UPI QR code in a camera frame?

By looking for corner squares whose stripes run 1:1:3:1:1. The finder patterns give that ratio at any angle. Error-correcting maths then reads the code. No learning needed.

What does a vision transformer do first with a picture?

Cuts it into small patches and treats them like words. Each patch becomes a token, and attention lets every patch look at every other, as in language models.

Words worth knowing

Pixel
One tiny square of a picture, stored as red, green and blue numbers.
Convolution
Sliding a small grid of weights over a picture and adding up at every spot.
Feature map
The picture of a filter's answers: bright where it found its pattern.
Max pooling
Keeping the biggest number in each small block, which shrinks a map and tolerates small shifts.
CNN
A convolutional neural network: layers of learned filters and pooling, then an answer.
Data augmentation
Randomly turning, flipping, re-lighting and adding noise to training pictures so a network can't memorise them.
IoU
Intersection over union: how much two boxes overlap, from 0 to 1.
Non-max suppression
Keeping the most confident box and deleting the boxes that overlap it too much.
Adversarial example
A picture changed on purpose, often invisibly, so that a model gets it wrong.

Fork it. Teach with it.

This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/visionclear”.

git clone https://github.com/bdeeps/visionclear.git

Built with three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).

←→ previous / next box · / search

Would you like to see the full page, with the interactive model?