How does image recognition work?
To a computer, a mango is 2,304 numbers. Here a tiny camera renders one, hand-made filters find its edges, and a real neural network learns to tell mangoes from bananas, chillies and coins right in your browser. Then it finds them in a scene, draws boxes and masks, and gets fooled by invisible noise.
Read how it works: How does image recognition work? · The history of image recognition · More boxes on Glassbox
Chapters in this interactive model
- To a computer, a picture is just numbers: A camera turns a mango into a grid of red, green and blue values. That grid is all the computer gets.
- Filters find edges, then shapes: A small grid of weights slides over the picture. Stack layers of them and the network sees parts, then objects.
- Train a recogniser, live: Show a network 1,200 labelled pictures, again and again, and it learns to tell a mango from a chilli.
- Where is it? Boxes, then masks: Slide the recogniser across a whole scene, keep the confident boxes, remove the duplicates, then colour every pixel.
- How image recognition fails: Invisible noise, lazy shortcuts, bad light and unbalanced data. Knowing the failures is part of knowing the tool.
- Where computers see, and what comes next: From photo search to eye screening, farms, roads, factories and UPI payments. Plus the models that look and talk.
This interactive model needs JavaScript and WebGL.