How does image recognition work?

To a computer, a mango is 2,304 numbers. Here a tiny camera renders one, hand-made filters find its edges, and a real neural network learns to tell mangoes from bananas, chillies and coins right in your browser. Then it finds them in a scene, draws boxes and masks, and gets fooled by invisible noise.

Read how it works: How does image recognition work? · The history of image recognition · More boxes on Glassbox

Chapters in this interactive model

  1. To a computer, a picture is just numbers: A camera turns a mango into a grid of red, green and blue values. That grid is all the computer gets.
  2. Filters find edges, then shapes: A small grid of weights slides over the picture. Stack layers of them and the network sees parts, then objects.
  3. Train a recogniser, live: Show a network 1,200 labelled pictures, again and again, and it learns to tell a mango from a chilli.
  4. Where is it? Boxes, then masks: Slide the recogniser across a whole scene, keep the confident boxes, remove the duplicates, then colour every pixel.
  5. How image recognition fails: Invisible noise, lazy shortcuts, bad light and unbalanced data. Knowing the failures is part of knowing the tool.
  6. Where computers see, and what comes next: From photo search to eye screening, farms, roads, factories and UPI payments. Plus the models that look and talk.
Glassbox VisionClear

Drag to orbit · scroll or pinch to zoom