How do neural networks work?

A neural network is just multiply, add and squash, done millions of times with weights it learns from examples. An artificial neuron multiplies each input by a weight, adds a bias, and passes the total through an activation function such as a step, a sigmoid or ReLU.

A neural network is just multiply, add and squash, done millions of times with weights it learns from examples. Here they really compute and train in your browser: tune a neuron by hand, watch a hidden layer bend a boundary, roll a ball down a loss valley until it flies out, draw a digit for a network trained on the spot, and catch a big network memorising noise.

NeuralNetClearOpened 18 Sept 202618 min to playFree · no sign-up

In 60 seconds

  1. One neuron: multiply, add, squash

    An artificial neuron multiplies each input by a weight, adds a bias, and passes the total through an activation function such as a step, a sigmoid or ReLU. On a 2D board its decision is one straight line. Rosenblatt's 1958 perceptron rule nudges the weights after every mistake until the line splits the classes. Real brain cells are far more complex: the name is a loose inspiration.

  2. Layers bend the line

    No single straight line can split the XOR pattern, so one neuron can't learn it. A hidden layer fixes that: each hidden neuron draws its own line and the output neuron combines them, so the boundary bends. A network with 2 inputs, h hidden neurons and 1 output has 4h + 1 weights and biases, its parameters.

  3. Learning is walking downhill

    The loss measures how wrong the network is. Its gradient says which way is uphill for every weight, and gradient descent steps the other way, by the learning rate. Backpropagation sends the output error backwards through the links to find every gradient in one sweep. Too big a learning rate overshoots the valley and the loss explodes.

  4. Pictures are numbers

    A 16 by 16 drawing is 256 pixel values. A 256 → 32 → 10 network, 8,554 parameters, trains in your browser on 1,200 made-up digits and gives a probability for each digit. Image networks slide small filters over the picture, a convolution, and the right 3 × 3 weights make an edge detector.

  5. Memorising is not learning

    A network too big for its data bends around every noisy point: training loss falls while test loss rises. More data, a smaller network, early stopping or regularisation help. Data also decides who a network works for: a group that is rare in the training data gets more mistakes, as studies of face analysis and speech recognition have found.

  6. Everywhere, with limits

    The same maths tags photos, types speech, translates Indian languages, screens eye scans with doctors, predicts traffic, and helps banks flag suspicious UPI payments, where catching more fraud also blocks more honest payments. Networks find patterns; they don't understand, they need lots of data and energy, and people must stay responsible.

The history

About 80 years from a neuron drawn on paper to networks that read, see and talk.

Read the full history
  1. 1943A neuron made of maths
  2. 1969Perceptrons, the book
  3. 1989Reading handwritten zip codes
  4. 2016AlphaGo beats Lee Sedol
  5. 2024The IndiaAI Mission

The full explanation

NeuralNetClear, chapter by chapter

Chapter 1

One artificial neuron

Multiply, add, squash: the tiny sum inside every neural network.

An artificial neuron is a tiny sum. It takes some numbers in (the inputs), multiplies each by its own weight, adds them up with one extra number called the bias, and passes the total through an activation function. That's it.

Here it judges mangoes. Input x₁ is how yellow the skin is, x₂ is how soft it feels. The total is z = w₁·x₁ + w₂·x₂ + b. A step activation says "ripe" (1) if z is above zero and "not ripe" (0) if not. A sigmoid gives a smooth score between 0 and 1. ReLU gives 0 for anything below zero and z itself above it.

On the board, every spot where z = 0 lies on one straight line. The weights tilt the line and the bias slides it. So one neuron can only split things with a straight cut.

The name comes from brain cells, which BrainClear and NervousClear show. But the likeness is loose: a real neuron is a living cell with thousands of connections, chemical signals and timing. This one is multiply, add, squash.

How does it find good weights? In 1958 Frank Rosenblatt's perceptron rule: show it one example; if its guess is wrong, nudge each weight by rate × error × input. Repeat. If one straight line can split the data, this is guaranteed to find one.

Try “One neuron” in the interactive model →

Chapter 2

Layers: why one neuron is not enough

Each hidden neuron draws a line. Together, their lines bend.

Try the XOR pattern: orange in two opposite corners, blue in the other two. No single straight line can split them. One neuron draws one line, so one neuron cannot learn XOR. In 1969 Marvin Minsky and Seymour Papert proved limits like this in a famous book, and interest in neural networks cooled for years.

The fix is a hidden layer: a row of neurons between the inputs and the output. Each hidden neuron looks at the same inputs and draws its own line (the small squares show what each one sees). The output neuron then adds up their answers with its own weights. Lines combined this way make a boundary that bends, and with enough hidden neurons it can wrap around almost any shape.

Working out the answer, layer by layer from inputs to output, is called the forward pass. Watch the little lights travel along the links: bigger lights carry more (weight × input).

Every weight and bias is a parameter: a number the network learns. Here, 2 inputs → h hidden → 1 output has 4h + 1 of them. Real image networks have millions; large language models (LLMClear, coming next) have billions.

Try “Layers” in the interactive model →

Chapter 3

Learning: rolling downhill

Measure the error, find which way is down, take a small step. Repeat thousands of times.

A new network starts with random weights, so its answers are nonsense. Learning means changing the weights, a little at a time, until the answers get good.

First we need a score for "how wrong": the loss. Here it is cross-entropy: near 0 when the network is confidently right, big when it is confidently wrong. Imagine the loss as a landscape: every possible set of weights is a spot on the map, and the loss is the height. Learning is walking downhill.

The gradient says which way is uphill for every weight at once. Gradient descent takes a step the other way. The step size is the learning rate. Too small and learning crawls. Too big and every step overshoots the valley floor, bounces higher up the far side, and the loss explodes.

How does the network find the gradient for a weight deep inside? Backpropagation. The error at the output is passed backwards through the same links, each neuron getting its share of the blame in proportion to its weights. That gives every weight its gradient in one backward sweep (the red lights). One pass through all the training data is an epoch.

This is just calculus done on a computer, very many times. There is no understanding inside, only numbers moving downhill.

Try “Learning” in the interactive model →

Chapter 4

Seeing images: pixels in, guesses out

Draw a digit. A network trained here, in your browser, has a go.

To a computer, a picture is just numbers: one brightness value per pixel (CameraClear shows how a camera sensor makes them). This drawing pad is 16 × 16 = 256 pixels, so the network has 256 inputs.

They feed 32 hidden neurons, which feed 10 outputs, one for each digit 0 to 9. A softmax turns the 10 outputs into probabilities that add up to 100%. That is 8,554 weights and biases.

Nobody wrote rules like "a 7 has a line across the top". When you opened this chapter, the page drew 1,200 made-up digits, each with a random tilt, stretch and pen width, and trained the network on them with backpropagation. Then it tested it on 300 new ones it had never seen. Your drawing is centred and scaled first, as the famous MNIST digits were.

Big image networks don't connect every pixel to every neuron. They slide small filters across the picture: a 3 × 3 grid of weights, multiplied with the pixels under it and added up, at every position. This is convolution. The right numbers make an edge detector. In a convolutional neural network (CNN), like Yann LeCun's LeNet, the filter numbers are learned too, and the first layer's filters often end up looking like edge detectors.

It only knows digits. Draw a cat and it will still name a digit, sometimes confidently. It isn't seeing a cat and lying: it has no idea what a cat is.

Try “Seeing images” in the interactive model →

Chapter 5

Overfitting, and why data matters

A network can memorise instead of learn, and it can only be as fair as its data.

A network that is too big for its data can memorise it: it bends its boundary around every single training point, including the ones that are just wrong or unlucky (noise). It then scores brilliantly on the examples it has seen and badly on new ones. This is overfitting.

That is why we always keep some test data back. Watch the two loss curves: training loss keeps falling, but when test loss turns and climbs, the network has stopped learning the pattern and started learning the noise.

Fixes: more data (harder to memorise 300 points than 30), a smaller network, stopping early, or regularisation, which gently pulls every weight towards zero so the boundary stays smooth.

Data also decides who a network works for. If one group is rare in the training data, the network learns the common group's pattern and makes more mistakes on the rare one. It is not being mean; it simply never saw enough examples. This really happens: a 2018 study found commercial face-analysis systems wrong up to 34.7% of the time for darker-skinned women, against 0.8% for lighter-skinned men, and a 2020 study found speech recognisers made about twice as many errors for Black speakers as for white speakers. In India, a voice app trained mostly on one accent or language can struggle with others.

The fix is not magic: collect data that represents everyone, measure accuracy for each group, and keep people responsible for decisions that matter.

Try “Overfitting and data” in the interactive model →

Chapter 6

Where they are used, and their limits

Pattern finders everywhere: useful, powerful, and not the same as understanding.

Every network in this box does the same thing: numbers in, numbers out, with weights learned from examples. Change what the numbers mean and you get very different tools.

Photos: pixels in, labels out (your gallery search). Speech: sound in, words out (voice typing). Translation: words in, words out; India's Bhashini project builds these for Indian languages. Medicine: a network that screens eye photos for diabetic damage was tested at Aravind and Sankara Nethralaya eye hospitals in India, and matched specialist graders; a doctor still decides. Maps: networks predict traffic and arrival times.

UPI payments: as reported to Parliament, NPCI gives banks an AI/ML-based fraud monitoring system that raises alerts and can decline suspicious payments, and the RBI's innovation hub built MuleHunter.AI (announced in 2024) to spot "mule" accounts used to move stolen money. The details are not public. Every such system trades catching fraud against blocking honest payments: slide the threshold and see.

Limits. A network finds patterns in its training data. It doesn't understand the world, and it can be confidently wrong on anything unlike what it saw. It needs lots of data, which can carry unfairness in with it. Big ones need lots of energy: training GPT-3 was estimated at about 1,287 MWh of electricity, and data centres (for all uses, AI included) took about 1.5% of the world's electricity in 2024, a share that is growing.

How to think about AI: treat it as a powerful, sometimes brilliant pattern-matcher, not an oracle or a mind. Ask what data it learned from and who might be missing. Check anything important. Keep a person responsible for decisions about people. Used that way it helps doctors, farmers, students and banks every day.

Next, LLMClear opens up the biggest networks of all: large language models that predict the next word.

Try “Where they are used” in the interactive model →

Test yourself

Frequently asked

What does one artificial neuron calculate?

Inputs times weights, plus a bias, passed through an activation function. Multiply each input by its weight, add them and the bias, then squash the total.

On a 2D board, what shape is one neuron's decision boundary?

A straight line. The boundary is where w₁·x₁ + w₂·x₂ + b = 0, which is the equation of a straight line.

What does the perceptron rule do after a wrong guess?

Nudges each weight by rate × error × input. The error says which way to move; the input says which weights were responsible.

Why can't a single neuron learn XOR?

Its boundary is one straight line, and no straight line splits opposite corners. One neuron always splits the board with a single straight line.

What does a hidden layer add?

Several lines that the output neuron combines into a bent boundary. Each hidden neuron draws its own line; the output mixes them.

How many parameters does a network with 2 inputs, 3 hidden neurons and 1 output have?

13. 2×3 weights + 3 biases + 3×1 weights + 1 bias = 13, that is 4h + 1 with h = 3.

What does the learning rate control?

How big each downhill step is. Each weight moves by learning rate × its gradient on every step.

What happens if the learning rate is far too high?

Steps overshoot the valley and the loss can explode. Every step jumps past the bottom and lands higher up the far side, so the loss grows.

What does backpropagation compute?

The gradient of the loss for every weight, working backwards from the output. It passes the output error back through the links to share out the blame.

How does a 16 × 16 picture go into the network?

As 256 numbers, one per pixel. Each pixel's brightness is one input.

Why test the network on digits it never trained on?

To check it learned the pattern, not just memorised its examples. Only new examples show whether it will work on your drawing.

What does a convolution filter do?

Slides a small grid of weights over the image, multiplying and adding at every spot. With the right weights, the sum is big wherever there is an edge.

What is overfitting?

Memorising the training examples, noise and all, so it fails on new ones. Training loss falls while test loss rises: it learned the examples, not the pattern.

Which of these helps against overfitting?

More data, a smaller network or regularisation. All three make it harder to bend the boundary around every noisy point.

A network is trained on data with very few examples from group B. What usually happens?

It makes more mistakes for group B. It mostly learns the common group's pattern, so the rare group gets more errors. Measure accuracy per group.

What do a photo tagger and a speech recogniser have in common?

Both turn numbers in into numbers out, with weights learned from examples. Same maths; the numbers just mean pixels in one case and sound in the other.

A fraud system lowers its threshold to catch more fraud. What else happens?

It blocks more honest payments by mistake. Catching more fraud and making fewer false alarms pull in opposite directions.

Which is a real limit of neural networks?

They can be confidently wrong on things unlike their training data. They match patterns from their data; they have no understanding to fall back on.

Words worth knowing

Artificial neuron
A tiny calculation: inputs times weights, plus a bias, passed through an activation function.
Weight
A learned number that says how much one input counts, and in which direction.
Activation function
The squash at the end of a neuron, such as a step, a sigmoid or ReLU.
Hidden layer
Neurons between the inputs and the output, whose lines combine into a bent boundary.
Loss
One number that says how wrong the network is on its training examples.
Gradient descent
Moving every weight a small step downhill on the loss, again and again.
Learning rate
The size of each step; too big and the loss explodes.
Backpropagation
Passing the output error backwards through the network to find every weight's gradient.
Convolution
Sliding a small grid of weights over an image to find patterns such as edges.
Overfitting
Memorising the training examples, noise and all, instead of the general pattern.

Fork it. Teach with it.

This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/neuralnetclear”.

git clone https://github.com/bdeeps/neuralnetclear.git

Built with three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).

←→ previous / next box · / search

Would you like to see the full page, with the interactive model?