How does AI inference work?

Training a model happens once; using it happens billions of times. Training writes a model's weights, slowly and once. Inference freezes them and runs the inputs forward to answer a question. Our spiral network cost 192 million multiply-adds to train and 304 per answer, so after about 630,000 questions answering has cost more than learning.

Training a model happens once; using it happens billions of times. Watch a frozen network answer questions, light up a grid of multiply-adds, see why a chatbot's speed is set by memory, shrink a model to 4 bits and break it at 2, and follow one question into a 120 kW rack and back.

InferenceClearOpened 20 Sept 202620 min to playFree · no sign-up

In 60 seconds

  1. Train once, answer forever

    Training writes a model's weights, slowly and once. Inference freezes them and runs the inputs forward to answer a question. Our spiral network cost 192 million multiply-adds to train and 304 per answer, so after about 630,000 questions answering has cost more than learning.

  2. It is all multiply-add

    A forward pass is mostly one step: multiply an input by a weight and add it to a total. A layer is a matrix times a vector. An LLM needs about 2 × parameters FLOPs per token, and GPUs, TPUs and phone NPUs do thousands of these sums at once.

  3. Memory is the real limit

    To write each token, a chatbot must read every weight from memory. So tokens per second ≈ memory bandwidth ÷ model size in bytes: about 200 for an 8B model in 16-bit on an H100. A growing KV cache adds more to read; batching many users shares the reads.

  4. Fewer bits, smaller models

    Quantisation rounds weights onto a few levels: FP32 → FP16 → INT8 → INT4. Size and read time fall; accuracy holds until too few bits are left, and our net breaks at 2 bits. Pruning and distillation shrink models too, which is how phones run AI.

  5. The journey of one question

    Your prompt travels to a data centre, is tokenised, then prefilled in one big pass. The answer is decoded one token at a time and streamed back. Racks draw around 120 kW and need liquid cooling; a median chatbot text prompt was reported at about 0.24 Wh.

  6. Phone or cloud

    On-device AI keeps data private, works offline and has no network wait, but the model must be small. The cloud runs far bigger models but needs a connection and servers. Each answer uses little energy, yet data centres used about 1.5% of the world's electricity in 2024.

The history

From motor-driven knobs in 1960 to data centres answering billions of questions a day: how we learned to run trained models fast, small and cheap.

Read the full history
  1. 1980XCON: expert rules at work in a factory
  2. 2009GPUs make deep learning practical
  3. 2016TPU: a chip built for inference
  4. 2023A chatbot model on a laptop
  5. 2023Speculative decoding: guess ahead, check in one go

The full explanation

InferenceClear, chapter by chapter

Chapter 1

Train once, answer millions of times

Training writes the numbers. Inference just uses them.

An AI model has two lives. First it is trained: shown example after example, its weights nudged a little each time (NeuralNetClear shows this step by step, and LLMClear shows it for chatbots). That is slow and costly, and it happens once. Then the weights are frozen, and the model is put to work. Every time someone asks it something, it runs its numbers forwards once to give an answer. That is inference.

Think of a recipe. Writing and testing the recipe takes weeks, but it is written once. Cooking it takes minutes, and it is cooked every day, in thousands of kitchens. Training writes the recipe. Inference is the cooking.

The network on the left is real. It learned the two spiral arms on the right when this page opened, then its weights were locked. Now it just answers: each glowing dot is a new question, "which arm is this point on?". To answer, it does 304 multiply-adds, one for each weight, and nothing else. It never learns from your questions.

Here is the twist. Training this net cost about 192 million multiply-adds, which sounds huge next to 304 per answer. But answers add up. After about 630,000 questions, the answering has cost more than the learning. For a popular chatbot used by millions, most of the lifetime computing and energy goes on inference, not training.

Try “Train vs infer” in the interactive model →

Chapter 2

Multiply and add, billions of times

A forward pass is mostly one kind of sum, done over and over.

Look inside any neural network while it answers and you find the same step, repeated: take an input, multiply it by a weight, and add the result to a running total. That is a multiply-add. A whole layer of them is a matrix times a vector.

The grid is the real middle layer of the chapter 1 network: 16 inputs across, 16 neurons down, 256 weights. The input column slides across; each cell multiplies its weight by the input above it (orange for plus, blue for minus) and the row's total grows on the right. One whole answer from that network costs 304 multiply-adds, or 608 FLOPs (a FLOP is one multiply or one add).

Big models just do far more of it. For a language model with N parameters, each new token costs about 2 × N FLOPs: every weight is multiplied once and added once. An 8-billion-parameter model needs about 16 billion FLOPs for every token it writes.

CPU or GPU? A CPU has a few very clever, very fast cores, built to do many different jobs one after another. A GPU has thousands of simpler cores that all do the same sum on different numbers at the same time. Since every cell in the grid is independent, the GPU can do the whole layer in one go. That is why AI runs on GPUs and on special chips like TPUs and phone NPUs, which are built almost entirely out of multiply-add units. (ComputerClear shows how a CPU steps through instructions.)

Try “Multiply-add” in the interactive model →

Chapter 3

The real bottleneck is memory

For every word it writes, a chatbot must read all of its weights.

A chatbot writes its answer one token at a time (LLMClear explains tokens). Each new token needs a full forward pass, and a full forward pass needs every weight. An 8-billion-parameter model stored in 16-bit numbers is 16 GB of weights, so for every single token the chip must read 16 GB from memory.

That makes memory bandwidth, how many bytes per second can flow from memory to the chip, the real speed limit. A rough rule: tokens per second ≈ bandwidth ÷ model size in bytes. An H100 GPU reads about 3,350 GB/s, so about 3,350 ÷ 16 ≈ 200 tokens/s at best for that model, and a bit less in practice. The sums themselves would take a fraction of that time. The chip spends most of its life waiting for data, like a super-fast cook with one slow helper fetching ingredients. (ComputerClear shows the same memory hierarchy in an ordinary computer.)

There is a second thing to read. To avoid redoing work, the model keeps notes about every earlier token in the conversation: the KV cache (the keys and values from attention). It grows with every word, so long chats get slower and use more memory.

The trick servers use is batching. Read the weights once, and use them for many users' next tokens at the same time. The bytes of weights are shared, so the total speed rises almost in step with the number of users, until the KV caches or the maths catch up.

Try “Memory wall” in the interactive model →

Chapter 4

Making models smaller: quantisation

Store each weight with fewer bits, and the model shrinks and speeds up.

Models are usually trained with each weight stored as a 32-bit or 16-bit decimal number. But the answers rarely need that much detail. Quantisation rounds every weight onto a small set of allowed values, so each one fits in fewer bits: FP32 → FP16 → INT8 → INT4. That is 4, 2, 1 and half a byte per weight.

Smaller weights help twice. The model takes less memory, and (as chapter 3 showed) fewer bytes to read means more tokens per second. An 8-billion-parameter model is 32 GB in FP32 but only about 4 GB in INT4.

The bars are the 256 real weights of our frozen network's middle layer. The flat sheets are the levels you are allowed to use. Watch the bars snap to them, then look at the map: at 8 bits nothing changes, at 4 bits it barely changes, but at 3 or 2 bits the spiral falls apart. Big models behave the same way, and careful methods (like GPTQ, and quantisation-aware training) push the safe limit down to about 4 bits.

Two more ways to shrink a model: pruning removes the smallest weights altogether (try the slider: our tiny net has little to spare, but big nets often have many near-useless weights), and distillation trains a small "student" model to copy a big "teacher" model's answers.

This is why your phone can run AI. A 3-billion-parameter model at about 4 bits needs roughly 1.5 GB, which fits beside your apps. Keyboard suggestions, photo search and offline speech all use small, quantised models on the phone's own chip (MobileClear shows the phone around it).

Try “Quantisation” in the interactive model →

Chapter 5

The journey of one question

Your prompt, a data centre, prefill, decode, and the words streaming back.

When you send a question to a chatbot, it travels to a data centre, a building full of computers. There it is cut into tokens, waits briefly in a queue, and is handed to a group of GPUs that hold the model's weights.

Then come two very different steps. Prefill: the whole prompt goes through the model at once. All its tokens are known, so this is one big matrix multiply that keeps the chips busy, and the KV cache is filled in. Decode: the answer is written one token at a time, each needing a full read of the weights (chapter 3). Each token is sent to you as soon as it is ready. That is why answers appear word by word: it is streaming, not typing for show.

So there are two numbers to watch. Time to first token (network + queue + prefill) is how long before anything appears. Tokens per second is how fast the rest arrives. A long prompt mostly adds to the first; a long answer mostly adds to the second.

The building behind it. A modern AI rack packs 72 GPUs and is reported to draw about 120 kW, as much as dozens of homes. Air can't carry that much heat away, so water-based liquid cooling runs pipes right onto the chips. A whole data centre adds cooling and power losses on top (CurrentClear explains the power side). The world's data centres used about 415 TWh in 2024, around 1.5% of all electricity (IEA), and AI is a fast-growing share.

Per question, the reported energy is small: Google reported 0.24 Wh for a median text prompt to its Gemini app (2025), about the same as a 10 W LED bulb for a minute and a half. But billions of questions a day add up, and long answers, pictures, video and "thinking" models use much more.

Try “Serving” in the interactive model →

Chapter 6

On your phone, or in the cloud?

Privacy, speed, cost and offline use: where the maths runs matters.

Inference can happen in two places. In the cloud, your request travels to a data centre (chapter 5) where big models run on big GPUs. On the device, a small model runs on your phone's own chip, often on a part built just for neural networks, the NPU. Your phone does this more than you might think: keyboard suggestions, finding "dog" in your photos, face unlock, and turning speech into text can all run without leaving the phone (MobileClear shows the phone's parts).

On the device: your data stays with you (better privacy), there is no network wait, it works offline, and it costs the company nothing per question. But the model must be small (chapter 4), and it uses your battery. In the cloud: much bigger, cleverer models, and your phone stays cool, but you need a connection, you wait for the round trip, your data leaves your phone, and someone pays for the servers and electricity.

So many apps mix both: quick, private jobs on the phone, and hard questions sent to the cloud.

In India this choice matters a lot. Many phones are budget models with less memory, signal can be patchy, and people speak many languages. The government's Bhashini platform (2022) offers translation and speech tools for India's 22 scheduled languages. In 2023 NPCI launched Hello! UPI, voice payments in Hindi and English, on a payments system that handled about 20 billion transactions in August 2025. Running AI at that scale, in every language, cheaply and reliably, is an open engineering challenge.

A balanced view. Each answer uses little energy, but there are billions of them. The IEA estimates data centres used about 1.5% of the world's electricity in 2024 and may more than double that by 2030, with AI a big part of the growth. Chips, smaller models and better cooling keep making each answer cheaper, yet demand keeps growing. Using AI where it genuinely helps, and small models where they are enough, is part of using it well.

Try “Device or cloud” in the interactive model →

Test yourself

Frequently asked

What happens to a model's weights during inference?

They stay frozen and are only read. Inference only uses the weights. Learning happens in training.

In the recipe picture, what is inference?

Cooking the dish, again and again. The recipe (training) is written once; the cooking (inference) happens every day.

Why can inference cost more than training over a model's life?

Each answer is tiny, but there are millions or billions of them. A tiny cost times a huge number of questions can pass the one-time training cost.

What is the basic step a neural network repeats while answering?

Multiply a weight by an input and add it to a total. Almost all the work of a forward pass is multiply-adds.

Roughly how many FLOPs does an LLM with 8 billion parameters need per token?

About 16 billion. About 2 × parameters: one multiply and one add per weight.

Why are GPUs good at this?

They have thousands of cores doing the same sum on different numbers at once. The multiply-adds in a layer are independent, so thousands of simple cores can do them together.

When a chatbot writes one token, what must it read from memory?

All of its weights (plus its KV cache). Every token needs a full forward pass, and that uses every weight.

Rough rule: tokens per second ≈ …

memory bandwidth ÷ model size in bytes. For one user, decoding is limited by how fast the weights can be read.

Why does batching many users help a server?

One read of the weights serves every user's next token. The weight bytes are shared across the batch, so throughput rises.

What does quantisation do?

Stores each weight with fewer bits by rounding it to a few allowed levels. Rounded weights fit in 8 or 4 bits instead of 16 or 32.

An 8-billion-parameter model in INT4 (half a byte per weight) is about…

4 GB. 8 billion × 0.5 bytes = 4 billion bytes, about 4 GB.

Why does a smaller model also answer faster?

Fewer bytes to read from memory for every token. Decoding is memory-bound, so fewer bytes means more tokens per second.

What happens in prefill?

The whole prompt is processed at once and the KV cache filled. Prefill takes all prompt tokens in one big pass. Decode then writes the answer.

Why do chatbot answers appear word by word?

Each token is made one at a time and sent as soon as it is ready. Decoding makes one token per step, and streaming sends each straight away.

Why do big AI racks use liquid cooling?

They make so much heat (around 100 kW a rack) that air can't carry it away. Liquids carry far more heat than air, which such dense racks need.

Which is a real advantage of running AI on your phone?

Your data can stay on the phone, and it works offline. On-device models keep data local and need no network, but they must be small.

Why can't your phone run the very biggest models?

They don't have enough memory or bandwidth for hundreds of GB of weights. Big models need far more memory and bandwidth than a phone has (chapters 3 and 4).

Which statement about AI energy use is fairest?

Each answer uses a little, but billions of answers add up to a noticeable share of electricity. Small per question, large in total, and growing: the IEA tracks it closely.

Words worth knowing

Inference
Using a trained model, with its weights frozen, to answer a new question.
Multiply-add (MAC)
Multiply one input by one weight and add it to a running total; two FLOPs.
FLOP
One floating-point operation, such as a single multiply or add.
Memory bandwidth
How many bytes per second can move from memory to the chip; it sets LLM decoding speed.
KV cache
Saved keys and values for every earlier token, so they need not be recomputed.
Batching
Serving many users' next tokens with one read of the weights.
Quantisation
Storing weights with fewer bits by rounding them onto a few allowed levels.
Prefill and decode
Processing the whole prompt at once, then writing the answer one token at a time.
On-device AI
Running a model on your own phone or laptop instead of in a data centre.

Fork it. Teach with it.

This box is plain HTML, CSS and JavaScript, with no build step and no accounts. Run it yourself and it sends nothing anywhere. The code is MIT. The words, images and videos are CC BY 4.0, so you can reuse them anywhere if you credit “Glassbox, glassbox.how/e/inferenceclear”.

git clone https://github.com/bdeeps/inferenceclear.git

Built with three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).

←→ previous / next box · / search

Would you like to see the full page, with the interactive model?