Chapter 1
Train once, answer millions of times
Training writes the numbers. Inference just uses them.
An AI model has two lives. First it is trained: shown example after example, its weights nudged a little each time (NeuralNetClear shows this step by step, and LLMClear shows it for chatbots). That is slow and costly, and it happens once. Then the weights are frozen, and the model is put to work. Every time someone asks it something, it runs its numbers forwards once to give an answer. That is inference.
Think of a recipe. Writing and testing the recipe takes weeks, but it is written once. Cooking it takes minutes, and it is cooked every day, in thousands of kitchens. Training writes the recipe. Inference is the cooking.
The network on the left is real. It learned the two spiral arms on the right when this page opened, then its weights were locked. Now it just answers: each glowing dot is a new question, "which arm is this point on?". To answer, it does 304 multiply-adds, one for each weight, and nothing else. It never learns from your questions.
Here is the twist. Training this net cost about 192 million multiply-adds, which sounds huge next to 304 per answer. But answers add up. After about 630,000 questions, the answering has cost more than the learning. For a popular chatbot used by millions, most of the lifetime computing and energy goes on inference, not training.


