How does AI inference work?

Training a model happens once; using it happens billions of times. Training writes a model's weights, slowly and once. Inference freezes them and runs the inputs forward to answer a question. Our spiral network cost 192 million multiply-adds to train and 304 per answer, so after about 630,000 questions answering has cost more than learning.

Read how it works: How does AI inference work? · The history of AI inference · More boxes on Glassbox

Chapters in this interactive model

  1. Train once, answer millions of times: Training writes the numbers. Inference just uses them.
  2. Multiply and add, billions of times: A forward pass is mostly one kind of sum, done over and over.
  3. The real bottleneck is memory: For every word it writes, a chatbot must read all of its weights.
  4. Making models smaller: quantisation: Store each weight with fewer bits, and the model shrinks and speeds up.
  5. The journey of one question: Your prompt, a data centre, prefill, decode, and the words streaming back.
  6. On your phone, or in the cloud?: Privacy, speed, cost and offline use: where the maths runs matters.
Glassbox InferenceClear

Drag to orbit · scroll or pinch to zoom