Chapter 1
Predict the next letter
A language model is a machine for guessing what comes next.
Type "the monsoon reaches ker" and you can guess the next letter. So can a machine. A language model is a machine that, given some text, gives a probability for every possible next piece. Pick one, add it to the text, and ask again. Do that a few thousand times and you have written a page.
This one is real and tiny. When the page loaded, it read a short text about the monsoon and chai (about 2,100 letters, written for this box) and simply counted: after "th", how often came "e"? how often "a"? Those counts become the bars. The engineer Claude Shannon played this game in 1948, building fake English from letter counts.
The context is how many letters back it looks. With 1 letter it writes gibberish. With 3 or 4 it writes real-looking words, but it starts to copy its training text word for word, because it has seen so little. Big LLMs look back thousands of tokens and learn patterns instead of storing counts.
Temperature changes how it picks. Near 0 it always takes the tallest bar: safe, but it loops. At 1 it follows the true odds. Above 1 the odds flatten and it gets wild. Top-k throws away all but the k best choices first. Chatbots use the same dials.


