From Markov counting letters in a poem in 1913 to chatbots used by hundreds of millions: a century of predicting the next word.
In 1913 Andrey Markov counted which letters followed which in a Russian poem, and in 1948 Claude Shannon used such counts to generate fake English. For decades the models stayed small. Then words became vectors, attention let networks look back, and the 2017 transformer could be scaled up almost without limit, until ChatGPT put an LLM in everyone's pocket.
Neural network language model with learned word vectors
Yoshua Bengio and colleagues, Canada
2017
Transformer
Vaswani and colleagues, Google
2023
Hindi LLM from an Indian startup
OpenHathi, Sarvam AI with AI4Bharat, India
2024
Broad law regulating AI
EU AI Act, European Union
23 January 1913Counting letters and words
1913 – 1989
Counting letters and words
Mathematicians count which letters and words follow which, and the first chatbot fakes a conversation with simple rules.
1913
23 January 1913
Counting letters in a poem
Andrey MarkovSt Petersburg, Russia
Markov went through the first 20,000 letters of Pushkin's poem Eugene Onegin, counting how often a vowel followed a vowel or a consonant. He showed that each letter's chance depends on the one before it.
Why it mattered. It was the first chain of probabilities built from real text: the seed of every language model.
Claude ShannonBell Labs, Murray Hill, United States
In the paper that founded information theory, Shannon generated text by picking each letter or word according to how often it followed the previous ones. With longer contexts, the nonsense looked more and more like English.
Why it mattered. It is exactly the idea in chapter 1: a model that predicts the next piece from counts.
Claude ShannonBell Labs, Murray Hill, United States
Shannon asked people to guess a hidden text one letter at a time. From how often they guessed right, he estimated that English carries only about 1 bit of information per letter, because so much of it is predictable.
Why it mattered. It showed how much of language can be predicted from what came before.
ELIZA matched keywords in what you typed and replied from simple scripts, most famously as a therapist who turned your words into questions. It learned nothing. Yet many users felt understood, which alarmed its creator.
Why it mattered. The "ELIZA effect", trusting fluent text too much, is still the key caution with LLMs.
Karen Spärck JonesUniversity of Cambridge, United Kingdom
Spärck Jones proposed giving rare words more weight than common ones when matching documents to a search, an idea now called inverse document frequency.
Why it mattered. Her statistical view of words underpins search engines and fed into later word vectors.
Translation learns from millions of example sentences, and neural networks start predicting the next word.
1990
June 1990
Translation learned from examples
Peter Brown, Robert Mercer and colleagues at IBMYorktown Heights, United States
IBM's team translated French to English using statistics learned from Canadian parliament records in both languages, with no hand-written grammar. A language model scored how natural each English sentence was.
Why it mattered. Statistical translation proved that learning from lots of text could beat rules.
Sepp Hochreiter and Jürgen SchmidhuberMunich, Germany
The long short-term memory (LSTM) network added gates that let a network keep information over many steps. Two decades later, LSTMs powered speech recognition and translation on phones.
Why it mattered. It let networks read sequences, the job transformers later took over.
Yoshua Bengio and colleaguesUniversité de Montréal, Canada
Bengio's team trained a neural network to predict the next word, learning a vector for every word along the way. Similar words got similar vectors, so the model could handle sentences it had never seen.
Why it mattered. Next-word prediction plus learned word vectors is the recipe LLMs still follow.
Google Translate launched using statistical machine translation, trained on huge collections of translated documents from the United Nations and the European Parliament. It moved to neural networks in 2016.
Why it mattered. Millions of people began using language models every day without knowing it.
Words become points in space, networks learn to translate whole sentences, and attention lets them look back at the source.
2013
January 2013
word2vec: words as points in space
Tomas Mikolov and colleagues at GoogleMountain View, United States
word2vec learned word vectors quickly from billions of words. Famously, king − man + woman landed near queen. Later studies found such sums work less neatly than the headlines said, and that the vectors carry social biases.
Why it mattered. Embeddings became the standard first layer of every language model.
Ilya Sutskever, Oriol Vinyals and Quoc LeGoogle, Mountain View, United States
An LSTM read a whole sentence into one vector, and a second LSTM wrote the translation from it. It worked surprisingly well, but long sentences were hard to squeeze into one vector.
Why it mattered. It showed that one network could learn to map any text to any text.
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua BengioUniversité de Montréal, Canada
Instead of one squeezed vector, the translator learned to look back at the most relevant source words for each word it wrote. The paper was presented at ICLR in 2015.
Why it mattered. Attention solved the long-sentence problem and became the heart of the transformer.
Rico Sennrich, Barry Haddow and Alexandra BirchUniversity of Edinburgh, United Kingdom
The team borrowed a 1994 compression trick, byte-pair encoding, to split rare words into common pieces for translation. Published at ACL in 2016, it became how most LLMs make their tokens.
Why it mattered. Subword tokens let models handle any word, even ones they never saw.
One architecture, built on attention, is scaled from millions to hundreds of billions of parameters.
2017
12 June 2017
"Attention Is All You Need"
Ashish Vaswani and seven co-authors at GoogleMountain View, United States
The transformer dropped step-by-step reading and used only attention, so every word could look at every other word at once. It trained much faster on GPUs and beat earlier translators.
Why it mattered. Almost every modern LLM, from GPT to Claude, Gemini and Llama, is a transformer.
Paul Christiano, Jan Leike and colleaguesOpenAI and DeepMind
Researchers trained an agent from people's choices between pairs of short video clips, instead of a written score. This reinforcement learning from human feedback (RLHF) later shaped chatbots.
Why it mattered. It gave builders a way to tune models towards what people actually prefer.
Alec Radford and colleagues at OpenAISan Francisco, United States
OpenAI pretrained a 117-million-parameter transformer on books to predict the next token, then fine-tuned it for tasks. The Generative Pre-trained Transformer was born.
Why it mattered. Pretraining on plain text, then adapting, became the standard recipe.
Jacob Devlin and colleagues at GoogleMountain View, United States
BERT learned by filling in hidden words using the text on both sides. Fine-tuned, it topped many language tests, and Google began using it in Search in 2019.
Why it mattered. It showed how much pretrained transformers knew about language.
GPT-2, with 1.5 billion parameters, wrote fluent paragraphs from a one-line prompt. OpenAI first released only smaller versions, citing possible misuse, and published the full model in November 2019.
Why it mattered. It started the public debate about what fluent machine text could be used for.
Tom Brown and colleagues at OpenAISan Francisco, United States
GPT-3 had 175 billion parameters and was trained on about 300 billion tokens. Given a few examples in the prompt, it could do new tasks with no extra training. The same year, OpenAI researchers reported smooth "scaling laws" for size, data and compute.
Why it mattered. Scale itself became the strategy, and so did its cost in chips and energy.
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Margaret MitchellFAccT conference
This paper argued that ever-bigger language models carry costs: energy use, biased and hard-to-audit training data, and fluent text that people may wrongly trust. Its publication became part of a public dispute at Google.
Why it mattered. It put bias, energy and over-trust at the centre of the debate.
About a thousandfold growth in five years. Some later models, like GPT-4, did not disclose their size.
2018 GPT-1: 117 million
2018 BERT-Large: 340 million
2019 GPT-2: 1.5 billion
2019 T5: 11 billion
2020 GPT-3: 175 billion
2021 Gopher: 280 billion
2022 PaLM: 540 billion
2024 Llama 3.1: 405 billion
2024 DeepSeek-V3: 671 billion (37 billion active per token)
2022 – today
Chatbots for everyone
Tuned with human feedback, LLMs reach hundreds of millions of people; open models, Indian-language models and new laws follow.
2022
March 2022
Tuning models to follow instructions
Long Ouyang and colleagues at OpenAI; DeepMind's Chinchilla teamSan Francisco and London
InstructGPT used human-written examples and RLHF to make GPT-3 follow instructions; people preferred a 1.3-billion-parameter tuned model to the raw 175-billion one. The same month, DeepMind's Chinchilla showed models needed far more training text, about 20 tokens per parameter.
Why it mattered. Teaching and data, not only size, turned out to matter.
OpenAI released ChatGPT as a free research preview. It passed a million users within about five days, and was reported to reach 100 million monthly users by January 2023, one of the fastest-growing apps ever (an analyst estimate).
Why it mattered. LLMs went from research papers to everyday tools for students, workers and families.
In March 2023 OpenAI released GPT-4, without disclosing its size, and Anthropic released Claude; Google launched Bard, followed by Gemini models in December. Companies began competing on reasoning, safety and cost.
Why it mattered. The largest models also became less open about how they were built.
Meta shared LLaMA's weights with researchers; they leaked online within a week. Llama 2 followed in July 2023 under a licence allowing most commercial use, and thousands of tuned versions appeared, including many for Indian languages.
Why it mattered. Anyone with a good computer could now run, study and adapt a capable LLM.
Sarvam AI, Ola Krutrim, AI4Bharat, NPCIBengaluru and Chennai, India
Sarvam AI released OpenHathi, a Hindi model built on Llama 2 with AI4Bharat at IIT Madras, and Ola announced its Krutrim models on 15 December. Earlier, in September, NPCI had announced Hello! UPI voice payments using Hindi and English language models.
Why it mattered. India's many languages, often costly to tokenize, got models built with them in mind.
Department of Science and Technology, IIT Bombay; Sarvam AINew Delhi and Bengaluru, India
The government launched BharatGen, a publicly funded effort led by IIT Bombay to build multilingual, multimodal models for Indian languages. In October, Sarvam released Sarvam-1, a 2-billion-parameter model for 10 Indian languages and English. In April 2025 the IndiaAI Mission chose Sarvam to build a national model.
Why it mattered. Public money and startups both began building Indian foundation models.
The world's first broad AI law entered into force, sorting AI uses by risk. Rules for general-purpose models like LLMs apply from August 2025, including transparency about training, with most other rules from August 2026.
Why it mattered. Governments began setting binding rules for how AI is built and used.
DeepSeek-V3 used a mixture of experts: 671 billion parameters in total but only 37 billion used for each token. Its report claimed a much lower training cost than rivals, and its weights were published openly.
Why it mattered. Efficiency, not just size, became a race of its own.
Google reported that a median text prompt to its Gemini app used about 0.24 watt-hours of energy and 0.26 ml of water. Independent estimates for other models vary widely, and total AI electricity demand keeps rising.
Why it mattered. Energy use per answer is small, but billions of answers add up.