The history

The history of large language models

From Markov counting letters in a poem in 1913 to chatbots used by hundreds of millions: a century of predicting the next word.

In 1913 Andrey Markov counted which letters followed which in a Russian poem, and in 1948 Claude Shannon used such counts to generate fake English. For decades the models stayed small. Then words became vectors, attention let networks look back, and the 2017 transformer could be scaled up almost without limit, until ChatGPT put an LLM in everyone's pocket.

112
years
29
moments
8
people
9
places

1913

Probabilities of letters counted from real text

Andrey Markov, Russia

1948

Text generated from n-gram counts

Claude Shannon, United States

1966

Chatbot

ELIZA, Joseph Weizenbaum, United States

2003

Neural network language model with learned word vectors

Yoshua Bengio and colleagues, Canada

2017

Transformer

Vaswani and colleagues, Google

2023

Hindi LLM from an Indian startup

OpenHathi, Sarvam AI with AI4Bharat, India

2024

Broad law regulating AI

EU AI Act, European Union

23 January 1913Counting letters and words

1913 – 1989

Counting letters and words

Mathematicians count which letters and words follow which, and the first chatbot fakes a conversation with simple rules.

1913

23 January 1913

Counting letters in a poem

Andrey MarkovSt Petersburg, Russia

Markov went through the first 20,000 letters of Pushkin's poem Eugene Onegin, counting how often a vowel followed a vowel or a consonant. He showed that each letter's chance depends on the one before it.

Why it mattered. It was the first chain of probabilities built from real text: the seed of every language model.

1948

July 1948

Fake English from counts

Claude ShannonBell Labs, Murray Hill, United States

In the paper that founded information theory, Shannon generated text by picking each letter or word according to how often it followed the previous ones. With longer contexts, the nonsense looked more and more like English.

Why it mattered. It is exactly the idea in chapter 1: a model that predicts the next piece from counts.

1951

January 1951

People guess the next letter

Claude ShannonBell Labs, Murray Hill, United States

Shannon asked people to guess a hidden text one letter at a time. From how often they guessed right, he estimated that English carries only about 1 bit of information per letter, because so much of it is predictable.

Why it mattered. It showed how much of language can be predicted from what came before.

1966

January 1966

ELIZA, the first chatbot

Joseph WeizenbaumMIT, Cambridge, United States

ELIZA matched keywords in what you typed and replied from simple scripts, most famously as a therapist who turned your words into questions. It learned nothing. Yet many users felt understood, which alarmed its creator.

Why it mattered. The "ELIZA effect", trusting fluent text too much, is still the key caution with LLMs.

1972

Weighing words by how rare they are

Karen Spärck JonesUniversity of Cambridge, United Kingdom

Spärck Jones proposed giving rare words more weight than common ones when matching documents to a search, an idea now called inverse document frequency.

Why it mattered. Her statistical view of words underpins search engines and fed into later word vectors.

1990 – 2012

Statistics and small networks

Translation learns from millions of example sentences, and neural networks start predicting the next word.

1990

June 1990

Translation learned from examples

Peter Brown, Robert Mercer and colleagues at IBMYorktown Heights, United States

IBM's team translated French to English using statistics learned from Canadian parliament records in both languages, with no hand-written grammar. A language model scored how natural each English sentence was.

Why it mattered. Statistical translation proved that learning from lots of text could beat rules.

1997

Networks with a longer memory

Sepp Hochreiter and Jürgen SchmidhuberMunich, Germany

The long short-term memory (LSTM) network added gates that let a network keep information over many steps. Two decades later, LSTMs powered speech recognition and translation on phones.

Why it mattered. It let networks read sequences, the job transformers later took over.

2003

A neural network predicts the next word

Yoshua Bengio and colleaguesUniversité de Montréal, Canada

Bengio's team trained a neural network to predict the next word, learning a vector for every word along the way. Similar words got similar vectors, so the model could handle sentences it had never seen.

Why it mattered. Next-word prediction plus learned word vectors is the recipe LLMs still follow.

2006

April 2006

Statistical translation for everyone

GoogleMountain View, United States

Google Translate launched using statistical machine translation, trained on huge collections of translated documents from the United Nations and the European Parliament. It moved to neural networks in 2016.

Why it mattered. Millions of people began using language models every day without knowing it.

2013 – 2016

Words as vectors, and attention

Words become points in space, networks learn to translate whole sentences, and attention lets them look back at the source.

2013

January 2013

word2vec: words as points in space

Tomas Mikolov and colleagues at GoogleMountain View, United States

word2vec learned word vectors quickly from billions of words. Famously, king − man + woman landed near queen. Later studies found such sums work less neatly than the headlines said, and that the vectors carry social biases.

Why it mattered. Embeddings became the standard first layer of every language model.

2014

September 2014

Sequence to sequence

Ilya Sutskever, Oriol Vinyals and Quoc LeGoogle, Mountain View, United States

An LSTM read a whole sentence into one vector, and a second LSTM wrote the translation from it. It worked surprisingly well, but long sentences were hard to squeeze into one vector.

Why it mattered. It showed that one network could learn to map any text to any text.

2014

September 2014

Attention: look back at the source

Dzmitry Bahdanau, Kyunghyun Cho and Yoshua BengioUniversité de Montréal, Canada

Instead of one squeezed vector, the translator learned to look back at the most relevant source words for each word it wrote. The paper was presented at ICLR in 2015.

Why it mattered. Attention solved the long-sentence problem and became the heart of the transformer.

2015

August 2015

Byte-pair encoding for words

Rico Sennrich, Barry Haddow and Alexandra BirchUniversity of Edinburgh, United Kingdom

The team borrowed a 1994 compression trick, byte-pair encoding, to split rare words into common pieces for translation. Published at ACL in 2016, it became how most LLMs make their tokens.

Why it mattered. Subword tokens let models handle any word, even ones they never saw.

2017 – 2021

The transformer and scale

One architecture, built on attention, is scaled from millions to hundreds of billions of parameters.

2017

12 June 2017

"Attention Is All You Need"

Ashish Vaswani and seven co-authors at GoogleMountain View, United States

The transformer dropped step-by-step reading and used only attention, so every word could look at every other word at once. It trained much faster on GPUs and beat earlier translators.

Why it mattered. Almost every modern LLM, from GPT to Claude, Gemini and Llama, is a transformer.

2017

June 2017

Learning from human preferences

Paul Christiano, Jan Leike and colleaguesOpenAI and DeepMind

Researchers trained an agent from people's choices between pairs of short video clips, instead of a written score. This reinforcement learning from human feedback (RLHF) later shaped chatbots.

Why it mattered. It gave builders a way to tune models towards what people actually prefer.

2018

June 2018

GPT: pretrain, then fine-tune

Alec Radford and colleagues at OpenAISan Francisco, United States

OpenAI pretrained a 117-million-parameter transformer on books to predict the next token, then fine-tuned it for tasks. The Generative Pre-trained Transformer was born.

Why it mattered. Pretraining on plain text, then adapting, became the standard recipe.

2018

October 2018

BERT reads in both directions

Jacob Devlin and colleagues at GoogleMountain View, United States

BERT learned by filling in hidden words using the text on both sides. Fine-tuned, it topped many language tests, and Google began using it in Search in 2019.

Why it mattered. It showed how much pretrained transformers knew about language.

2019

14 February 2019

GPT-2 writes convincing paragraphs

OpenAISan Francisco, United States

GPT-2, with 1.5 billion parameters, wrote fluent paragraphs from a one-line prompt. OpenAI first released only smaller versions, citing possible misuse, and published the full model in November 2019.

Why it mattered. It started the public debate about what fluent machine text could be used for.

2020

28 May 2020

GPT-3 and learning from examples in the prompt

Tom Brown and colleagues at OpenAISan Francisco, United States

GPT-3 had 175 billion parameters and was trained on about 300 billion tokens. Given a few examples in the prompt, it could do new tasks with no extra training. The same year, OpenAI researchers reported smooth "scaling laws" for size, data and compute.

Why it mattered. Scale itself became the strategy, and so did its cost in chips and energy.

2021

March 2021

"Stochastic parrots": the risks of scale

Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Margaret MitchellFAccT conference

This paper argued that ever-bigger language models carry costs: energy use, biased and hard-to-audit training data, and fluent text that people may wrongly trust. Its publication became part of a public dispute at Google.

Why it mattered. It put bias, energy and over-trust at the centre of the debate.

By the numbers

Parameters in notable language models (reported)

About a thousandfold growth in five years. Some later models, like GPT-4, did not disclose their size.

100,000,0001,000,000,00010,000,000,000100,000,000,0001,000,000,000,000 2020 2018: GPT-1: 117 million20182018: BERT-Large: 340 million2019: GPT-2: 1.5 billion20192019: T5: 11 billion2020: GPT-3: 175 billion20202021: Gopher: 280 billion20212022: PaLM: 540 billion20222024: Llama 3.1: 405 billion20242024: DeepSeek-V3: 671 billion (37 billion active per token)
  1. 2018 GPT-1: 117 million
  2. 2018 BERT-Large: 340 million
  3. 2019 GPT-2: 1.5 billion
  4. 2019 T5: 11 billion
  5. 2020 GPT-3: 175 billion
  6. 2021 Gopher: 280 billion
  7. 2022 PaLM: 540 billion
  8. 2024 Llama 3.1: 405 billion
  9. 2024 DeepSeek-V3: 671 billion (37 billion active per token)

2022 – today

Chatbots for everyone

Tuned with human feedback, LLMs reach hundreds of millions of people; open models, Indian-language models and new laws follow.

2022

March 2022

Tuning models to follow instructions

Long Ouyang and colleagues at OpenAI; DeepMind's Chinchilla teamSan Francisco and London

InstructGPT used human-written examples and RLHF to make GPT-3 follow instructions; people preferred a 1.3-billion-parameter tuned model to the raw 175-billion one. The same month, DeepMind's Chinchilla showed models needed far more training text, about 20 tokens per parameter.

Why it mattered. Teaching and data, not only size, turned out to matter.

2022

30 November 2022

ChatGPT opens to the public

OpenAISan Francisco, United States

OpenAI released ChatGPT as a free research preview. It passed a million users within about five days, and was reported to reach 100 million monthly users by January 2023, one of the fastest-growing apps ever (an analyst estimate).

Why it mattered. LLMs went from research papers to everyday tools for students, workers and families.

2023

March 2023

A crowded field

OpenAI, Anthropic, GoogleUnited States

In March 2023 OpenAI released GPT-4, without disclosing its size, and Anthropic released Claude; Google launched Bard, followed by Gemini models in December. Companies began competing on reasoning, safety and cost.

Why it mattered. The largest models also became less open about how they were built.

2023

24 February 2023

Open-weight models: Llama

Meta AIMenlo Park, United States

Meta shared LLaMA's weights with researchers; they leaked online within a week. Llama 2 followed in July 2023 under a licence allowing most commercial use, and thousands of tuned versions appeared, including many for Indian languages.

Why it mattered. Anyone with a good computer could now run, study and adapt a capable LLM.

2023

December 2023

LLMs for Indian languages

Sarvam AI, Ola Krutrim, AI4Bharat, NPCIBengaluru and Chennai, India

Sarvam AI released OpenHathi, a Hindi model built on Llama 2 with AI4Bharat at IIT Madras, and Ola announced its Krutrim models on 15 December. Earlier, in September, NPCI had announced Hello! UPI voice payments using Hindi and English language models.

Why it mattered. India's many languages, often costly to tokenize, got models built with them in mind.

2024

September–October 2024

BharatGen and Sarvam-1

Department of Science and Technology, IIT Bombay; Sarvam AINew Delhi and Bengaluru, India

The government launched BharatGen, a publicly funded effort led by IIT Bombay to build multilingual, multimodal models for Indian languages. In October, Sarvam released Sarvam-1, a 2-billion-parameter model for 10 Indian languages and English. In April 2025 the IndiaAI Mission chose Sarvam to build a national model.

Why it mattered. Public money and startups both began building Indian foundation models.

2024

1 August 2024

The EU AI Act comes into force

European UnionBrussels, Belgium

The world's first broad AI law entered into force, sorting AI uses by risk. Rules for general-purpose models like LLMs apply from August 2025, including transparency about training, with most other rules from August 2026.

Why it mattered. Governments began setting binding rules for how AI is built and used.

2024

December 2024

Cheaper training, open weights

DeepSeekHangzhou, China

DeepSeek-V3 used a mixture of experts: 671 billion parameters in total but only 37 billion used for each token. Its report claimed a much lower training cost than rivals, and its weights were published openly.

Why it mattered. Efficiency, not just size, became a race of its own.

2025

August 2025

Counting the energy of one prompt

GoogleMountain View, United States

Google reported that a median text prompt to its Gemini app used about 0.24 watt-hours of energy and 0.26 ml of water. Independent estimates for other models vary widely, and total AI electricity demand keeps rising.

Why it mattered. Energy use per answer is small, but billions of answers add up.

By the numbers

Training text for notable models (reported)

After 2022, builders grew the training text faster than the model size.

100,000,000,0001,000,000,000,00010,000,000,000,000100,000,000,000,000 2020 2020: GPT-3: about 300 billion tokens20202022: Chinchilla: 1.4 trillion20222023: Llama 2: 2 trillion20232024: Llama 3.1: over 15 trillion2024
  1. 2020 GPT-3: about 300 billion tokens
  2. 2022 Chinchilla: 1.4 trillion
  3. 2023 Llama 2: 2 trillion
  4. 2024 Llama 3.1: over 15 trillion

Did you know?

Shannon estimated that English carries only about 1 bit of information per letter, because most letters can be guessed from what came before.

Training GPT-3 took about 3.14 × 10²³ arithmetic operations, and was estimated to use about 1,287 MWh of electricity.

In UTF-8 an English letter takes 1 byte and a Devanagari letter 3, one reason Hindi can cost more tokens.

ELIZA had no understanding at all, yet its creator's secretary reportedly asked him to leave the room so she could talk to it in private.

The transformer was invented for translation, not chat.

The people

Who figured it out

Andrey Markov

1856 – 1922 · Mathematician · Russia

Counted letter sequences in Eugene Onegin in 1913, founding the theory of Markov chains.

Claude Shannon

1916 – 2001 · Engineer and mathematician · United States

Founded information theory and generated English-like text from letter and word counts.

Joseph Weizenbaum

1923 – 2008 · Computer scientist · Germany and United States

Built ELIZA, then spent decades warning people not to over-trust computers.

Karen Spärck Jones

1935 – 2007 · Computer scientist · United Kingdom

Pioneered statistical ways of weighing words that still power search.

Yoshua Bengio

born 1964 · Computer scientist · Canada

Built an early neural language model (2003) and co-wrote the 2014 attention paper.

Ashish Vaswani

· Computer scientist · India

Lead author of "Attention Is All You Need", the 2017 transformer paper, written with seven co-authors.

Timnit Gebru

· Computer scientist · Ethiopia and United States

Co-wrote the 2021 "Stochastic Parrots" paper on the risks of large language models.

Mitesh Khapra

· Computer scientist, IIT Madras · India

A co-founder of AI4Bharat, which builds open datasets and models for Indian languages.

Where it happened

9 places, one idea

Sources

Where this comes from

Dates marked “c.” are approximate, and historians sometimes disagree about who was first. If you spot a mistake, tell us.

  1. First Links in the Markov Chain American Scientist (Brian Hayes, 2013)
  2. Markov chain Wikipedia
  3. A Mathematical Theory of Communication (C. E. Shannon, 1948) Bell System Technical Journal, via Harvard
  4. Prediction and Entropy of Printed English (C. E. Shannon, 1951) Bell System Technical Journal, via Princeton
  5. n-gram Wikipedia
  6. ELIZA: a computer program for the study of natural language communication between man and machine (Weizenbaum, 1966) Communications of the ACM
  7. ELIZA Wikipedia
  8. Karen Spärck Jones Wikipedia
  9. A Statistical Approach to Machine Translation (Brown et al., 1990) Computational Linguistics (ACL Anthology)
  10. Statistical machine translation Wikipedia
  11. Long short-term memory Wikipedia
  12. A Neural Probabilistic Language Model (Bengio et al., 2003) Journal of Machine Learning Research
  13. Google Translate Wikipedia
  14. Efficient Estimation of Word Representations in Vector Space (Mikolov et al., 2013) arXiv
  15. Word2vec Wikipedia
  16. Sequence to Sequence Learning with Neural Networks (Sutskever, Vinyals & Le, 2014) arXiv
  17. Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau, Cho & Bengio, 2014) arXiv
  18. Neural Machine Translation of Rare Words with Subword Units (Sennrich, Haddow & Birch, 2016) ACL Anthology
  19. Attention Is All You Need (Vaswani et al., 2017) arXiv
  20. Transformer (deep learning architecture) Wikipedia
  21. Deep reinforcement learning from human preferences (Christiano et al., 2017) arXiv
  22. Improving Language Understanding by Generative Pre-Training (Radford et al., 2018) OpenAI
  23. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (Devlin et al., 2018) arXiv
  24. Better language models and their implications OpenAI
  25. Language Models are Few-Shot Learners (Brown et al., 2020) arXiv
  26. Scaling Laws for Neural Language Models (Kaplan et al., 2020) arXiv
  27. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? (Bender, Gebru et al., 2021) ACM FAccT
  28. Carbon Emissions and Large Neural Network Training (Patterson et al., 2021) arXiv
  29. Training language models to follow instructions with human feedback (Ouyang et al., 2022) arXiv
  30. Training Compute-Optimal Large Language Models (Hoffmann et al., 2022) arXiv
  31. Introducing ChatGPT OpenAI
  32. ChatGPT Wikipedia
  33. LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023) arXiv
  34. Llama (language model) Wikipedia
  35. Introducing Llama 3.1: Our most capable models to date Meta AI
  36. GPT-4 Technical Report arXiv (OpenAI)
  37. Claude (language model) Wikipedia
  38. Gemini (language model) Wikipedia
  39. Sarvam AI Wikipedia
  40. Ola launches Krutrim, AI model trained on Indian languages MediaNama
  41. NPCI launches new products; users can now make voice UPI payments YourStory
  42. Launch of BharatGen: the first Government supported Multimodal Large Language Model Initiative Department of Science & Technology, Government of India
  43. Sarvam AI launches homegrown multilingual LLM, Sarvam-1 YourStory
  44. Sarvam to build India's sovereign large language model Sarvam AI
  45. AI Act enters into force European Commission
  46. Artificial Intelligence Act Wikipedia
  47. DeepSeek-V3 Technical Report arXiv
  48. Measuring the environmental impact of AI inference Google Cloud Blog
  49. Language Model Tokenizers Introduce Unfairness Between Languages (Petrov et al., 2023) arXiv / NeurIPS
  50. PaLM: Scaling Language Modeling with Pathways (Chowdhery et al., 2022) arXiv
  51. Scaling Language Models: Methods, Analysis & Insights from Training Gopher (Rae et al., 2021) arXiv
  52. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (Raffel et al., 2019) arXiv
  53. Llama 2: Open Foundation and Fine-Tuned Chat Models (Touvron et al., 2023) arXiv

That's the history. Now see how it works.