Tech●●●●●Difficulty 4 of 5

How does a large language model write a sentence, one piece at a time?

A chatbot never plans its whole answer: it repeatedly guesses the next little chunk of text, and then the next.

▶ Start the story

By guessing the next piece over and over. A large language model is a neural network trained on a vast amount of text, and at heart it does one thing: given the text so far, it estimates how likely each possible next piece is. Is "I like to eat" more likely to be followed by "bread" or "rocks"? It picks a piece, adds it to the text, and asks the same question again. A whole answer is just that loop, run again and again.

The pieces are called tokens. Since the model only handles numbers, text is first split into tokens from a fixed vocabulary, often whole words or fragments of words, and each one becomes a number and then a list of numbers called an embedding.

The clever part is how the model uses context. Its transformer design has an attention mechanism that relates every token to every other one at once, however far apart they are, so the end of a long sentence can depend on its beginning. At the end, the model turns its calculation into a probability for every token in its vocabulary. A setting called temperature decides how adventurous the pick is: low means almost always the top choice, high means more randomness.

This explains a famous weakness. Training rewards a good guess at the next word even when the model lacks the information, so it can produce fluent, confident text that is false. When the data scientist Teresa Kubacka asked ChatGPT about a phenomenon she had made up, it answered with plausible-looking citations.

How a language model writes
  1. Step 1: Split into tokens

    Text becomes numbered pieces from a fixed vocabulary.

  2. Step 2: Embed

    Each token becomes a list of numbers.

  3. Step 3: Attend

    Attention relates every token to every other in the window.

  4. Step 4: Score every next token

    The model gives a probability to each token in the vocabulary.

  5. Step 5: Pick, append, repeat

    One token is chosen, with temperature setting how adventurous, and the loop runs again.

A wide chart: along the horizontal axis, each generated token of an AI's answer; above each, translucent discs show the top 16 candidate tokens and their probabilities, with the chosen one marked.
Inside a real model's writing: for each token it produced (left to right), the discs show its 16 likeliest candidates and their probabilities, run at temperature 0 and at temperature 1.Photo: Hjhornbeck · CC BY-SA 4.0

Quiz me

0/3

  1. 1.What does a language model actually output at each step?
  2. 2.What does lowering the temperature do to the text a model writes?
  3. 3.Why can a language model state something false with confidence?

Recap

Tokens in, probabilities out, pick one, append, repeat.

Surprising fact · A model invented a plausible answer, with citations, about a physics phenomenon that a data scientist had simply made up.

Sources (5)

No source, no claim. Every fact in this lesson (26 claims) cites at least one of these.

  1. [1]Large language model · Wikipedia
  2. [2]Transformer (deep learning) · Wikipedia
  3. [3]Attention Is All You Need · Wikipedia
  4. [4]Hallucination (artificial intelligence) · Wikipedia
  5. [5]Softmax function · Wikipedia
More lessons in 💻 Tech (3) See all tech lessons →

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson ↗