Skip to main content

What is a token in AI? Tokens, tokenizers and byte-pair encoding

What a token is in a language model like ChatGPT, how a tokenizer turns text into numbers, and how byte-pair encoding (BPE) builds its vocabulary, with a worked example.

Updated 10 Oct 2026Examples checked by running the code

A token is the unit of text a language model reads and writes. Before a model like ChatGPT sees your message, a tokenizer cuts it into tokens and turns each one into a number. The model only ever works with those numbers. When it answers, it produces one token at a time, and the tokenizer turns the numbers back into text.

A token can be a single letter, a piece of a word, a whole word, or a space plus a word. Which one depends on the tokenizer.

The simplest tokenizer: one token per letter

Andrej Karpathy's microgpt, a complete GPT in 200 lines of Python, uses the simplest possible tokenizer. It reads first names, and every letter is its own token:

  • the 26 letters a to z get the ids 0 to 25;
  • one extra token, BOS (beginning of sequence), marks where a name starts and ends.

That is a vocabulary of 27 tokens. The name "emma" becomes five steps for the model: start marker, e, m, m, a, and then it should predict the end marker.

In Python, the whole tokenizer is a sorted list of the letters and a dictionary from letter to id:

docs = ['emma', 'olivia', 'ava']
uchars = sorted(set(''.join(docs)))   # the unique letters, in order
stoi = {ch: i for i, ch in enumerate(uchars)}
print(uchars, [stoi[ch] for ch in 'ava'])
# ['a', 'e', 'i', 'l', 'm', 'o', 'v'] [0, 6, 0]

Why real models don't use letters

One token per letter is easy, but wasteful: every word costs many steps, and a model can only look at a fixed number of tokens at once (its context window). Real models use bigger pieces, so the same text takes far fewer tokens. GPT-2's vocabulary has 50,257 tokens: frequent words, common word pieces like "ing", and pieces with a leading space like " the".

A rough rule of thumb from OpenAI for English text: one token is about four characters, or about three quarters of a word.

Byte-pair encoding (BPE), step by step

Most GPT tokenizers build their vocabulary with byte-pair encoding (BPE). The idea:

  1. Start with every character as its own token.
  2. Find the most common pair of neighbouring tokens in the training text.
  3. Merge that pair into one new token, and add it to the vocabulary.
  4. Repeat until the vocabulary is as big as you want.

Here it is on one short sentence: the cat sat on the mat. the hat is on the cat. (46 characters, so 46 tokens to start). Spaces are written as ␣.

mergepairtimes it appearstokens after
1a + t541
2t + h437
3th + e433
4the + ␣429
5at + ␣326
6␣ + the␣323

After six merges the same sentence takes 23 tokens instead of 46, and the vocabulary has learned "at", "the␣" and "␣the␣". On a huge pile of text, the merges end up capturing common words and word pieces. GPT-2's 50,257 tokens are 256 single bytes, 50,000 merges and one special end-of-text token.

What tokens explain about chatbots

  • Odd spelling mistakes. A model sees "strawberry" as a few tokens, not ten letters, so counting the letters inside a word is surprisingly hard for it.
  • Limits and prices. Context windows and API prices are measured in tokens, not words.
  • Other languages. Text in languages that were rare in the tokenizer's training data often needs more tokens for the same meaning.

Try it yourself

Zero to microGPT builds microgpt's tokenizer in lesson 2.2, uses it for a first language model in 5.1, and covers BPE in 11.3. Lesson 11.3 has a widget where you type any text and press "merge the most common pair" to watch the token count fall.

Common questions

Is a token the same as a word? Sometimes. Common words are often one token, but longer or rarer words are split into several pieces, and spaces and punctuation are usually part of tokens too.

Why does the model need numbers at all? Everything inside a neural network is arithmetic. Each token id picks out a row of numbers (its embedding) that the model learns during training, and that is where the maths starts.

Does every model use the same tokenizer? No. Each model family trains its own, with its own vocabulary size. The same sentence can be a different number of tokens in different models.