Um momento
0x30Lesson 4 of 15

Learn subword tokens with byte-pair encoding

See how modern tokenizers learn reusable word pieces from data.

20 min 6-question quiz 1 code exercise
By the end of this lesson you can
  • Explain the out-of-vocabulary problem and why subwords solve it
  • Run byte-pair encoding merges by hand
  • Apply learned merges to split a word the tokenizer has never seen

Every tokenizer faces a dilemma about its vocabulary - the set of tokens it knows:

  • Whole words: a word missing from the vocabulary (“unfriendliest”, a new product name, a typo) becomes an unknown token, and its meaning is lost. This is the out-of-vocabulary problem.
  • Single characters: nothing is ever unknown, but sentences become very long sequences, and single letters carry little meaning.

Subword tokenizers sit in between: common words stay whole, and rare words split into familiar pieces like un + friend + li + est. The models behind today’s chat assistants all read subword tokens.

How BPE learns its pieces

BPE learns its vocabulary from a corpus, bottom-up:

  1. Start with every word split into single characters.
  2. Count every pair of adjacent symbols across the corpus (a word seen 6 times counts 6 times).
  3. Merge the most frequent pair everywhere, creating one new symbol, and remember that merge.
  4. Repeat until the vocabulary is the size you want.

Try it

Be the tokenizer

This tiny corpus is the classic example from the BPE paper. Each round, pick the adjacent pair you think is most frequent - remember to multiply by how often each word was seen. Watch the green merged symbols grow into useful pieces like est.

Round 1 of 4Score 0/0
Corpus words split into their current symbols
WordSymbols nowSeen
lowlow×5
lowerlower×2
newestnewest×6
widestwidest×3

Which adjacent pair appears most often in the corpus? (Count each word as many times as it was seen.)

Splitting a new word

To tokenize text, apply the learned merges in the order they were learned. Here the word “lowest” never appeared in training, yet it splits into two known pieces:

apply_merges.py
1merges = [("e", "s"), ("es", "t"), ("l", "o"), ("lo", "w")]
2symbols = list("lowest")
3for first, second in merges:
4    merged = []
5    position = 0
6    while position < len(symbols):
7        if position + 1 < len(symbols) and (symbols[position], symbols[position + 1]) == (first, second):
8            merged.append(first + second)
9            position += 2
10        else:
11            merged.append(symbols[position])
12            position += 1
13    symbols = merged
14print(symbols)
Output
['low', 'est']

Key takeaways

  • Word vocabularies break on unseen words; character vocabularies make sequences long. Subwords balance the two.

  • BPE learns merges by repeatedly joining the most frequent adjacent pair.

  • New words are split by replaying the learned merges, so nothing is ever truly unknown.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply NLP with Python

Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.

Exercise 1

Find the first merge

+25 XP

Do the first step of BPE training. The first input line is the number of words N. Each of the next N lines has a word and how many times it was seen, like low 5.

Treat each word as a sequence of characters, count every adjacent pair (weighted by the word’s count), and print the most frequent pair as a b: count. If pairs tie, print the one that comes first alphabetically (compare them as "a b" strings).

  • The BPE paper corpus (a tie)
  • A clear winner
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: