Capstone: build a tiny language model
Train a word-level bigram model, measure its perplexity, and generate text with temperature.
- Train a smoothed bigram language model from text
- Evaluate it with perplexity on held-out text
- Generate text greedily and with temperature
Time to build the whole loop - training, evaluation and generation - in a model small enough to read. A bigram model predicts each word from just the previous one, using counts from training text. Real LLMs replace the counts with a Transformer that looks at thousands of previous tokens, but the shape of the job is the same:
- Train: count which words follow which.
- Evaluate: perplexity on text the model hasn’t seen.
- Generate: pick a next word, append, repeat.
Unseen word pairs would get probability 0 (and infinite perplexity), so we use add-one smoothing: P(word | previous) = (count(previous, word) + 1) / (count(previous) + V), where V is the vocabulary size.
count_pair, count_previous, vocabulary = 0, 4, 10
print((count_pair + 1) / (count_previous + vocabulary))0.07142857142857142
Key takeaways
Training learns next-token statistics; smoothing keeps unseen pairs possible.
Perplexity on held-out text measures how surprised the model is.
Generation is repeated next-token choice - greedy or sampled with a temperature.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Train and evaluate a bigram model
Line 1 is training text; line 2 is test text (lowercase words). V is the number of distinct words across both lines. For each consecutive pair in the test text, P = (count(prev, word) + 1) / (count(prev) + V), where count(prev) counts how often prev is followed by any word in training.
Print vocabulary: V and perplexity: X (exp of the mean −ln P, 2 decimals).
- In-domain sentence
- Out-of-domain sentence
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Generate text with temperature
Line 1 is training text; line 2 is START WORDS TEMPERATURE. Starting from START, generate up to WORDS more words. For the current word, collect its followers and counts in training (sorted alphabetically); stop early if it has none.
- Temperature 0: pick the most frequent follower (alphabetically first on ties).
- Otherwise: call
random.seed(0)once at the start, and pick withrandom.choices(followers, weights=[count ** (1 / temperature) ...])[0].
Print the generated words joined by spaces.
- Greedy
- Sampled at temperature 0.5
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…