Model short sequences with n-grams
Count nearby word sequences and explore what context can tell us.
- Count bigrams and explain the limits of short-context models.
An n-gram is a sequence of n consecutive tokens; a bigram has two. N-gram language models estimate a next-token probability from nearby sequence counts. They are easy to understand and can produce locally plausible text, but sparse data and short context limit them.
words = "we learn language".split()
bigrams = list(zip(words, words[1:]))
print(bigrams)[('we', 'learn'), ('learn', 'language')]A bigram model estimates P(next | previous) from observed counts. An unseen context can get zero probability without smoothing. Longer contexts capture more patterns but also make data sparsity worse. Autoregressive language models extend next-token prediction with learned representations and larger contexts.
Key takeaways
Count bigrams and explain the limits of short-context models.
Simple baselines help make ideas concrete.
Interpret language tools in context and check important results.
Lesson quiz
4 questions · pass with 3 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Count adjacent word pairs
Read a line, lowercase it, extract alphabetic words, count adjacent bigrams, and print each pair alphabetically as first second: count.
- Repeated pair
- Two pairs
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…