Understand tokens and next-token prediction
See how text becomes model input and how a language model generates a sequence.
- Explain tokens, context, and autoregressive next-token generation.
A language model receives text as tokens: units that may represent whole words, word pieces, punctuation, or other symbols. An autoregressive language model estimates a distribution over the next token given the preceding context. Generation repeats this process, appending a selected token until a stopping condition is reached. This predicts likely continuations; it does not mean the model verifies every statement against reality.
A small example
tokens = ["The", " cat", " sat"]
next_token = " down"
print("".join(tokens) + next_token)The cat sat down
Token boundaries are not the same as words, and different tokenizers can split the same text differently. Context windows limit how much input a model can process at once. The generated continuation depends on the prompt, model parameters, and decoding choices.
Key takeaways
Explain tokens, context, and autoregressive next-token generation.
Measure behavior with realistic examples and inspect important failures.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…