Choose a tokenization strategy
Split sentences consistently and make punctuation handling explicit.
- Extract word tokens while describing what a tokenizer keeps or discards.
Tokenization divides text into units. A whitespace split is quick, but leaves punctuation attached ("hello,"). A regular expression can extract word-like spans. Real tokenizers handle contractions, writing systems, and subwords with more detailed rules; there is no single best boundary for every task.
1import re
2text = "NLP is useful, isn't it?"
3tokens = re.findall(r"[a-z]+(?:'[a-z]+)?", text.lower())
4print(tokens)['nlp', 'is', 'useful', "isn't", 'it']
Whether to keep punctuation or split contractions depends on what the model should learn. The regular expression in this example keeps a simple apostrophe contraction together.
Key takeaways
Extract word tokens while describing what a tokenizer keeps or discards.
Simple baselines help make ideas concrete.
Interpret language tools in context and check important results.
Lesson quiz
4 questions · pass with 3 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Extract word tokens
Read a line, lowercase it, extract words with re.findall(r"[a-z]+", text), and print them separated by one space. Ignore punctuation.
- Punctuation
- Repeated spaces
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…