Clean text: stop words, stems and lemmas
Decide which words to drop and how to reduce words to a common form.
- Remove stop words and explain when that helps or hurts
- Tell stemming from lemmatization
- Write a tiny rule-based stemmer
Count the words in almost any English text and the winners are always the same: the, and, of, to, a. A handful of tiny words make up a large share of all text - an effect known as Zipf’s law. They hold sentences together but say little about the topic.
from collections import Counter
text = "the cat and the dog and the bird"
print(Counter(text.split()).most_common(2))[('the', 3), ('and', 2)]Stop words
Stop words are very common words that a pipeline throws away so the meaningful words stand out. It is a classic trick for keyword search and word counts. But “very common” is not the same as “meaningless”, as you are about to see.
Try it
Clean a review
This review is negative. Turn on Remove stop words and read what is left. Which word that flips the meaning disappeared?
Then turn on Stem and look at how studies, running and learned change. Are the results always real words?
- I
- am
- not
- happy
- with
- the
- results
- ,
- but
- the
- studies
- were
- running
- smoothly
- and
- the
- researchers
- learned
- a
- lot
- .
Stemming vs lemmatization
Both reduce related forms - study, studies, studying - to one form, so a count or a search treats them as the same word.
| Stemming | Lemmatization | |
|---|---|---|
| How | Chops off endings by rule | Looks words up, using grammar |
| studies | studi | study |
| better | better | good |
| mice | mice | mouse |
| Speed | Very fast | Slower, needs a dictionary |
A stemmer (like the classic Porter stemmer) is fast and simple, but its output (studi) need not be a real word. A lemmatizer returns the dictionary form, the lemma, and can handle irregular words like mice and better - but it needs a vocabulary and often the word’s part of speech.
Try it
Stem or lemma?
Each line shows a word being reduced. Was it a rule-based stemmer or a dictionary-based lemmatizer? Hint: real words and irregular forms point to one of them.
“studies → studi”
“mice → mouse”
“universal → univers”
“better → good”
“was → be”
“connection → connect”
Key takeaways
Stop words are very common words; removing them can sharpen counts and search - or delete meaning like “not”.
Stemming chops endings by rule (fast, crude); lemmatization finds the dictionary form (slower, smarter).
Every cleaning step is a trade-off: decide based on the task, then check what it does to real examples.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply NLP with Python
Try each text-processing idea in Python, run it against sample inputs, and use the results to see where the method works or falls short.
Remove stop words
Read a line, lowercase it, split it on whitespace, and drop every token in stop_words. Print the remaining tokens separated by one space.
- Cat on a mat
- Keeps other words
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Write a tiny stemmer
Read a line of lowercase words and stem each one with these rules, checked in order (apply only the first that matches):
- ends with
ingand is longer than 5 letters: removeing - ends with
edand is longer than 4 letters: removeed - ends with
sbut notss: remove thes
Print the stems separated by one space.
- Every rule
- Short words are left alone
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…