Know the vision tasks: classify, detect, segment
Tell classification, detection and segmentation apart, and work with bounding-box formats.
- Match a problem to classification, detection or segmentation
- Describe what each task outputs
- Convert between common bounding-box formats
Vision problems differ in what they must output:
| Task | Question it answers | Output |
|---|---|---|
| Classification | What is in this image? | One label |
| Object detection | What objects are where? | A box + label per object |
| Semantic segmentation | Which class is each pixel? | A label for every pixel |
| Instance segmentation | Which object does each pixel belong to? | A mask per object |
| Keypoint detection | Where are the joints/landmarks? | Points per object |
Detection and segmentation take more effort to label - drawing boxes and masks costs far more than typing one label - so pick the simplest task that answers your question.
Try it
Which task is it?
For each product idea, choose the vision task that fits best. Ask yourself: does it need where, and how precisely?
“Sort uploaded photos into “food”, “pets” and “landscapes””
“Count the cars in a parking-lot camera feed”
“Blur the background behind a person in a video call”
“Tell whether a chest X-ray looks normal or abnormal”
“Draw boxes around pedestrians for a self-driving car”
“Measure the exact area of a tumor in a scan”
Bounding-box formats
A box is four numbers, but libraries disagree on which four:
- corner + size:
[x, y, width, height](x, y = top-left corner), - two corners:
[x1, y1, x2, y2](top-left and bottom-right), - center + size:
[center_x, center_y, width, height], often divided by the image size so values run 0–1.
Mixing formats produces boxes in the wrong place without any error message, so convert carefully.
x, y, width, height = 10, 20, 30, 40
print([x, y, x + width, y + height])
print([x + width / 2, y + height / 2, width, height])[10, 20, 40, 60] [25.0, 40.0, 30, 40]
Key takeaways
Classification → one label; detection → boxes + labels; segmentation → pixel-level masks.
Choose the simplest task that answers the question - labels for harder tasks cost more.
Know your box format:
[x, y, w, h],[x1, y1, x2, y2]or center-based.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: apply computer vision with Python
Use small pixel arrays to explore vision concepts, run your code against sample images, and connect each result to the larger computer vision idea.
Convert a box to corners
Read four numbers x y width height (a box’s top-left corner and size). Print the same box as two corners: x1 y1 x2 y2.
- A box
- At the origin
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…