Um momento
0xA0Lesson 11 of 16

Know the vision tasks: classify, detect, segment

Tell classification, detection and segmentation apart, and work with bounding-box formats.

18 min 6-question quiz 1 code exercise
By the end of this lesson you can
  • Match a problem to classification, detection or segmentation
  • Describe what each task outputs
  • Convert between common bounding-box formats

Vision problems differ in what they must output:

TaskQuestion it answersOutput
ClassificationWhat is in this image?One label
Object detectionWhat objects are where?A box + label per object
Semantic segmentationWhich class is each pixel?A label for every pixel
Instance segmentationWhich object does each pixel belong to?A mask per object
Keypoint detectionWhere are the joints/landmarks?Points per object

Detection and segmentation take more effort to label - drawing boxes and masks costs far more than typing one label - so pick the simplest task that answers your question.

Try it

Which task is it?

For each product idea, choose the vision task that fits best. Ask yourself: does it need where, and how precisely?

0 of 6 sortedScore 0/0
  • “Sort uploaded photos into “food”, “pets” and “landscapes””

  • “Count the cars in a parking-lot camera feed”

  • “Blur the background behind a person in a video call”

  • “Tell whether a chest X-ray looks normal or abnormal”

  • “Draw boxes around pedestrians for a self-driving car”

  • “Measure the exact area of a tumor in a scan”

Bounding-box formats

A box is four numbers, but libraries disagree on which four:

  • corner + size: [x, y, width, height] (x, y = top-left corner),
  • two corners: [x1, y1, x2, y2] (top-left and bottom-right),
  • center + size: [center_x, center_y, width, height], often divided by the image size so values run 0–1.

Mixing formats produces boxes in the wrong place without any error message, so convert carefully.

box_formats.py
x, y, width, height = 10, 20, 30, 40
print([x, y, x + width, y + height])
print([x + width / 2, y + height / 2, width, height])
Output
[10, 20, 40, 60]
[25.0, 40.0, 30, 40]

Key takeaways

  • Classification → one label; detection → boxes + labels; segmentation → pixel-level masks.

  • Choose the simplest task that answers the question - labels for harder tasks cost more.

  • Know your box format: [x, y, w, h], [x1, y1, x2, y2] or center-based.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: apply computer vision with Python

Use small pixel arrays to explore vision concepts, run your code against sample images, and connect each result to the larger computer vision idea.

Exercise 1

Convert a box to corners

+25 XP

Read four numbers x y width height (a box’s top-left corner and size). Print the same box as two corners: x1 y1 x2 y2.

  • A box
  • At the origin
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: