Um momento
0x10Lesson 2 of 10

Load data and make it tidy

Read CSV files with pandas, inspect their shape, and reshape messy tables into tidy ones.

22 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Load a CSV into a pandas DataFrame and inspect it
  • Apply the three rules of tidy data
  • Reshape a wide table into a long one with melt

pandas is Python’s standard tool for tables. A DataFrame is a table with named columns; each column is a Series. The first things to do with any new dataset:

1import pandas as pd
2trips = pd.read_csv("trips.csv")
3trips.shape          # (rows, columns)
4trips.head()         # first five rows
5trips.dtypes         # the type of each column
6trips.isna().sum()   # missing values per column
7trips.describe()     # summary statistics

The code here runs in your browser: the first program that imports pandas downloads it, which takes a few seconds.

first_dataframe.py
1import io
2import pandas as pd
3csv_text = """trip_id,minutes,mode
41,12,bus
52,18,walk
63,9,bus"""
7trips = pd.read_csv(io.StringIO(csv_text))
8print(trips.shape)
9print(trips["minutes"].mean())
10print(trips["mode"].value_counts().to_dict())
Output
(3, 3)
13.0
{'bus': 2, 'walk': 1}

Tidy data

Hadley Wickham’s “Tidy Data” (2014) gave a simple target shape that makes analysis easy:

  1. Each variable is a column.
  2. Each observation is a row.
  3. Each type of observational unit is a table.

Spreadsheets are often “wide”: one column per year, say. That breaks rule 1 - the year is a variable hiding in the column names. pd.melt turns wide into long (tidy); pivot goes back for display.

Try it

Tidy or messy?

Is each table layout tidy? If not, which rule does it break?

0 of 5 sortedScore 0/0
  • “city | 2024 | 2025 | 2026 (riders in each year column)”

  • “city | year | riders”

  • “station | "lat,lon" in one text column”

  • “trip rows mixed with station-address rows in one sheet”

  • “trip_id | start_time | minutes | member_type”

Key takeaways

  • Start every dataset with shape, head, types, missing values and describe().

  • Tidy: variables in columns, observations in rows, one unit per table.

  • pd.melt turns wide tables long; pivot turns them back.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

First look at a dataset

+25 XP

Read the CSV from standard input with pandas and print:

1R rows, C columns
2columns: a, b, c
3missing: col=n, col=n      (columns with missing values, in column order; or "missing: none")
4mean minutes: X            (1 decimal, ignoring missing values)
  • Gaps in two columns
  • Complete data
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Melt a wide table

+25 XP

Read a wide CSV (a city column plus one column per year) and melt it into city, year, riders. Sort by city, then year, and print each row as city year riders.

  • Two cities, two years
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: