Load data and make it tidy
Read CSV files with pandas, inspect their shape, and reshape messy tables into tidy ones.
- Load a CSV into a pandas DataFrame and inspect it
- Apply the three rules of tidy data
- Reshape a wide table into a long one with melt
pandas is Python’s standard tool for tables. A DataFrame is a table with named columns; each column is a Series. The first things to do with any new dataset:
1import pandas as pd
2trips = pd.read_csv("trips.csv")
3trips.shape # (rows, columns)
4trips.head() # first five rows
5trips.dtypes # the type of each column
6trips.isna().sum() # missing values per column
7trips.describe() # summary statisticsThe code here runs in your browser: the first program that imports pandas downloads it, which takes a few seconds.
1import io
2import pandas as pd
3csv_text = """trip_id,minutes,mode
41,12,bus
52,18,walk
63,9,bus"""
7trips = pd.read_csv(io.StringIO(csv_text))
8print(trips.shape)
9print(trips["minutes"].mean())
10print(trips["mode"].value_counts().to_dict())(3, 3)
13.0
{'bus': 2, 'walk': 1}Tidy data
Hadley Wickham’s “Tidy Data” (2014) gave a simple target shape that makes analysis easy:
- Each variable is a column.
- Each observation is a row.
- Each type of observational unit is a table.
Spreadsheets are often “wide”: one column per year, say. That breaks rule 1 - the year is a variable hiding in the column names. pd.melt turns wide into long (tidy); pivot goes back for display.
Try it
Tidy or messy?
Is each table layout tidy? If not, which rule does it break?
“city | 2024 | 2025 | 2026 (riders in each year column)”
“city | year | riders”
“station | "lat,lon" in one text column”
“trip rows mixed with station-address rows in one sheet”
“trip_id | start_time | minutes | member_type”
Key takeaways
Start every dataset with shape, head, types, missing values and describe().
Tidy: variables in columns, observations in rows, one unit per table.
pd.melt turns wide tables long; pivot turns them back.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
First look at a dataset
Read the CSV from standard input with pandas and print:
1R rows, C columns
2columns: a, b, c
3missing: col=n, col=n (columns with missing values, in column order; or "missing: none")
4mean minutes: X (1 decimal, ignoring missing values)- Gaps in two columns
- Complete data
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Melt a wide table
Read a wide CSV (a city column plus one column per year) and melt it into city, year, riders. Sort by city, then year, and print each row as city year riders.
- Two cities, two years
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…