Reinforcement learning: learning by trial and error
Train a treasure-hunting robot with Q-learning, tune its curiosity and patience, and meet agents that cheat their own reward.
- Describe the agent-environment loop: states, actions and rewards
- Explain what a Q-value is and how the Q-learning update improves it
- Explain the roles of the learning rate, the discount and exploration - and what reward hacking is
In the slot-machine example in Types of machine learning, every pull paid out (or didn’t) right away. Real problems are harder: a chess move, a step in a maze or a turn of a steering wheel might only pay off - or backfire - many moves later. Reinforcement learning (RL) tackles exactly this.
An agent observes the state of its world, picks an action, and receives a reward and a new state. Nobody tells it the right move. Its goal is to maximize the total reward it collects over time - the return.
Rewards in the future are usually worth a little less than rewards now, so the return discounts them by a factor (gamma) per step:
With the agent only cares about the very next reward; close to 1, it plans far ahead. Watch how gamma changes whether a far-away treasure beats a quick snack:
1def discounted_return(rewards, gamma):
2 total = 0
3 for step, reward in enumerate(rewards):
4 total += gamma ** step * reward
5 return total
6
7slow_treasure = [-1, -1, -1, -1, -1, 10] # five steps, then the treasure
8quick_snack = [-1, 2] # one step, then a small snack
9for gamma in [0.5, 0.9, 1.0]:
10 print(f"gamma {gamma}: treasure {discounted_return(slow_treasure, gamma):6.2f} snack {discounted_return(quick_snack, gamma):5.2f}")gamma 0.5: treasure -1.62 snack 0.00 gamma 0.9: treasure 1.81 snack 0.80 gamma 1.0: treasure 5.00 snack 1.00
An impatient agent () grabs the snack; a patient one walks to the treasure. Neither is “right” - the discount is a design choice that shapes behavior.
Q-learning: a cheat sheet that writes itself
Q-learning keeps a table of Q-values: is the agent’s current estimate of the return it will get if it takes action in state and plays well afterwards. Once the table is good, acting is easy: pick the action with the highest Q.
The table starts at zero. After every step, the agent nudges one entry toward what it just experienced:
In words: the value of a move = the reward I just got + the (discounted) value of the best move from where I landed. The learning rate says how far to move toward that new estimate. Rewards found at the goal slowly “leak” backwards through the table, one step at a time, until even the start knows which way to go.
Try it
Train a treasure hunter
The robot starts knowing nothing. Each step costs −1, the diamond is worth +10 and the skulls are −10 pits. Click 1 episode a few times and watch values appear near the goal and spread backwards. Then press 50 episodes and watch the learning curve climb. Experiment: set exploration to 0 and reset - does it still learn? What happens with a discount of 0?
Arrows: best known action. Number: best Q-value. Green = promising, red = dangerous. Blue outline: the route the agent would take now. Blue dots: the last episode’s wander.
1import random
2
3random.seed(4)
4# A corridor: cells 0..4. The treasure (+10) is at cell 4; every step costs 1.
5ACTIONS = ["left", "right"]
6q = {(cell, action): 0.0 for cell in range(5) for action in ACTIONS}
7alpha, gamma, epsilon = 0.5, 0.9, 0.2
8
9for episode in range(200):
10 cell = 0
11 while cell != 4:
12 if random.random() < epsilon: # explore
13 action = random.choice(ACTIONS)
14 else: # exploit
15 action = max(ACTIONS, key=lambda a: q[(cell, a)])
16 nxt = max(0, cell - 1) if action == "left" else cell + 1
17 reward = 10 if nxt == 4 else -1
18 future = 0 if nxt == 4 else max(q[(nxt, a)] for a in ACTIONS)
19 q[(cell, action)] += alpha * (reward + gamma * future - q[(cell, action)])
20 cell = nxt
21
22for cell in range(4):
23 left, right = q[(cell, "left")], q[(cell, "right")]
24 print(f"cell {cell}: left {left:6.2f} right {right:6.2f} -> go {'right' if right > left else 'left'}")cell 0: left 3.12 right 4.58 -> go right cell 1: left 3.12 right 6.20 -> go right cell 2: left 4.58 right 8.00 -> go right cell 3: left 6.20 right 10.00 -> go right
Look at the “right” column: 10, then −1 + 0.9 × 10 = 8, then −1 + 0.9 × 8 = 6.2, and so on. The treasure’s value has flowed backwards through the corridor exactly as the update rule predicts.
Key takeaways
RL agents learn from rewards through a loop: state → action → reward + next state.
The return adds up future rewards, discounted by : low gamma = impatient, high gamma = plans ahead.
Q-learning improves toward ; values spread backwards from rewards.
Agents must explore to find good strategies - and they will exploit any loophole in the reward (reward hacking).
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Apply Q-learning updates
Line 1 holds the learning rate α and the discount γ. Each following line is one experienced transition: state action reward next_state, where next_state is end when the episode finishes.
All Q-values start at 0. For each transition, apply . The future term is 0 when the next state is end or has no Q-values yet.
Finally, for each state in the order it first appeared, print state: best_action value with the value to 2 decimal places. On ties, pick the action that state tried first.
- Value flows backwards
- Learning rate 1
- Avoid the pit
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Patient or impatient?
Line 1: the discount γ. Line 2: the rewards received at each step. Print return: G - the discounted return (first reward at t = 0) to 2 decimal places - and undiscounted: S, the plain sum formatted with :g.
- The slow treasure
- Halving
- No discount
- Only now matters
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…