Um momento
0xB0Lesson 12 of 18

Reinforcement learning: learning by trial and error

Train a treasure-hunting robot with Q-learning, tune its curiosity and patience, and meet agents that cheat their own reward.

25 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Describe the agent-environment loop: states, actions and rewards
  • Explain what a Q-value is and how the Q-learning update improves it
  • Explain the roles of the learning rate, the discount and exploration - and what reward hacking is

In the slot-machine example in Types of machine learning, every pull paid out (or didn’t) right away. Real problems are harder: a chess move, a step in a maze or a turn of a steering wheel might only pay off - or backfire - many moves later. Reinforcement learning (RL) tackles exactly this.

An agent observes the state of its world, picks an action, and receives a reward and a new state. Nobody tells it the right move. Its goal is to maximize the total reward it collects over time - the return.

🤖 Agentpolicy: state → action🌍 Worldgame, robot body, market…action aₜnext state sₜ₊₁ + reward rₜ₊₁repeat thousands(or millions) of times ↻
The reinforcement learning loop: the agent acts, the world responds with a new situation and a reward, and the agent adjusts so that future rewards add up to more.

Rewards in the future are usually worth a little less than rewards now, so the return discounts them by a factor γ\gamma (gamma) per step:

G=r1+γr2+γ2r3+…G = r_1 + \gamma r_2 + \gamma^2 r_3 + \dots

With γ=0\gamma = 0 the agent only cares about the very next reward; close to 1, it plans far ahead. Watch how gamma changes whether a far-away treasure beats a quick snack:

discounted_return.py
1def discounted_return(rewards, gamma):
2    total = 0
3    for step, reward in enumerate(rewards):
4        total += gamma ** step * reward
5    return total
6
7slow_treasure = [-1, -1, -1, -1, -1, 10]   # five steps, then the treasure
8quick_snack = [-1, 2]                       # one step, then a small snack
9for gamma in [0.5, 0.9, 1.0]:
10    print(f"gamma {gamma}: treasure {discounted_return(slow_treasure, gamma):6.2f}   snack {discounted_return(quick_snack, gamma):5.2f}")
Output
gamma 0.5: treasure  -1.62   snack  0.00
gamma 0.9: treasure   1.81   snack  0.80
gamma 1.0: treasure   5.00   snack  1.00

An impatient agent (γ=0.5\gamma = 0.5) grabs the snack; a patient one walks to the treasure. Neither is “right” - the discount is a design choice that shapes behavior.

Q-learning: a cheat sheet that writes itself

Q-learning keeps a table of Q-values: Q(s,a)Q(s, a) is the agent’s current estimate of the return it will get if it takes action aa in state ss and plays well afterwards. Once the table is good, acting is easy: pick the action with the highest Q.

The table starts at zero. After every step, the agent nudges one entry toward what it just experienced:

Q(s,a)←Q(s,a)+α[ r+γmax⁡a′Q(s′,a′)−Q(s,a) ]Q(s, a) \leftarrow Q(s, a) + \alpha \big[\, r + \gamma \max_{a'} Q(s', a') - Q(s, a) \,\big]

In words: the value of a move = the reward I just got + the (discounted) value of the best move from where I landed. The learning rate α\alpha says how far to move toward that new estimate. Rewards found at the goal slowly “leak” backwards through the table, one step at a time, until even the start knows which way to go.

Try it

Train a treasure hunter

The robot starts knowing nothing. Each step costs −1, the diamond is worth +10 and the skulls are −10 pits. Click 1 episode a few times and watch values appear near the goal and spread backwards. Then press 50 episodes and watch the learning curve climb. Experiment: set exploration to 0 and reset - does it still learn? What happens with a discount of 0?

Arrows: best known action. Number: best Q-value. Green = promising, red = dangerous. Blue outline: the route the agent would take now. Blue dots: the last episode’s wander.

0
Episodes
—
Avg reward (last 20)
Not yet
Greedy route
010-20episodes →Train a few episodes to draw the learning curve
How far each surprise moves the estimate.
How much future rewards count. 0 = live for now.
Chance of a random move instead of the best one.
q_learning_corridor.py
1import random
2
3random.seed(4)
4# A corridor: cells 0..4. The treasure (+10) is at cell 4; every step costs 1.
5ACTIONS = ["left", "right"]
6q = {(cell, action): 0.0 for cell in range(5) for action in ACTIONS}
7alpha, gamma, epsilon = 0.5, 0.9, 0.2
8
9for episode in range(200):
10    cell = 0
11    while cell != 4:
12        if random.random() < epsilon:                       # explore
13            action = random.choice(ACTIONS)
14        else:                                               # exploit
15            action = max(ACTIONS, key=lambda a: q[(cell, a)])
16        nxt = max(0, cell - 1) if action == "left" else cell + 1
17        reward = 10 if nxt == 4 else -1
18        future = 0 if nxt == 4 else max(q[(nxt, a)] for a in ACTIONS)
19        q[(cell, action)] += alpha * (reward + gamma * future - q[(cell, action)])
20        cell = nxt
21
22for cell in range(4):
23    left, right = q[(cell, "left")], q[(cell, "right")]
24    print(f"cell {cell}: left {left:6.2f}  right {right:6.2f}  -> go {'right' if right > left else 'left'}")
Output
cell 0: left   3.12  right   4.58  -> go right
cell 1: left   3.12  right   6.20  -> go right
cell 2: left   4.58  right   8.00  -> go right
cell 3: left   6.20  right  10.00  -> go right

Look at the “right” column: 10, then −1 + 0.9 × 10 = 8, then −1 + 0.9 × 8 = 6.2, and so on. The treasure’s value has flowed backwards through the corridor exactly as the update rule predicts.

Key takeaways

  • RL agents learn from rewards through a loop: state → action → reward + next state.

  • The return adds up future rewards, discounted by γ\gamma: low gamma = impatient, high gamma = plans ahead.

  • Q-learning improves Q(s,a)Q(s, a) toward r+γmax⁡Q(s′,a′)r + \gamma \max Q(s', a'); values spread backwards from rewards.

  • Agents must explore to find good strategies - and they will exploit any loophole in the reward (reward hacking).

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Apply Q-learning updates

+25 XP

Line 1 holds the learning rate α and the discount γ. Each following line is one experienced transition: state action reward next_state, where next_state is end when the episode finishes.

All Q-values start at 0. For each transition, apply Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s, a) \leftarrow Q(s, a) + \alpha [r + \gamma \max_{a'} Q(s', a') - Q(s, a)]. The future term is 0 when the next state is end or has no Q-values yet.

Finally, for each state in the order it first appeared, print state: best_action value with the value to 2 decimal places. On ties, pick the action that state tried first.

  • Value flows backwards
  • Learning rate 1
  • Avoid the pit
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Patient or impatient?

+25 XP

Line 1: the discount γ. Line 2: the rewards received at each step. Print return: G - the discounted return ∑tγtrt\sum_t \gamma^t r_t (first reward at t = 0) to 2 decimal places - and undiscounted: S, the plain sum formatted with :g.

  • The slow treasure
  • Halving
  • No discount
  • Only now matters
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: