Um momento
0x40Lesson 5 of 6

Evaluate the whole loop

Measure task outcomes and diagnose failures across model, tools, and control logic.

12 min 5-question quiz
By the end of this lesson you can
  • Create multi-step evaluations that reveal where a loop succeeds or fails.

A loop’s quality depends on the complete path, not only one model response. Evaluate whether it reaches the goal, uses the right tools, handles failures, respects limits, and avoids harmful side effects. Save traces of decisions, calls, results, and stop reasons with sensitive data minimized. Include ordinary tasks, edge cases, timeouts, malformed tool responses, and cases that require a human handoff.

A small example

Illustrative Python
1expected = {"goal_met": True, "steps_at_most": 3}
2run = {"goal_met": True, "steps": 2}
3passed = run["goal_met"] == expected["goal_met"] and run["steps"] <= expected["steps_at_most"]
4print("pass" if passed else "review")
Output
pass

Use clear task success criteria and repeat runs where model behavior varies. Inspect representative traces to locate whether a failure came from planning, tool selection, execution, observation parsing, or stopping logic. Test changes against the same cases so improvements do not hide regressions.

Key takeaways

  • Create multi-step evaluations that reveal where a loop succeeds or fails.

  • Bound the loop, validate actions, and make its outcome observable.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: