---
theme: seriph
title: "Lesson 32 — How Do Failed Trajectories Become Learning Signals?"
info: "English video course for AI Agents in Depth"
author: Bojie Li
transition: slide-left
mdc: true
lineNumbers: false
monaco: false
aspectRatio: 16/9
canvasWidth: 980
layout: cover
class: cover
---
Improve · Chapter 8 · Continual Evolution
# How Do Failed Trajectories Become Learning Signals?
Outcome verification, process rules, Rubrics, and cross-trajectory experience
Lesson 32 of 42 · 19 minutes · Deriving Learning Signals from Operational Trajectories; Consolidating Experience into Knowledge
---
Improve · Chapter 8 · Continual Evolution
# Problems this chapter will solve
Lesson 32
How Do Failed Trajectories Become Learning Signals?
Lesson 33
Where Should an Agent Store What It Learns?
Lesson 34
How Can a Self-Modifying Agent Change Without Drifting?
---
# Why this problem matters
Outcome
Read what changed in the environment.
Process
Locate rule violations and ineffective decisions.
Meaning
Use a Rubric for dimensions that code cannot settle.
---
# Three ideas to keep in view
Trajectory verifier
Outcome checks + process rules + language Rubric
Contrastive evidence
Compare success, partial success, and failure
Experience document
Mechanism + conditions + evidence + exceptions
---
# The book's visual model
Three-layer trajectory verification from outcomes to an LLM Rubric
---
# Save the trajectory vs. Consolidate experience
Save the trajectory
- High detail
- Hard to retrieve
- Incidental actions become noise
Consolidate experience
- Cross-run pattern
- Explicit applicability
- Evidence and counterexamples
A trajectory is evidence; it is not yet a lesson.
---
# Diagnose before updating
~~~python
outcome = environment_verifier(trajectory)
violations = process_verifier(trajectory)
rubric = semantic_judge(trajectory, outcome)
diagnosis = triangulate(outcome, violations, rubric)
experience = consolidate(similar_diagnoses)
~~~
---
# Test the claim
8-12 min
Diagnose customer-service trajectories with three evidence layers
Observe: False promises, privacy violations, over-refusal, and cited evidence
8-22 min
Consolidate several trajectories into experience documents
Observe: Transfer gain, retrieval cost, negative transfer, and applicability conditions
Demo budget: 4 minutes · one contiguous terminal block
---
class: course-terminal
---
Live demo
# Switching to the terminal
~~~bash
$ cd chapter8/trajectory-verifier && python demo.py
$ cd chapter8/gaia-experience && python demo_documents.py
~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
---
# What the evidence supports
Finding 1
Environment outcomes constrain what a language judge may claim.
Finding 2
Failures and partial successes reveal conditions hidden by successful runs.
Finding 3
Cross-trajectory documents can transfer while using fewer tokens than raw history.
---
layout: center
---
Where the claim stops
# Boundary condition
A pattern supported by past trajectories may become obsolete after an API, policy, or environment change.
---
layout: center
---
Engineering takeaway
# Design rule
Promote experience only with provenance, applicability conditions, counterevidence, and a revalidation trigger.
---
# Continue the experiment
---
layout: center
class: text-center
---
Pause and apply
# Your turn
Which detail in a successful trajectory was causal, and how would you distinguish it from coincidence?
---
layout: center
class: text-center
---
Next · Lesson 33
Choose the artifact that should change: knowledge, instructions, programs, or parameters.
→