--- theme: seriph title: "Lesson 24 — Which Agent Should You Ship?" info: "English video course for AI Agents in Depth" author: Bojie Li transition: slide-left mdc: true lineNumbers: false monaco: false aspectRatio: 16/9 canvasWidth: 980 layout: cover class: cover ---
Improve · Chapter 6 · Agent Evaluation
# Which Agent Should You Ship?

Model behavior, latency, cost, and evaluation-driven selection

Lesson 24 of 42 · 19 minutes · Evaluation-Driven Model Selection; Model Behavior; Cost Analysis; Continuous Iteration
--- layout: center class: text-center ---
The central question
Why is the highest benchmark score not enough to choose a production Agent?
--- # Why this problem matters

Quality

Success, boundary behavior, and variance across task slices

Behavior

When the model searches, edits, retries, or stops

Economics

Latency, cache use, tokens, availability, and total task cost

--- # Three ideas to keep in view

Fixed Harness

Swap models to locate a model-side bottleneck

Ablation

Remove one Harness component to measure its contribution

Pareto frontier

Choose a non-dominated quality/cost/latency point

--- # The book's visual model Loop from benchmark results to system improvements
Loop from benchmark results to system improvements
--- # Leaderboard choice vs. Deployment choice

Leaderboard choice

Deployment choice

The unit of selection is model + context + tools + runtime.
--- # Filter before ranking ~~~python eligible = [r for r in runs if r.safety_pass] eligible = [r for r in eligible if r.p95_latency < sla] frontier = pareto(eligible, maximize='success', minimize='cost') winner = validate_on_holdout(frontier) ship_with_feature_flag(winner) ~~~ --- # Test the claim
6-82 min

Recompute a full Agent cost breakdown

Observe: Per-step cost, cache savings, compression savings, and non-additive interactions

6-72 min

Validate a fixed-Harness action-threshold experiment

Observe: Event-boundary accounting, first-edit timing, rework, and independent final tests

Demo budget: 4 minutes · one contiguous terminal block
--- class: course-terminal ---
Live demo
# Switching to the terminal ~~~bash $ cd chapter6/agent-cost-analysis && python demo.py --offline --scenario all $ python -m unittest discover -s chapter6/model-action-threshold/tests -v ~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
--- # What the evidence supports

Finding 1

Different models carry different default tool-use policies inside the same Harness.

Finding 2

Cache-friendly context and compression change cost without changing the task.

Finding 3

A model upgrade is a hypothesis that must clear your own gates.

--- layout: center ---
Where the claim stops
# Boundary condition
A smoke run verifies integration, not steady-state availability, tail latency, or statistical superiority.
--- layout: center ---
Engineering takeaway
# Design rule
Select on a domain-specific Pareto frontier after safety and reliability gates—not on a global leaderboard rank.
--- # Continue the experiment
Experiment 6-9: provider/model benchmark chapter6/model-benchmark/ Experiment 6-10: full memory component matrix chapter6/user-memory-system-evaluation/ Action-threshold canonical campaign chapter6/model-action-threshold/results/ Cost-analysis sample trace chapter6/agent-cost-analysis/sample_trace.json
--- layout: center class: text-center ---
Pause and apply
# Your turn
What is the first non-quality gate that would eliminate a model from your production shortlist?
--- layout: center class: text-center ---
Next · Lesson 25
Build evaluation infrastructure that can distinguish a real improvement from noise.