---
theme: seriph
title: "Lesson 24 — Which Agent Should You Ship?"
info: "English video course for AI Agents in Depth"
author: Bojie Li
transition: slide-left
mdc: true
lineNumbers: false
monaco: false
aspectRatio: 16/9
canvasWidth: 980
layout: cover
class: cover
---
Improve · Chapter 6 · Agent Evaluation
# Which Agent Should You Ship?
Model behavior, latency, cost, and evaluation-driven selection
Lesson 24 of 42 · 19 minutes · Evaluation-Driven Model Selection; Model Behavior; Cost Analysis; Continuous Iteration
---
layout: center
class: text-center
---
The central question
Why is the highest benchmark score not enough to choose a production Agent?
---
# Why this problem matters
Quality
Success, boundary behavior, and variance across task slices
Behavior
When the model searches, edits, retries, or stops
Economics
Latency, cache use, tokens, availability, and total task cost
---
# Three ideas to keep in view
Fixed Harness
Swap models to locate a model-side bottleneck
Ablation
Remove one Harness component to measure its contribution
Pareto frontier
Choose a non-dominated quality/cost/latency point
---
# The book's visual model
Loop from benchmark results to system improvements
---
# Leaderboard choice vs. Deployment choice
Leaderboard choice
- One public score
- Unknown Harness
- Average case
Deployment choice
- Your task distribution
- Your complete Harness
- Cost and failure boundaries
The unit of selection is model + context + tools + runtime.
---
# Filter before ranking
~~~python
eligible = [r for r in runs if r.safety_pass]
eligible = [r for r in eligible if r.p95_latency < sla]
frontier = pareto(eligible, maximize='success', minimize='cost')
winner = validate_on_holdout(frontier)
ship_with_feature_flag(winner)
~~~
---
# Test the claim
6-82 min
Recompute a full Agent cost breakdown
Observe: Per-step cost, cache savings, compression savings, and non-additive interactions
6-72 min
Validate a fixed-Harness action-threshold experiment
Observe: Event-boundary accounting, first-edit timing, rework, and independent final tests
Demo budget: 4 minutes · one contiguous terminal block
---
class: course-terminal
---
Live demo
# Switching to the terminal
~~~bash
$ cd chapter6/agent-cost-analysis && python demo.py --offline --scenario all
$ python -m unittest discover -s chapter6/model-action-threshold/tests -v
~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
---
# What the evidence supports
Finding 1
Different models carry different default tool-use policies inside the same Harness.
Finding 2
Cache-friendly context and compression change cost without changing the task.
Finding 3
A model upgrade is a hypothesis that must clear your own gates.
---
layout: center
---
Where the claim stops
# Boundary condition
A smoke run verifies integration, not steady-state availability, tail latency, or statistical superiority.
---
layout: center
---
Engineering takeaway
# Design rule
Select on a domain-specific Pareto frontier after safety and reliability gates—not on a global leaderboard rank.
---
# Continue the experiment
---
layout: center
class: text-center
---
Pause and apply
# Your turn
What is the first non-quality gate that would eliminate a model from your production shortlist?
---
layout: center
class: text-center
---
Next · Lesson 25
Build evaluation infrastructure that can distinguish a real improvement from noise.
→