Files
ai-agent-book/slides/lesson-24.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

7.9 KiB

theme, title, info, author, transition, mdc, lineNumbers, monaco, aspectRatio, canvasWidth, layout, class
theme title info author transition mdc lineNumbers monaco aspectRatio canvasWidth layout class
seriph Lesson 24 — Which Agent Should You Ship? English video course for AI Agents in Depth Bojie Li slide-left true false false 16/9 980 cover cover
Improve · Chapter 6 · Agent Evaluation

Which Agent Should You Ship?

Model behavior, latency, cost, and evaluation-driven selection

Lesson 24 of 42 · 19 minutes · Evaluation-Driven Model Selection; Model Behavior; Cost Analysis; Continuous Iteration

layout: center class: text-center

The central question
Why is the highest benchmark score not enough to choose a production Agent?

Why this problem matters

Quality

Success, boundary behavior, and variance across task slices

Behavior

When the model searches, edits, retries, or stops

Economics

Latency, cache use, tokens, availability, and total task cost


Three ideas to keep in view

Fixed Harness

Swap models to locate a model-side bottleneck

Ablation

Remove one Harness component to measure its contribution

Pareto frontier

Choose a non-dominated quality/cost/latency point


The book's visual model

Loop from benchmark results to system improvements
Loop from benchmark results to system improvements

Leaderboard choice vs. Deployment choice

Leaderboard choice

  • One public score
  • Unknown Harness
  • Average case

Deployment choice

  • Your task distribution
  • Your complete Harness
  • Cost and failure boundaries
The unit of selection is model + context + tools + runtime.

Filter before ranking

eligible = [r for r in runs if r.safety_pass]
eligible = [r for r in eligible if r.p95_latency < sla]
frontier = pareto(eligible, maximize='success', minimize='cost')
winner = validate_on_holdout(frontier)
ship_with_feature_flag(winner)

Test the claim

6-82 min

Recompute a full Agent cost breakdown

Observe: Per-step cost, cache savings, compression savings, and non-additive interactions

6-72 min

Validate a fixed-Harness action-threshold experiment

Observe: Event-boundary accounting, first-edit timing, rework, and independent final tests

Demo budget: 4 minutes · one contiguous terminal block

class: course-terminal

Live demo

Switching to the terminal

$ cd chapter6/agent-cost-analysis && python demo.py --offline --scenario all

$ python -m unittest discover -s chapter6/model-action-threshold/tests -v
Run the command(s), narrate decisions, and point to the observation—not just the output.

What the evidence supports

Finding 1

Different models carry different default tool-use policies inside the same Harness.

Finding 2

Cache-friendly context and compression change cost without changing the task.

Finding 3

A model upgrade is a hypothesis that must clear your own gates.


layout: center

Where the claim stops

Boundary condition

A smoke run verifies integration, not steady-state availability, tail latency, or statistical superiority.

layout: center

Engineering takeaway

Design rule

Select on a domain-specific Pareto frontier after safety and reliability gates—not on a global leaderboard rank.

Continue the experiment


layout: center class: text-center

Pause and apply

Your turn

What is the first non-quality gate that would eliminate a model from your production shortlist?

layout: center class: text-center

Next · Lesson 25
Build evaluation infrastructure that can distinguish a real improvement from noise.