Files
ai-agent-book/slides/lesson-04.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

7.1 KiB

theme, title, info, author, transition, mdc, lineNumbers, monaco, aspectRatio, canvasWidth, layout, class
theme title info author transition mdc lineNumbers monaco aspectRatio canvasWidth layout class
seriph Lesson 04 — Why Doesn't a Stronger Model Make a Reliable Agent? English video course for AI Agents in Depth Bojie Li slide-left true false false 16/9 980 cover cover
Build · Chapter 1 · Agent Fundamentals

Why Doesn't a Stronger Model Make a Reliable Agent?

Harness engineering, orchestration, and guardrails

Lesson 04 of 42 · 17 minutes · Harness Engineering; Model Choice; Orchestration Patterns; Guardrails and Safety

layout: center class: text-center

The central question
If models keep improving, why does the software around them keep getting more important?

Why this problem matters

Constrain

Permissions, budgets, and valid action boundaries

Verify

Independent evidence that work is actually complete

Recover

Retries, fallbacks, checkpoints, and termination paths


Three ideas to keep in view

Context engineering

Control what the model can see.

Loop engineering

Control when the system continues or stops.

Harness engineering

Control the complete runtime around the model.


The book's visual model

The execution loop of an autonomous Agent
The execution loop of an autonomous Agent

Workflow vs. Autonomous Agent

Workflow

  • Known stages
  • Predictable control flow
  • Easy to inspect

Autonomous Agent

  • Open-ended plan
  • Adaptive tool use
  • Needs stronger verification
Use the least autonomous pattern that can solve the task.

Verification must observe the world

proposal = agent.execute(task)
evidence = environment.inspect(proposal)
if not verifier.accepts(evidence):
    agent.revise(evidence)
guardrails.check_before_commit()

Test the claim

1-32 min

Inspect a search-and-code execution plan

Observe: Which work belongs to search, code, validation, and stopping logic

Demo budget: 2 minutes · one contiguous terminal block

class: course-terminal

Live demo

Switching to the terminal

$ uv run python chapter1/search-codegen/main.py --backend openai --dry-run --request "Compare ASEAN capitals"
Run the command(s), narrate decisions, and point to the observation—not just the output.

What the evidence supports

Finding 1

Most production code handles boundaries and failures rather than the happy path.

Finding 2

Independent observations add information that self-reflection cannot.

Finding 3

Model selection should follow an evaluation, not a reputation.


layout: center

Where the claim stops

Boundary condition

A Harness can patch unstable behavior, but it cannot make an unverifiable goal objectively verifiable.

layout: center

Engineering takeaway

Design rule

Prompts first, workflows second, autonomous Agents only where adaptation creates real value.

Continue the experiment


layout: center class: text-center

Pause and apply

Your turn

Which failure in your Agent should be prevented, detected, recovered, or escalated?

layout: center class: text-center

Chapter 1 complete · Next · Lesson 05
Move inside the context window and inspect what the API actually sends.