--- theme: seriph title: "Lesson 04 — Why Doesn't a Stronger Model Make a Reliable Agent?" info: "English video course for AI Agents in Depth" author: Bojie Li transition: slide-left mdc: true lineNumbers: false monaco: false aspectRatio: 16/9 canvasWidth: 980 layout: cover class: cover ---
Build · Chapter 1 · Agent Fundamentals
# Why Doesn't a Stronger Model Make a Reliable Agent?

Harness engineering, orchestration, and guardrails

Lesson 04 of 42 · 17 minutes · Harness Engineering; Model Choice; Orchestration Patterns; Guardrails and Safety
--- layout: center class: text-center ---
The central question
If models keep improving, why does the software around them keep getting more important?
--- # Why this problem matters

Constrain

Permissions, budgets, and valid action boundaries

Verify

Independent evidence that work is actually complete

Recover

Retries, fallbacks, checkpoints, and termination paths

--- # Three ideas to keep in view

Context engineering

Control what the model can see.

Loop engineering

Control when the system continues or stops.

Harness engineering

Control the complete runtime around the model.

--- # The book's visual model The execution loop of an autonomous Agent
The execution loop of an autonomous Agent
--- # Workflow vs. Autonomous Agent

Workflow

Autonomous Agent

Use the least autonomous pattern that can solve the task.
--- # Verification must observe the world ~~~python proposal = agent.execute(task) evidence = environment.inspect(proposal) if not verifier.accepts(evidence): agent.revise(evidence) guardrails.check_before_commit() ~~~ --- # Test the claim
1-32 min

Inspect a search-and-code execution plan

Observe: Which work belongs to search, code, validation, and stopping logic

Demo budget: 2 minutes · one contiguous terminal block
--- class: course-terminal ---
Live demo
# Switching to the terminal ~~~bash $ uv run python chapter1/search-codegen/main.py --backend openai --dry-run --request "Compare ASEAN capitals" ~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
--- # What the evidence supports

Finding 1

Most production code handles boundaries and failures rather than the happy path.

Finding 2

Independent observations add information that self-reflection cannot.

Finding 3

Model selection should follow an evaluation, not a reputation.

--- layout: center ---
Where the claim stops
# Boundary condition
A Harness can patch unstable behavior, but it cannot make an unverifiable goal objectively verifiable.
--- layout: center ---
Engineering takeaway
# Design rule
Prompts first, workflows second, autonomous Agents only where adaptation creates real value.
--- # Continue the experiment
Workflow patterns book-en/images/fig1-wf-routing.svg Evaluator-optimizer workflow book-en/images/fig1-wf-evaluator.svg n8n workflow example book-en/images/n8n-workflow.png
--- layout: center class: text-center ---
Pause and apply
# Your turn
Which failure in your Agent should be prevented, detected, recovered, or escalated?
--- layout: center class: text-center ---
Chapter 1 complete · Next · Lesson 05
Move inside the context window and inspect what the API actually sends.