--- theme: seriph title: "Lesson 15 — How Do You Let an Agent Act Without Letting It Cause Damage?" info: "English video course for AI Agents in Depth" author: Bojie Li transition: slide-left mdc: true lineNumbers: false monaco: false aspectRatio: 16/9 canvasWidth: 980 layout: cover class: cover ---
Build · Chapter 4 · Tools
# How Do You Let an Agent Act Without Letting It Cause Damage?

Execution tools, independent checks, and fail-closed design

Lesson 15 of 42 · 18 minutes · Execution Tools; Security; Proposer-Reviewer; Sidecar
--- layout: center class: text-center ---
The central question
Where should safety checks live when the model can write files, run code, and call external systems?
--- # Why this problem matters

Risk classification

Read, reversible write, irreversible action

Pre-approval

Review intent and parameters before execution

Post-validation

Inspect the actual resulting state

--- # Three ideas to keep in view

Fail closed

Unknown or malformed operations are denied

Independent evidence

Use data the proposer cannot forge

Sidecar

Keep enforcement outside the Agent's own mutable process

--- # The book's visual model Synchronous model training versus asynchronous deployment
Synchronous model training versus asynchronous deployment
--- # Model self-report vs. Independent gate

Model self-report

Independent gate

The final boundary must not trust the model's own claim.
--- # Server truth is the gatekeeper ~~~python request = agent.propose_action() facts = database.read_ground_truth(request.target) policy.validate(request, facts) result = executor.run(request) validator.inspect(result) ~~~ --- # Test the claim
4-3A1 min

Run an allowed code action

Observe: Validation, sandbox execution, and bounded output

4-3B2 min

Inspect execution-tool safety behavior

Observe: Approval, rejection, syntax checks, and output handling

Demo budget: 3 minutes · one contiguous terminal block
--- class: course-terminal ---
Live demo
# Switching to the terminal ~~~bash $ uv run python chapter4/execution-tools/cli.py code --language python --code "print(2 ** 10)" $ uv run python chapter4/execution-tools/cli.py demo ~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
--- # What the evidence supports

Finding 1

Risk depends on parameters and environment, not only the tool name.

Finding 2

Pre-approval reduces harmful attempts; validation catches harmful results.

Finding 3

Long outputs need truncation plus durable storage, not silent loss.

--- layout: center ---
Where the claim stops
# Boundary condition
A second model is not independent if it sees the same injected context and trusts the same unverified facts.
--- layout: center ---
Engineering takeaway
# Design rule
Guide the model with instructions, but enforce irreversible constraints with independent code and data.
--- # Continue the experiment
Execution-tool tests chapter4/execution-tools/ Sidecar design book-en/chapter4.md
--- layout: center class: text-center ---
Pause and apply
# Your turn
What is the trusted root in your Agent system, and can the Agent modify it?
--- layout: center class: text-center ---
Next · Lesson 16
Some tasks require another Agent or a human rather than another tool.