---
theme: seriph
title: "Lesson 37 — How Does an Agent Act Through Pixels?"
info: "English video course for AI Agents in Depth"
author: Bojie Li
transition: slide-left
mdc: true
lineNumbers: false
monaco: false
aspectRatio: 16/9
canvasWidth: 980
layout: cover
class: cover
---
Expand · Chapter 9 · Multimodal Interaction
# How Does an Agent Act Through Pixels?
GUI action spaces, visual grounding, and bounded interaction
Lesson 37 of 42 · 18 minutes · Computer Use; Action Space Design; Visual Grounding; Real-Time Performance
---
layout: center
class: text-center
---
The central question
How does an Agent turn a screenshot and a goal into the right interface action?
---
# Why this problem matters
Observation
A screenshot is a partial, time-sensitive view of application state.
Grounding
The Agent must map language to an element ID or coordinate.
Interaction
Every click or keystroke changes the next observation.
---
# Three ideas to keep in view
Structured tree
Use DOM or accessibility elements when they are reliable
Visual grounding
Locate targets directly in pixels when structure is absent
Bounded loop
Observe → one guarded action → observe again
---
# The book's visual model
Visual grounding with annotated interface elements
---
# Structured grounding vs. Visual grounding
Structured grounding
- DOM/accessibility IDs
- Closed-set selection
- Fails on custom drawing
Visual grounding
- Works from pixels
- General interface coverage
- Coordinate and scale errors
Production systems need both paths, coordinate transforms, and a confidence-aware fallback.
---
# Every action creates a new observation
~~~python
while budget.remaining:
screenshot, tree = browser.observe()
target = agent.ground(task, screenshot, tree)
action = agent.choose_action(target)
receipt = browser.execute(guard(action))
if verifier.done(receipt): break
~~~
---
# Test the claim
9-6 preflight2 min
Run offline Computer Use contract checks
Observe: Endpoint identity, screenshot retention, manifest integrity, and redaction boundaries
9-6 retained status1 min
Inspect the retained open-model acceptance pointer
Observe: Experiment arm, model scope, status, and the hashes required to trace the full run
Demo budget: 3 minutes · one contiguous terminal block
---
class: course-terminal
---
Live demo
# Switching to the terminal
~~~bash
$ python -m pytest chapter9/computer-use-open-model/tests -q
$ python -m json.tool chapter9/computer-use-open-model/validation/latest.json
~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
---
# What the evidence supports
Finding 1
A Computer Use result is a trajectory of state changes—not a final textual answer.
Finding 2
Structured elements improve precision, while pixel grounding expands interface coverage.
Finding 3
Scaling, coordinate transforms, stale screenshots, and hidden state create distinct failure modes.
---
layout: center
---
Where the claim stops
# Boundary condition
Offline contract tests and an acceptance pointer do not reproduce the retained browser trajectory or complete the separate Anthropic-native arm.
---
layout: center
---
Engineering takeaway
# Design rule
Execute one bounded GUI action at a time, re-observe after every state change, and retain screenshots plus external outcome evidence.
---
# Continue the experiment
---
layout: center
class: text-center
---
Pause and apply
# Your turn
Which state change would prove that your GUI action succeeded, even if the Agent claims it did?
---
layout: center
class: text-center
---
Next · Lesson 38
Cross from visual interfaces into physical control, where latency and mistakes have mechanical consequences.
→