--- theme: seriph title: "Lesson 37 — How Does an Agent Act Through Pixels?" info: "English video course for AI Agents in Depth" author: Bojie Li transition: slide-left mdc: true lineNumbers: false monaco: false aspectRatio: 16/9 canvasWidth: 980 layout: cover class: cover ---
Expand · Chapter 9 · Multimodal Interaction
# How Does an Agent Act Through Pixels?

GUI action spaces, visual grounding, and bounded interaction

Lesson 37 of 42 · 18 minutes · Computer Use; Action Space Design; Visual Grounding; Real-Time Performance
--- layout: center class: text-center ---
The central question
How does an Agent turn a screenshot and a goal into the right interface action?
--- # Why this problem matters

Observation

A screenshot is a partial, time-sensitive view of application state.

Grounding

The Agent must map language to an element ID or coordinate.

Interaction

Every click or keystroke changes the next observation.

--- # Three ideas to keep in view

Structured tree

Use DOM or accessibility elements when they are reliable

Visual grounding

Locate targets directly in pixels when structure is absent

Bounded loop

Observe → one guarded action → observe again

--- # The book's visual model Visual grounding with annotated interface elements
Visual grounding with annotated interface elements
--- # Structured grounding vs. Visual grounding

Structured grounding

Visual grounding

Production systems need both paths, coordinate transforms, and a confidence-aware fallback.
--- # Every action creates a new observation ~~~python while budget.remaining: screenshot, tree = browser.observe() target = agent.ground(task, screenshot, tree) action = agent.choose_action(target) receipt = browser.execute(guard(action)) if verifier.done(receipt): break ~~~ --- # Test the claim
9-6 preflight2 min

Run offline Computer Use contract checks

Observe: Endpoint identity, screenshot retention, manifest integrity, and redaction boundaries

9-6 retained status1 min

Inspect the retained open-model acceptance pointer

Observe: Experiment arm, model scope, status, and the hashes required to trace the full run

Demo budget: 3 minutes · one contiguous terminal block
--- class: course-terminal ---
Live demo
# Switching to the terminal ~~~bash $ python -m pytest chapter9/computer-use-open-model/tests -q $ python -m json.tool chapter9/computer-use-open-model/validation/latest.json ~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
--- # What the evidence supports

Finding 1

A Computer Use result is a trajectory of state changes—not a final textual answer.

Finding 2

Structured elements improve precision, while pixel grounding expands interface coverage.

Finding 3

Scaling, coordinate transforms, stale screenshots, and hidden state create distinct failure modes.

--- layout: center ---
Where the claim stops
# Boundary condition
Offline contract tests and an acceptance pointer do not reproduce the retained browser trajectory or complete the separate Anthropic-native arm.
--- layout: center ---
Engineering takeaway
# Design rule
Execute one bounded GUI action at a time, re-observe after every state change, and retain screenshots plus external outcome evidence.
--- # Continue the experiment
Experiment 9-5: Claude Computer Use chapter9/claude-quickstarts/computer-use-demo/ Experiment 9-6: open-model Computer Use chapter9/computer-use-open-model/ Screenshot-action loop book-en/images/fig9-7.svg Coordinate scaling book-en/images/fig9-10.svg
--- layout: center class: text-center ---
Pause and apply
# Your turn
Which state change would prove that your GUI action succeeded, even if the Agent claims it did?
--- layout: center class: text-center ---
Next · Lesson 38
Cross from visual interfaces into physical control, where latency and mistakes have mechanical consequences.