Files
ai-agent-book/chapter10/talkact-reproduction/README.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

105 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Experiment 10-3 · Fixed-topology TalkAct baseline
This is the fixed-topology comparison arm of Chapter Experiment 10-3. Its
validation artifacts use the current `10-3` identifier; the autonomous arm is
implemented in [`autonomous-phone-registration`](../autonomous-phone-registration/).
This record covers the pinned external TalkAct reproduction used by current
Experiment 10-3. The comparison runs concurrent fast/slow agents (`duplex`) against a
single-model control (`strawman`) over four hermetic tasks and two labeled
repetitions per task and condition.
Status: **complete for the retained Anthropic-caller configuration**. The
[canonical run](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/)
contains all 16 episode logs, aggregate and per-episode analysis, the exact
protocol and environment, console logs, and a manifest. The independent
validator passes all 17 gates.
## Configuration and deviation
The campaign used the official `19PINE-AI/TalkAct` source at commit
`7d70007f72d45ddfc1a14e8e229b6d444e4919a2`, Python 3.12.11, Playwright
Chromium 149, and the hermetic Flask task server. The concurrent arm kept the
pinned source's fast `claude-haiku-4-5` and slow `claude-opus-4-8` agents; the
strawman arm kept the slow model alone.
TalkAct normally uses Gemini for its simulated caller, but the configured
Gemini credential returned `400 API_KEY_INVALID`. The campaign therefore used
the source-supported `CUV_USER_MODEL=claude-sonnet-4-5-20250929` override. This
preserves the task and agent topology but makes the caller Anthropic-based too,
which can introduce same-family bias. These results are a distinct
Anthropic-caller configuration and must not be silently pooled with upstream
results that use the default Gemini caller.
The exact campaign was:
```bash
CUV_USER_MODEL=claude-sonnet-4-5-20250929 python bench/run_bench.py \
--tasks forms-insurance booking-flight webmail-report meeting-helper \
--conditions duplex strawman \
--seeds 2
```
## Results
| Metric | Duplex | Strawman |
| --- | ---: | ---: |
| Episodes | 8 | 8 |
| Task success | 1.000 | 1.000 |
| Partial credit | 1.000 | 1.000 |
| Probe correctness | 0.833 | 0.917 |
| Voice latency p50 | 2.32 s | 12.52 s |
| Voice latency p90 | 2.85 s | 21.14 s |
| Voice latency maximum | 4.03 s | 37.29 s |
| Mean episode wall time | 207.2 s | 178.0 s |
| Voice-latency samples | 47 | 44 |
All 12 action-scored episodes achieved full deterministic success and partial
credit. The four `meeting-helper` episodes intentionally have no action-success
field and instead earned perfect retained probe-answer scores. There were no
episode or provider errors.
The concurrent arm reduced median voice latency by about **5.40×**
(`12.52 / 2.32`), materially below the upstream roughly 15× result. It did not
improve task success: the arms tied. The control also had higher aggregate
probe correctness (0.917 versus 0.833) and lower mean wall time (178.0 seconds
versus 207.2 seconds). The retained run therefore supports a response-latency
advantage for duplex, not a blanket quality or total-runtime advantage.
## Evidence and validation
The 16 raw episodes retain 39 fast-to-slow relays, 33 slow-to-fast events, 91
voice-latency samples, deterministic task checks, probe grades, transcripts,
environment state, and aggregate provider usage. The duplex slow tier used
19,124 input and 26,105 output tokens across 166 steps; its fast tier used
98,945 input and 5,782 output tokens across 87 turns. The strawman slow tier
used 19,208 input and 28,966 output tokens across 177 steps. Cache-read and
cache-creation counts are retained in the episode records.
Run the validator from the repository root:
```bash
python chapter10/talkact-reproduction/validate_campaign.py \
chapter10/talkact-reproduction/validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2
```
The generated [acceptance report](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/acceptance.json)
passes source-pin, campaign-shape, model, usage, error, concurrency, bridge,
latency, task-check, judge, aggregate, and credential-scan gates. The
[manifest](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/manifest.json)
hashes the 23 inputs and outputs from which those generated files are derived.
The earlier [authentication preflight](validation/exp10-3-anthropic-auth-20260803-v1/preflight.json)
is retained as failure history, not as the final result.
## Limitations
- The pinned runner's `--seeds` option labels repetitions but does not inject a
deterministic random seed into the episode runner or provider calls.
- The pinned source retains model labels and aggregate fast/slow token usage
per episode, but not individual provider response IDs.
- Simulated-caller token usage is not retained by the pinned source.
- `meeting-helper` is evaluated through retained probe grades rather than the
action-based success and partial-credit fields used by the other tasks.
- The Anthropic caller deviation preserves the benchmark topology but is not
directly comparable to the upstream default-Gemini caller configuration.