ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
This commit is contained in:
@@ -0,0 +1,104 @@
|
||||
# Experiment 10-3 · Fixed-topology TalkAct baseline
|
||||
|
||||
This is the fixed-topology comparison arm of Chapter Experiment 10-3. Its
|
||||
validation artifacts use the current `10-3` identifier; the autonomous arm is
|
||||
implemented in [`autonomous-phone-registration`](../autonomous-phone-registration/).
|
||||
|
||||
This record covers the pinned external TalkAct reproduction used by current
|
||||
Experiment 10-3. The comparison runs concurrent fast/slow agents (`duplex`) against a
|
||||
single-model control (`strawman`) over four hermetic tasks and two labeled
|
||||
repetitions per task and condition.
|
||||
|
||||
Status: **complete for the retained Anthropic-caller configuration**. The
|
||||
[canonical run](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/)
|
||||
contains all 16 episode logs, aggregate and per-episode analysis, the exact
|
||||
protocol and environment, console logs, and a manifest. The independent
|
||||
validator passes all 17 gates.
|
||||
|
||||
## Configuration and deviation
|
||||
|
||||
The campaign used the official `19PINE-AI/TalkAct` source at commit
|
||||
`7d70007f72d45ddfc1a14e8e229b6d444e4919a2`, Python 3.12.11, Playwright
|
||||
Chromium 149, and the hermetic Flask task server. The concurrent arm kept the
|
||||
pinned source's fast `claude-haiku-4-5` and slow `claude-opus-4-8` agents; the
|
||||
strawman arm kept the slow model alone.
|
||||
|
||||
TalkAct normally uses Gemini for its simulated caller, but the configured
|
||||
Gemini credential returned `400 API_KEY_INVALID`. The campaign therefore used
|
||||
the source-supported `CUV_USER_MODEL=claude-sonnet-4-5-20250929` override. This
|
||||
preserves the task and agent topology but makes the caller Anthropic-based too,
|
||||
which can introduce same-family bias. These results are a distinct
|
||||
Anthropic-caller configuration and must not be silently pooled with upstream
|
||||
results that use the default Gemini caller.
|
||||
|
||||
The exact campaign was:
|
||||
|
||||
```bash
|
||||
CUV_USER_MODEL=claude-sonnet-4-5-20250929 python bench/run_bench.py \
|
||||
--tasks forms-insurance booking-flight webmail-report meeting-helper \
|
||||
--conditions duplex strawman \
|
||||
--seeds 2
|
||||
```
|
||||
|
||||
## Results
|
||||
|
||||
| Metric | Duplex | Strawman |
|
||||
| --- | ---: | ---: |
|
||||
| Episodes | 8 | 8 |
|
||||
| Task success | 1.000 | 1.000 |
|
||||
| Partial credit | 1.000 | 1.000 |
|
||||
| Probe correctness | 0.833 | 0.917 |
|
||||
| Voice latency p50 | 2.32 s | 12.52 s |
|
||||
| Voice latency p90 | 2.85 s | 21.14 s |
|
||||
| Voice latency maximum | 4.03 s | 37.29 s |
|
||||
| Mean episode wall time | 207.2 s | 178.0 s |
|
||||
| Voice-latency samples | 47 | 44 |
|
||||
|
||||
All 12 action-scored episodes achieved full deterministic success and partial
|
||||
credit. The four `meeting-helper` episodes intentionally have no action-success
|
||||
field and instead earned perfect retained probe-answer scores. There were no
|
||||
episode or provider errors.
|
||||
|
||||
The concurrent arm reduced median voice latency by about **5.40×**
|
||||
(`12.52 / 2.32`), materially below the upstream roughly 15× result. It did not
|
||||
improve task success: the arms tied. The control also had higher aggregate
|
||||
probe correctness (0.917 versus 0.833) and lower mean wall time (178.0 seconds
|
||||
versus 207.2 seconds). The retained run therefore supports a response-latency
|
||||
advantage for duplex, not a blanket quality or total-runtime advantage.
|
||||
|
||||
## Evidence and validation
|
||||
|
||||
The 16 raw episodes retain 39 fast-to-slow relays, 33 slow-to-fast events, 91
|
||||
voice-latency samples, deterministic task checks, probe grades, transcripts,
|
||||
environment state, and aggregate provider usage. The duplex slow tier used
|
||||
19,124 input and 26,105 output tokens across 166 steps; its fast tier used
|
||||
98,945 input and 5,782 output tokens across 87 turns. The strawman slow tier
|
||||
used 19,208 input and 28,966 output tokens across 177 steps. Cache-read and
|
||||
cache-creation counts are retained in the episode records.
|
||||
|
||||
Run the validator from the repository root:
|
||||
|
||||
```bash
|
||||
python chapter10/talkact-reproduction/validate_campaign.py \
|
||||
chapter10/talkact-reproduction/validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2
|
||||
```
|
||||
|
||||
The generated [acceptance report](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/acceptance.json)
|
||||
passes source-pin, campaign-shape, model, usage, error, concurrency, bridge,
|
||||
latency, task-check, judge, aggregate, and credential-scan gates. The
|
||||
[manifest](validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/manifest.json)
|
||||
hashes the 23 inputs and outputs from which those generated files are derived.
|
||||
The earlier [authentication preflight](validation/exp10-3-anthropic-auth-20260803-v1/preflight.json)
|
||||
is retained as failure history, not as the final result.
|
||||
|
||||
## Limitations
|
||||
|
||||
- The pinned runner's `--seeds` option labels repetitions but does not inject a
|
||||
deterministic random seed into the episode runner or provider calls.
|
||||
- The pinned source retains model labels and aggregate fast/slow token usage
|
||||
per episode, but not individual provider response IDs.
|
||||
- Simulated-caller token usage is not retained by the pinned source.
|
||||
- `meeting-helper` is evaluated through retained probe grades rather than the
|
||||
action-based success and partial-credit fields used by the other tasks.
|
||||
- The Anthropic caller deviation preserves the benchmark topology but is not
|
||||
directly comparable to the upstream default-Gemini caller configuration.
|
||||
Reference in New Issue
Block a user