Files
ai-agent-book/docs/EXPERIMENT_STATUS.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

91 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Experiment status and evidence
This operational record is kept separate from the book README. It tracks the
cross-chapter experiments that require special local implementations, evidence
gates, external services, or hardware. Cloning a pinned source repository,
installing its dependencies, or passing a smoke test does not establish that an
experiment is complete.
Statuses in this file describe evidence retained in the repository. A local
checkpoint or run directory that is still untracked is useful work-in-progress,
but it does not close the clean-clone audit until a reviewable evidence package
is committed.
Status meanings:
- **Complete**: the manuscript's execution and evidence gates have substantive
saved evidence. A complete experiment may still produce a negative result.
- **Incomplete**: some implementation, execution, or evidence gates remain
unsatisfied.
- **Reader exercise**: completion requires a reader-operated campaign and its
retained evidence rather than a repository checkout alone.
This table is a selective operational ledger, not a second numbered index.
Each chapter README remains the authoritative ordered experiment list. The
WebRTC phone project is listed as an unnumbered add-on and uses a stable,
non-numeric evidence identifier.
## Tracked experiments
| Experiment | Track | Current status and evidence |
| --- | --- | --- |
| 4-1 | External-service acceptance | **Incomplete.** The real MCP catalog covers the public-data, multimodal, and filesystem gates, but Google Calendar and Notion remain blocked by missing authorized credentials. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). |
| 4-2 | Multimodal comparison | **Complete.** The canonical run compares native vision, extract-to-text, and tool-on-demand processing over the same chart/PDF questions with retained model receipts, latency/usage, tool traces, and an external judge. See the [canonical evidence](../chapter4/multimodal-agent/validation/latest.json). |
| 4-3 | External-service acceptance | **Incomplete only at Calendar/email authorization.** The canonical 20-call MCP campaign passes 13/15 gates: core safety, sandboxing, spreadsheet rendering, webhook, headless browser, real GitHub PR mutation, Xvfb Computer Use, and KVM-backed Android execution. Calendar and real email-provider mutations remain unsatisfied, so `official_complete` stays false. See the [canonical manifest](../chapter4/execution-tools/validation/experiment_4_3/real_mcp_gui_20260802T093657Z/manifest.json). |
| 4-4 | Human/external-channel acceptance | **Incomplete only at real notification delivery.** The canonical run retains a live repository-user decision on the same pending MCP request, a separate conservative timeout, six real Kimi K3 receipts, 55 MCP calls, and 61 verified manifest files. Email, Telegram, and Slack remain unconfigured and are not simulated, so `official_complete` stays false. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). |
| 4-5 | Active tool discovery | **Complete; accuracy uplift not observed.** The 126-tool campaign completed both control and active-discovery arms with 3/3 task success in each arm. Active discovery reduced exposed schema text and elapsed time, while retained detours and malformed actions remain visible. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). |
| 5-12 | Local implementation | **Incomplete.** The PEDO core, deterministic PostgreSQL demo, and tests are included in [the companion project](../chapter5/permission-embedded-data-objects/). The optional live Agent-generated-code campaign has not been run as a canonical evidence package. |
| 5-13 | Local experiment | **Complete; strict joint-advantage hypothesis not observed.** Both Agent-creation arms passed their acceptance gates. The [formal comparison](../chapter5/agent-creator/runs/exp5-12-kimi-k3-20260730-v1/comparison.json) found equal deterministic quality and greater template-arm efficiency, but not strictly higher quality and efficiency together. |
| 6-1 | External mailbox experiment | **Incomplete.** The Unipile credential probe returned 401 before mailbox mutation, so the three required real inbound-mail cases have not run. See the [project evidence](../chapter6/agent-with-event-trigger/validation/experiment_6_1/). |
| 6-2 | Interruptible asynchronous Agent | **Complete.** All four manuscript scenarios passed with real subprocesses: non-blocking work, queued instruction integration, interruption and recovery, and progress-aware cancellation. See the [canonical summary](../chapter6/async-agent/validation/experiment_6_2/real_subprocess_20260730T052500Z/summary.json). |
| 附加 | Local WebRTC speech project | **Complete.** Direct and ReAct arms each pass 20/20 gates over browser-microphone RTP, local Whisper, a real external LLM, TTS, and downlink RTP. PSTN/E.164 is outside the manuscript's local call-user gate. See the [stable-ID manifest](../chapter6/phone-agent/validation/runs/phone-agent-webrtc-audio-20260731-v1/manifest.json). |
| 6-5 | Local omni-speech experiment | **Complete; the two paths tie overall with complementary failures.** Pinned MiniCPM-o 4.5 ran locally on one RTX PRO 6000. Native end-to-end and same-model self-cascade each scored 3/4: self-cascade fixed one semantic perception error, while end-to-end preserved speaking-rate evidence erased by the transcript. The [canonical evidence](../chapter6/end-to-end-speech/validation/runs/exp6-5-minicpmo45-20260801-v1/evidence.json) also retains a real 24kHz speech output and passes all acceptance checks. |
| 6-6 | Local controllable-TTS experiment | **Complete; the full subjective ordering was not reproduced.** Fish Audio S1 produced the 24-reference library and A/B/C media, and three position-balanced Voxtral listening passes rated the multi-reference arm highest. C > B > A did not reproduce because A outscored B. See the [acceptance evidence](../chapter6/controllable-tts/validation/acceptance.json). |
| 6-7 | External reference implementation | **Complete for the bounded read-only task.** The official source and Dockerfile were pinned, the image was built locally with retained image/base digests, and Anthropic returned the requested `claude-sonnet-4-5-20250929` on 16/16 calls. The Agent executed 15 native `computer` actions, did not interact with Google reCAPTCHA, recovered through visible Open-Meteo JSON, and reported 70.2°F / clear sky with `end_turn`. The [canonical acceptance](../chapter6/claude-computer-use-native/validation/runs/exp6-7-anthropic-native-20260803-v2/acceptance.json) passes every source, receipt, action, screenshot, grounding, safety, hash, and credential gate; the historical 401 and two failed task attempts remain separately retained. |
| 6-8 | Provider-portable Computer Use | **Complete on the open-model arm.** OpenRouter returned `qwen/qwen3-vl-32b-instruct` for 16/16 real calls; the Agent recovered from a Google CAPTCHA through weather.com and completed in 16 one-action steps. The [canonical evidence](../chapter6/computer-use-open-model/validation/latest.json) retains 15 screenshots, raw responses, the action trajectory, deterministic answer grounding, hashes, and a clean credential scan. |
| 6-9 | External hardware track | **Incomplete.** The XLeRobot source and non-actuating preflight are pinned, but no authorized physical teleoperation or book task has run. See the [experiment record](../chapter6/xlerobot-teleoperation/README.md). |
| 6-11 | External API and hardware track | **Incomplete.** The exact Gemini Robotics-ER request failed authentication and no robot navigation occurred. A successful planning response, authorized navigation run, and the remaining evidence gates are still required. See the [experiment record](../chapter6/gemini-xlerobot-navigation/README.md). |
| 6-13 | External Sim2Real track | **Incomplete.** No local ManiSkill environment, RGB-only PPO checkpoint, >90% simulation evaluation, or real deployment exists. Stages 12 require real-scene hardware inputs, stages 34 require a suitable GPU environment, and stage 5 requires authorized SO-100 actuation. See the [experiment record](../chapter6/rgb-sim2real-grasping/README.md). |
| 7-1 | External benchmark execution | **Complete for the manuscript's bounded five-task campaign.** The pinned τ²-bench telecom run scored 4/5 (Pass@1 0.80), retained all raw trajectories and costs, and traces the failed task to a phone/line mismatch that skipped the required data refuel. Upstream format and trial-count verification pass; full-domain task coverage is explicitly outside this bounded claim. See the [manifest](../chapter7/tau2-bench-eval/validation/runs/exp7-1-openrouter-gpt41mini-telecom-20260802-v1/manifest.json). |
| 7-2 | Human benchmark | **Complete.** The retained case set covers easy, medium, and hard tasks from GAIA, AndroidWorld, SWE-bench Verified, τ²-bench, Terminal-Bench, and OSWorld-Verified, with 18/18 first-run trajectories and official verification outcomes. See the [completed case set](../chapter7/experiment-7-2-human-benchmark/README.md). |
| 7-3 | Local implementation | **Complete.** The four-grade memory rubric has 60 cases and 180/180 structured judgments with full scope in the [saved evidence](../chapter7/user-memory-system-evaluation/results/full_7_3_structured_rubric_evidence.json). |
| 7-4 | Local experiment | **Complete.** JSON Cards, RAG, and hybrid systems produced 180/180 real trajectories across 60 cases, with complete cost and failure analysis in the [saved campaign](../chapter7/user-memory-system-evaluation/results/full_7_4_60_cases_costed.json). |
| 7-5 | Local experiment | **Complete.** The known-memory trajectory-prefix campaign ran 33/33 real OpenRouter cells (11 production bad cases × JSON/Markdown/Python-like encodings), with zero API errors and 6/11 deterministic policy passes for each encoding. See the [manifest](../chapter7/user-memory-policy-eval/results/manifest.json) and [report](../chapter7/user-memory-policy-eval/results/policy_prefix_live.json). |
| 7-6 | Local experiment | **Complete.** The neutral TTS campaign retained 8/8 content-hashed OpenAI/Fish cells and direct-audio Voxtral judgments across four challenge categories. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). |
| 7-7 | Local experiment | **Complete.** The Arena Elo/BradleyTerry campaign processed 1,799,991 public records, retained rankings, matrices, snapshots, plots, and an independently passing manifest. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). |
| 7-8 | Local experiment | **Complete.** The neutral Coding Harness campaign retained 18/18 cells (two models × three tasks × three trials), zero API errors, full trajectories, summaries, and verified artifact hashes in the [manifest](../chapter7/model-action-threshold/results/exp7-8-action-threshold-20260731-v1/manifest.json). |
| 7-9 | Local experiment | **Complete.** The eight-turn Agent cost campaign retains four real token/cache/latency arms and the measured KV-cache/context-compression comparison. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). |
| 7-10 | Long-running provider benchmark | **Incomplete.** The runner and analyzer exist, but retained evidence contains only 29 smoke/readiness observations: no standard N=100 cells, rate ramp, Agent-cost phase, or 168-hour availability campaign. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). |
| 7-11 | Local experiment | **Complete.** The full 4 × 3 × 2 × 60 matrix retained 1,440/1,440 real trajectories with zero errors or unpriced usage, complete retrieval/task metrics and factorial analysis, and an independently passing verifier. See the [canonical matrix](../chapter7/user-memory-system-evaluation/results/full_7_11_60_case_matrix.json). |
| 7-12 | Emulator evaluation | **Complete evidence; deployment not approved.** The [canonical evidence](../chapter7/android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json) retains all 580/580 unique episodes (116 tasks × five trials), including evaluator failures, with zero runtime errors: 26 strict T3A successes (4.4828%) and mean evaluator reward 0.133621, comprising 77 full-reward states plus one `0.5` partial reward. The completed official Pixel 6/API-33 setup had all 24/24 required apps and ran local Qwen2.5-7B revision `a09a35458c702b33eeacc393d103063234e8bc28` via vLLM 0.19.0 on an RTX PRO 6000 Blackwell 96 GB. Candidate Qwen differs from the paired-source Doubao model, so the result establishes neither same-model uplift nor noninferiority. |
| 7-13 | Simulation evaluation | **Complete; action chunking improves a low-success policy.** Pinned OpenVLA-OFT and RoboTwin2 ran two real single-GPU `val_only` arms of 256 episodes each with three RGB views and 14-D proprio/action evidence. Chunk 1 scored 0/256 and chunk 25 scored 26/256 (13/128 IID and 13/128 OOD), a paired +10.15625 pp result. All 486 failures have evidence-backed timeout classifications; the [manifest](../chapter7/openvla-robotwin2-eval/validation/runs/exp7-13-localgpu-20260803-v1/manifest.json) binds 512 rollout-video hashes and passes the strict retained-package verifier. |
| 8-6 | Local speech training experiment | **Complete (bounded campaign).** Orpheus and Sesame LoRAs were each trained for 60 optimizer steps on the local RTX PRO 6000, evaluated on held-out examples, and compared against their base arms with 40 retained WAVs. Full adapter identities, hashes, automatic proxy results, and failures are in the [strict report](../chapter8/speech-sft-experiment/validation/exp8-6-20260804-v1/REPORT.md). |
| 8-7 | Local multilingual training experiment | **Incomplete.** The SFT implementation exists, but the repository retains no checkpoint or before/after multilingual benchmark. See the [Chapter 8 ledger](../chapter8/EXPERIMENT_LEDGER.md). |
| 8-8 | Local training experiment | **Complete.** The retained campaign contains 160/160 training and 80/80 held-out Kimi K3 teacher receipts, a real CUDA-trained SmolLM2-135M-Instruct LoRA checkpoint, and the paired comparison in [`validation/exp8-8-kimi3-smollm2-20260730/`](../chapter8/prompt-distillation/validation/exp8-8-kimi3-smollm2-20260730/). Held-out results: teacher 100%, baseline 0%, trained 95%; ~197× latency speedup; ~75% input-token reduction. All eight evidence gates pass. |
| 8-9 | Local training experiment | **Complete; the distillation uplift hypothesis was not supported.** All 24 Kimi K3 teacher cases now retain completed trajectories: 23 passed the deterministic answer verifier and entered SFT, while `aime-2016-9-I` completed under native low-reasoning control with the wrong answer and was correctly rejected. Real CUDA training produced a Qwen2.5-1.5B-Instruct LoRA checkpoint in [`checkpoints/exp8-9-qwen25-1.5b-kimi-k3-20260801-v1/`](../chapter8/cot-distillation/checkpoints/exp8-9-qwen25-1.5b-kimi-k3-20260801-v1/). The [completed three-arm report](../chapter8/cot-distillation/validation/experiment_8_9_complete_20260803_v2.json) records baseline 1/24, student 2/24, teacher 23/24, and nonsignificant paired improvement (p=1.0). |
| 8-118-16 | External training reproductions | **Incomplete.** The GeneralPoints, V-IRL, SimpleVLA-RL, RLVP, ReTool, and AWorld sources/entrypoints are mapped to varying degrees, but none has a retained full training-and-evaluation campaign satisfying its manuscript gate. See the [Chapter 8 ledger](../chapter8/EXPERIMENT_LEDGER.md). |
| 9-8 | External-repository self-evolution experiment | **Complete for the autonomous, review-driven self-update loop; downstream benefit not evaluated.** Pinned Hermes received all ten English chapters and its own source without any supplied candidate gap. It independently chose to add evidence-backed learning signals to persisted trajectories. Three fresh terminal-review rejections were fed back to the original Hermes proposer session; it corrected production-format parsing, persistence-path coverage, and counting-consistency defects until a fourth fresh reviewer returned `VERDICT: ACCEPT`. The accepted patch passes 6 new plus 38 existing focused tests and clean-clone application, but remains unmerged; the proposed downstream ablation campaign was not run. See the [credential-free manifest](../chapter9/hermes-self-evolution/validation/exp9-8-hermes-gpt56luna-autonomous-20260802-v2/manifest.json). |
| 9-9 | Longitudinal continual-evolution evaluation | **Complete.** The static, append-only, and evolving arms ran 3 seeds × 14 ordered tasks (126 real model calls). The retained evidence separates transfer, rule replacement, retention, obsolete-rule citation, and paired statistics; only the evolving arm replaces the obsolete 20 kg rule and retains the current 23 kg rule. See the [canonical evidence](../chapter9/self-evolution-eval/validation/latest.json). |
| 10-1 | Local architecture comparison | **Complete (bounded comparison).** The repaired Skill arm enforces `load_skill("triage")` before specialist tools while keeping the fixed schema prefix. The canonical v2 campaign retains 30 paired tasks, 12 boundary trajectories, 289 provider receipts, 31 real Tavily receipts, and 60 position-swapped independent judge receipts. Skill passes 15/30 deterministic gates versus Transfer 2/30; Skill is slower and uses more uncached input in this model/configuration. See the [acceptance manifest](../chapter10/multi-role-transfer/validation/comparison/runs/exp10-1-qwen35flash-20260809-v2/manifest.json) and [report](../chapter10/multi-role-transfer/validation/comparison/runs/exp10-1-qwen35flash-20260809-v2/REPORT.md). |
| 10-2 | Local multi-agent comparison | **Complete.** The four-role Manager and single-Agent arms translated all 26 units of the retained illustrated/code-heavy technical-book sample, with 12/12 acceptance gates and complete quality, context, token, latency, and resource comparisons. See the [canonical index](../chapter10/book-translation/validation/latest.json). |
| 10-3 | External concurrent-agent reproduction | **Complete for the retained Anthropic-caller configuration.** The pinned TalkAct campaign retains 16/16 episodes with no runtime/provider errors and passes all 17 gates. Duplex and strawman tie at 1.0 task success; duplex improves median voice latency from 12.52 s to 2.32 s (5.40×), but strawman has higher probe correctness and lower mean wall time. The invalid Gemini credential required TalkAct's supported Anthropic Sonnet caller override, so this same-family configuration must not be silently pooled with upstream default-Gemini results. See the [acceptance report](../chapter10/talkact-reproduction/validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/acceptance.json). |
| 10-3 | Local WebRTC orchestration experiment | **Complete.** A real LLM autonomously selected the Phone Agent; Playwright, bidirectional RTP, local TTS/Whisper, validation/re-asking, concurrent ask/fill, and one localhost submission pass all 9 gates. PSTN/E.164 is not required by the manuscript. See the [manifest](../chapter10/autonomous-phone-registration/validation/runs/exp10-3-webrtc-raw-20260731-v4/manifest.json). |
| 10-5 | External generative-agents reproduction | **Complete; the custom-event diffusion hypothesis was not supported.** The exact pinned 25-persona society completed three 17,280-step, two-virtual-day arms with 148,856 canonical provider receipts and zero logical errors. The custom climate workshop remained limited to its originator, while disabling reflection produced zero evidence-linked reflection thoughts and reduced all four blind plausibility scores; baseline was preferred for 17/25 personas. All 14 gates pass in the [canonical acceptance report](../chapter10/generative-agents/validation/runs/exp10-5-qwen37flash-20260804-v1/acceptance.json). |
| 10-6 | Local voice multi-agent experiment | **Complete.** One retained eight-seat v11 game completed three night/day/vote cycles and passed every gate in the same report: six real LLM-tool → macOS `say` → OpenRouter native-audio ASR round trips with exact action agreement, information isolation, a rule-determined winner, and all four strategy criteria. The report retains 13 unique response IDs, 1,650 audio-input tokens, 27 positive-byte TTS events, action history, usage, audio hashes, and judge-attempt provenance; the independent validator rechecked all six audio/action boundaries. See the [canonical report](../chapter10/voice-werewolf/validation/runs/exp10-6-simulated-user-openrouter-20260803-v11/acceptance_report.json). |
## Detailed ledgers
The chapter ledgers are the canonical detailed records for acceptance scope,
saved evidence, and audit findings:
- [Chapter 1 experiment ledger](../chapter1/EXPERIMENT_LEDGER.md)
- [Chapter 2 experiment ledger](../chapter2/EXPERIMENT_LEDGER.md)
- [Chapter 3 experiment ledger](../chapter3/EXPERIMENT_LEDGER.md)
- [Chapter 4 experiment ledger](../chapter4/EXPERIMENT_LEDGER.md)
- [Chapter 5 experiment ledger](../chapter5/EXPERIMENT_LEDGER.md)
- [Chapter 7 experiment coverage ledger](../chapter7/EXPERIMENT_LEDGER.md)
- [Chapter 8 experiment coverage ledger](../chapter8/EXPERIMENT_LEDGER.md)
Update this summary whenever one of the tracked completion gates changes. Git
history provides the status change log.