ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s

This commit is contained in:
2026-08-20 13:12:50 +00:00
commit b119135836
10275 changed files with 3284984 additions and 0 deletions
+23
View File
@@ -0,0 +1,23 @@
# Chapter 7 experiment coverage ledger
This ledger separates runnable code, pinned external sources, and direct
acceptance evidence. A repository checkout, smoke test, or mechanism demo is
never counted as completion of a broader manuscript experiment.
| Experiment | Manuscript acceptance scope | Current evidence | Audit status |
| --- | --- | --- | --- |
| 7-1 | Run τ²-bench, inspect multi-turn failures, and compare the dual-control telecom design with historical τ-bench | [`exp7-1-openrouter-gpt41mini-telecom-20260802-v1`](tau2-bench-eval/validation/runs/exp7-1-openrouter-gpt41mini-telecom-20260802-v1/manifest.json) retains the raw five-task telecom trajectory from pinned upstream `8d005b0…`, 4/5 Pass@1, exact costs, and a wrong-line failure analysis. Upstream format and trial-count checks pass; full-task coverage fails as expected because the manuscript command deliberately samples five tasks rather than a complete leaderboard submission. Historical τ-bench remains pinned at `59a200c…` for the design comparison. | **Complete saved bounded campaign** |
| 7-2 | Personally complete simple/medium/hard tasks from GAIA, AndroidWorld, SWE-bench Verified, τ²-bench, Terminal-Bench, and OSWorld-Verified, retaining trajectories and official verification | [`experiment-7-2-human-benchmark/results.json`](experiment-7-2-human-benchmark/results.json) binds the preregistered 18/18 Codex-as-human cases to their per-benchmark trajectories and first official results: 13 passed, 5 failed, 0 unscored. The [report](experiment-7-2-human-benchmark/README.md) explains every task, operator trajectory, score, failure, and AndroidWorld/τ² compatibility boundary. | **Complete saved bounded campaign** |
| 7-3 | Four-grade precision/recall/reasoning/proactivity rubric, examples/boundaries, hallucination veto | [`user-memory-system-evaluation/results/full_7_3_structured_rubric_evidence.json`](user-memory-system-evaluation/results/full_7_3_structured_rubric_evidence.json): 60 cases, 180/180 structured judgments, full scope, complete. | **Complete** |
| 7-4 | Run Advanced JSON Cards, RAG, and hybrid over the same 60 cases; compare quality, steps, tools, latency, cost, and failure boundaries | [`user-memory-system-evaluation/results/full_7_4_60_cases_costed.json`](user-memory-system-evaluation/results/full_7_4_60_cases_costed.json): 180/180 real trajectories, zero errors, complete pricing coverage and failure analysis. | **Complete** |
| 7-5 | Supply known user memories and trajectory prefixes; evaluate scoped use, current-instruction override, safe clarification, and forbidden next actions across JSON/Markdown/Python-like encodings | [`user-memory-policy-eval/results/manifest.json`](user-memory-policy-eval/results/manifest.json) binds 33/33 real OpenRouter cells (11 bad cases × 3 encodings), zero API errors, and the content-hashed report [`policy_prefix_live.json`](user-memory-policy-eval/results/policy_prefix_live.json): 6/11 passed for each encoding. | **Complete saved campaign** |
| 7-6 | Multiple TTS providers/configurations × diverse corpus; direct-audio judge scores accuracy, naturalness, emotion, and voice consistency against reference audio | [`mistral_multimodal_20260730`](tts-quality-eval/validation/mistral_multimodal_20260730/manifest.json) retains 8/8 content-hashed OpenAI/Fish MP3 cells over four challenge categories, the fixed reference hash, exact four-dimension Voxtral judgments, and a recomputed complete gate. Earlier Google/OpenRouter/account failures remain as historical negative evidence. | **Complete saved campaign** |
| 7-7 | Online Elo from real Arena votes, Bradley-Terry comparison, win matrix, official-ranking comparison, and historical animation | [`exp7-7-arena-20260731-v1`](elo-leaderboard/validation/runs/exp7-7-arena-20260731-v1/manifest.json) binds the 2.0 GB public snapshot by SHA-256 and processes all 1,799,991 source rows (1,670,250 accepted blind votes, 129 models). Chronological K=4 Elo and deterministic-bootstrap Bradley-Terry rankings have Spearman 0.787 / Kendall 0.606 agreement and 12/20 top-model overlap; empirical/predicted matrices, 17 monthly snapshots, three plots, and the D3 animation are content-hashed and independently revalidated. | **Complete saved campaign** |
| 7-8 | Hold a neutral coding harness fixed while swapping GPT/Claude models; repeat localized, cross-cutting, and contract-sensitive tasks; measure pre-edit exploration, first-patch acceptance, rework, final tests, latency, files, and tokens | [`model-action-threshold/results/exp7-8-action-threshold-20260731-v1/manifest.json`](model-action-threshold/results/exp7-8-action-threshold-20260731-v1/manifest.json) binds 18/18 real OpenRouter cells (2 models × 3 tasks × 3 trials), zero API errors, full trajectories, and independently passing artifact hashes. Both models pass all final tests; GPT-5.6-sol averages 6.89 pre-edit tool calls / 4.67 files versus Claude Sonnet 5 at 4.56 / 3.56. | **Complete saved campaign** |
| 7-9 | End-to-end multi-turn cost decomposition and measured KV-cache/context-compression A/B | `agent-cost-analysis/sample_trace.json` retains the real four-arm, eight-turn token/cache/latency observations; README reports per-step, percentile, component, and 2×2 results. | **Complete saved campaign** |
| 7-10 | Multi-provider 8K/32K/128K × 512/2048 workload at N≥100, p50/p95/p99, thinking, pricing, same-model provider pair, rate ramp, Agent trace, and 168-hour hourly availability | Campaign and strict analyzer implemented. [`model-benchmark/results/manifest.json`](model-benchmark/results/manifest.json) reports only 29 smoke/readiness observations, no standard N=100 cells, no rate ramp/Agent-cost phase, and no 168-hour campaign. | **Incomplete—long-running/costly campaign** |
| 7-11 | Full 4 embeddings × 3 rerankers × 2 main models × 60 cases with retrieval/task metrics and interaction analysis | [`user-memory-system-evaluation/results/full_7_11_60_case_matrix.json`](user-memory-system-evaluation/results/full_7_11_60_case_matrix.json): 60 cases × 24 cells = 1,440/1,440 real trajectories, zero error records, zero unpriced usage, complete retrieval/task metrics and factorial interaction analysis, top-level and completion status `complete`; independently rechecked by `user-memory-system-evaluation/validation/verify_full_matrix_20260731.py` (ALL CHECKS PASSED). Executed under documented backend substitutions (BGE-M3 and OpenAI embeddings via OpenRouter as identical models, Qwen3-Embedding-8B substituting the endpoint-gated Doubao embedding, Doubao-LLM reranker substituting the unreachable BGE cross-encoder); see [`user-memory-system-evaluation/results/full_matrix_backend_readiness_20260731.json`](user-memory-system-evaluation/results/full_matrix_backend_readiness_20260731.json) and `candidate_backend_probes_20260731.json`. | **Complete saved campaign** |
| 7-12 | Diagnose AndroidWorld, test layered hypotheses, make cost-benefit decision, rerun full suite, and iterate | [`android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json`](android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json) retains all 580/580 unique episodes (116 tasks × five trials), including evaluator failures, with zero runtime errors. Strict T3A success is 26/580 (4.4828%); mean evaluator reward is 0.133621, comprising 77 full-reward states plus one `0.5` partial reward. The five-shard run used the completed official setup and all 24/24 required apps on Pixel 6/API-33, with local Qwen2.5-7B revision `a09a35458c702b33eeacc393d103063234e8bc28` via vLLM 0.19.0 on an RTX PRO 6000 Blackwell 96 GB. Execution and evidence are complete, but deployment is not approved; because the candidate Qwen model differs from the paired-source Doubao model, this evidence supports neither a same-model uplift nor a noninferiority claim. | **Complete saved campaign—deployment not approved** |
| 7-13 | Real OpenVLA + RoboTwin2 `move_can_pot` evaluation with three RGB views, 14-D proprio/action, IID/OOD seeds, timing, failures, and action-chunk ablation | [`exp7-12-localgpu-20260803-v1`](openvla-robotwin2-eval/validation/runs/exp7-12-localgpu-20260803-v1/manifest.json) binds two real single-GPU `val_only` arms of 128 IID + 128 OOD episodes each, 512 rollout-video hashes, all process/config/checkpoint/source identities, and 486 evidence-backed timeout classifications. Chunk 1 scored 0/256; chunk 25 scored 26/256 (13/128 in both IID and OOD), a paired +10.15625 pp result. The strict analyzer and retained-package verifier both pass. | **Complete saved campaign; low absolute success retained** |
External source identities and commands are maintained in [README.md](README.md).
+48
View File
@@ -0,0 +1,48 @@
# الفصل السابع · تقييم الوكلاء
> يحول أداء الوكيل إلى إشارات قابلة للقياس والمقارنة. ويغطي بيئات التقييم، وتصميم مجموعات البيانات، والمقاييس، والدلالة الإحصائية، وقابلية الرصد، واختيار النموذج وفق النتائج، والتقييم الداخلي وبيئات المحاكاة في أنظمة الإنتاج.
← [العودة إلى الملف التمهيدي الرئيسي](../docs/ar/README.md) · 📖 [قراءة نص الفصل](../book-ar/chapter7.ar.md)
## كيفية قراءة التجارب
يستخدم النص هياكل آلية قصيرة لشرح تدفق التحكم؛ ويحتوي دليل التجارب على محولات SDK الكاملة والسجلات والاختبارات وأدلة القبول. لا حاجة لقراءة كل ملف سطرًا سطرًا.
- **Starter:** ابدأ بالهدف والأمر الأدنى وشروط القبول؛ وابدأ من [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** تتبّع نقطة الدخول والحلقة الأساسية ومخطط الحالة/الرسائل والأدوات وأداة التحقق.
- **Maintainer:** ثم اقرأ الاختبارات وmanifest الأدلة ومعالجة الأعطال ومسارات التراجع ومحولات المزوّد.
في القراءة الأولى يمكنك تجاوز بيانات الاعتماد وطبقة العرض وتوافق المزوّد؛ عُد إليها عند إعادة إنتاج رقم.
## المشاريع المصاحبة
| التجربة | المشروع | النوع | الوصف |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | يركز على تقييم قدرة الوكيل على استخدام الأدوات للاستدلال المعقد، بما في ذلك سيناريوهات مثل الحساب والبحث ومعالجة البيانات. |
| 6-2 | `tau2-bench/` | 📖 | إكمال مهام τ²-bench المتدرجة يدويًا وتسجيل المسارات. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | تشغيل قاعدة التقييم ذات المستويات الأربعة على 180 حكمًا منظمًا مع الأدلة ومنع الهلوسة. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | تشغيل 60 حالة على ثلاثة أنظمة مع محاسبة كاملة للتكلفة. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | تشغّل 11 حالة سيئة لبادئة المسار عبر تمثيلات ذاكرة JSON وMarkdown وشبيهة بـ Python، باستخدام استدعاءات OpenRouter حقيقية وفحوص سياسات حتمية. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | احتفظت المصفوفة الكاملة 4×3×2×60 بعدد 1,440/1,440 مسارًا حقيقيًا بلا أخطاء أو استخدام غير مسعّر، مع اكتمال مقاييس الاسترجاع والمهام وتحليل التفاعل واجتياز أداة التحقق المستقلة. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | أكملت الحملة الرسمية على GPU واحدة 256 حلقة لكل ذراع؛ حقق chunk 1 نتيجة 0/256 وchunk 25 نتيجة 26/256، مع حفظ تجزئات 512 rollout. |
| 6-2 | `terminal-bench/` | 📖 | يعد Terminal-Bench معيارًا لاختبار أداء AI Agent في البيئات الطرفية الحقيقية. بدءًا من تجميع الشفرة وحتى نماذج التدريب وإعداد الخوادم، يقوم بتقييم كيفية تعامل الوكلاء مع المهام الحقيقية الشاملة. يتضمن مجموعة بيانات مكونة من 100 مهمة تقريبًا وإطار عمل للتنفيذ، يدعم عمليات تنفيذ الوكيل المختلفة. |
| 6-2 | `SWE-bench/` | 📖 | يعد SWE-bench معيارًا لتقييم قدرة نماذج اللغات الكبيرة على حل مشكلات GitHub الحقيقية. بالنظر إلى قاعدة الشفرة ووصف المشكلة، يجب أن يقوم النموذج بإنشاء تصحيح يعمل على حل المشكلة. تتضمن إصدارات متعددة: SWE-bench، وSWE-bench Lite، وSWE-bench Verified، وSWE-bench Multimodal. |
| 6-2 | `GAIA/` | 📖 | تهدف GAIA إلى تقييم الجيل التالي من نماذج LLM (أولئك الذين لديهم أدوات تكبير، وتحفيز فعال، والوصول إلى البحث، وما إلى ذلك). فهو يحتوي على أكثر من 450 سؤالًا غير تافه يتطلب درجات متفاوتة من استخدام الأداة والاستقلالية، مع إجابات لا لبس فيها. مقسمة إلى 3 مستويات صعوبة. |
| 6-2 | `OSWorld/` | 📖 | يقيم قدرة الوكلاء على أداء المهام المعقدة ضمن بيئة نظام تشغيل كاملة، بما في ذلك إدارة الملفات وتشغيل التطبيق وتكوين النظام. |
| 6-2، 6-12 | `android_world/` | 📖 | يقوم بتقييم أداء الوكيل في بيئة الهاتف المحمول التي تعمل بنظام Android، بما في ذلك التنقل في التطبيق وتفاعل واجهة المستخدم وإمكانيات إكمال المهام (الريبو المعياري الخارجي). |
| 6-6 | [تقييم جودة تحويل النص إلى كلام](tts-quality-eval/) | ✅ | تجميع نفس المجموعة من النصوص الصعبة باستخدام تكوينات تحويل النص إلى كلام (TTS) المختلفة (نموذج/صوت/سرعة مختلفة)، ثم استخدام LLM-as-a-Judge متعدد الوسائط لتسجيل كل بُعد (الوضوح، والطبيعة، وما إلى ذلك) وفقًا لقاعدة تقييم، وتجميع النتائج في جدول مقارنة التكوين القابل للتكرار. |
| 6-7 | [elo-المتصدرين](elo-leaderboard/) | ✅ | تنفيذ لوحة صدارة لأداء الوكيل استنادًا إلى نظام تصنيف ELO، وتقييم القدرات النسبية للوكلاء المختلفين من خلال المقارنات الزوجية. |
| 6-8 | [عتبة انتقال النموذج إلى التنفيذ](model-action-threshold/) | ✅ | يقارن عتبة الانتقال من الاستكشاف إلى أول تعديل بين GPT-5.6-sol وClaude Sonnet 5 تحت Coding Harness محايد وثابت؛ أكملت الحملة 18/18 حالة بلا أخطاء API، وربط [البيان](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) المسارات والملخصات بهاشات قابلة للتحقق. |
| 6-9 | [تحليل تكلفة الوكيل](agent-cost-analysis/) | ✅ | ينفذ توزيعًا لتكلفة السلسلة الكاملة لمهمة وكيل نموذجية متعددة المنعطفات (استرداد خدمة العملاء): يستخدم نظام تتبع مخصص خفيف الوزن لتسجيل الرموز للإدخال/الإخراج/ذاكرة التخزين المؤقت، وزمن الوصول، والتكلفة لكل مكالمة LLM، والتجميع لتحديد "الخطوة الأكثر تكلفة"، ثم يستخدم اختبار A/B لتحديد المدخرات الحقيقية من التصميم الصديق لذاكرة التخزين المؤقت KV وضغط السياق. |
| 6-10 | [نموذج المعيار](model-benchmark/) | 🚧 | إجراء اختبار أفقي لموفري OpenAI المتعددين المتوافقين مع LLM API. يستخدم واجهة تدفق لقياس الوقت حتى أول رمز مميز (TTFT) بدقة، ويحسب النسب المئوية لزمن الوصول من طرف إلى طرف (p50/p95)، والإنتاجية، ومعدل النجاح في ظل التزامن. ينتج أمر واحد جدول مقارنة متعدد الأبعاد، يوضح أن اختيار النموذج هو عبارة عن مقايضة متعددة الأوجه بدلاً من مجرد النظر إلى لوحة المتصدرين. |
| 6-12 | [عالم الروبوت](android-world/) | 📖 | تقرير تقييم In-repo T3A وملاحظات تحليل الفشل على AndroidWorld (نقطة البداية للتجربة 6-12؛ وليس مصدر المعيار). |
| — | [تقييم تقارير الصحة العامة](public-health-reporting-eval/) | ✅ | يستخدم بيانات مجمعة على نمط DHIS2 لإجراء تقييم موضوعي لاستدعاءات أداة وكيل الإبلاغ عن الصحة العامة، ودقة الحساب، واستشهادات الأدلة، والمطالبات غير المدعومة. |
> يجب أن يتم استنساخ المعايير الخارجية المسماة Backtick بشكل منفصل. [`android-world/`](android-world/) (موصول) هو **ملاحظات تحليل تقييم T3A** الخاصة بهذا الريبو (راجع [README](android-world/README.md))، وليس نفس المسار مثل مصدر قياس الأداء الخارجي `android_world/`.
## أنواع المشاريع
| الأيقونة | النوع | المعنى |
| :--: | --- | --- |
| ✅ | **مستقل** | شفرة كاملة قابلة للتشغيل في هذا المستودع بعد إعداد مفتاح API |
| 📖 | **دليل إعادة الإنتاج** | وثائق تفصيلية تعتمد على مستودع خارجي يُجلب باستخدام `git clone` |
| 🚧 | **وثيقة التصميم** | وثيقة تصميم وخطة تنفيذ؛ أما الشفرة القابلة للتشغيل فما تزال قيد التطوير |
+48
View File
@@ -0,0 +1,48 @@
# Chapter 7 · Agent Evaluation
> Turns Agent performance into comparable signals. Covers evaluation environments, dataset design, metric systems, statistical significance, observability, evaluation-driven selection, and production-grade internal evaluation and simulation environments.
← [Back to main README](../docs/en/README.md) · 📖 [Read chapter text](../book-en/chapter7.md)
## How to Read the Experiments
The prose uses short mechanism skeletons to explain control flow; the experiment directory contains complete SDK adapters, logs, tests, and acceptance evidence. You do not need to read every file line by line.
- **Starter:** Start with the goal, minimum command, and acceptance conditions; begin with [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Follow the entry point, core loop, state/message schema, tools, and verifier.
- **Maintainer:** Then read tests, evidence manifests, failure handling, rollback paths, and provider adapters.
On a first pass, skip credential loading, presentation code, and provider-compatibility layers; return when reproducing a number.
## Companion Projects
| Exp. | Project | Type | Description |
| :--: | --- | :--: | --- |
| 6-1 | [tau2-bench-eval](tau2-bench-eval/) | ✅ | Retains a pinned five-task telecom campaign (4/5 passed), raw trajectories, costs, hashes, and analysis of the wrong-line failure that skipped data refueling. |
| 6-2 | `tau2-bench/` | 📖 | Manually completes graded τ²-bench tasks and records their trajectories. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | Runs the four-level rubric over 180 structured judgments with evidence and a hallucination veto. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Runs 60 cases across three systems with complete cost accounting. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Runs 11 trajectory-prefix bad cases across JSON, Markdown, and Python-like memory encodings with real OpenRouter calls and deterministic policy checks. |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | Synthesizes the same set of challenging texts using various TTS configurations (different model/voice/speed), then uses a multimodal LLM-as-a-Judge to score each dimension (clarity, naturalness, etc.) according to a Rubric, aggregating the results into a reproducible configuration comparison table. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Implements an agent performance leaderboard based on the ELO rating system, evaluating the relative abilities of different agents through pairwise comparisons. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | Compares GPT-5.6-sol and Claude Sonnet 5 at the transition from exploration to the first edit under the same neutral Coding Harness; all 18/18 cells completed without API errors, and the [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) binds the trajectories and summaries with verifiable hashes. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Performs a full-chain cost breakdown for a typical multi-turn agent task (customer service refund): uses a custom lightweight tracing system to record input/output/cache tokens, latency, and cost for each LLM call, aggregates to identify "which step is the most expensive," and then uses A/B testing to quantify the real savings from KV-cache-friendly design and context compression. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | Implements the multi-provider benchmark and strict analyzer, but retained evidence contains only smoke/readiness observations; the standard N=100 cells, rate ramp, Agent-cost phase, and 168-hour availability campaign remain incomplete. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | The full 4×3×2×60 matrix retained 1,440/1,440 real trajectories with zero errors or unpriced usage, complete retrieval/task metrics and interaction analysis, and an independently passing verifier. |
| 6-12 | [android-world](android-world/) | 📖 | In-repo T3A evaluation report and failure analysis notes on AndroidWorld (starting point for Experiment 6-12; not the benchmark source). |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | The retained single-GPU campaign completed 256 episodes per action-chunk arm; chunk 1 scored 0/256 and chunk 25 scored 26/256, with all 512 rollout identities hashed. |
| 6-2 | `terminal-bench/` | 📖 | Terminal-Bench is a benchmark for testing AI Agent performance in real terminal environments. From compiling code to training models and setting up servers, it evaluates how Agents handle real end-to-end tasks. Includes a dataset of ~100 tasks and an execution framework, supporting various Agent implementations. |
| 6-2 | `SWE-bench/` | 📖 | SWE-bench is a benchmark for evaluating the ability of large language models to solve real GitHub issues. Given a codebase and an issue description, the model must generate a patch that resolves the problem. Includes multiple versions: SWE-bench, SWE-bench Lite, SWE-bench Verified, and SWE-bench Multimodal. |
| 6-2 | `GAIA/` | 📖 | GAIA aims to evaluate next-generation LLMs (those with tool augmentation, efficient prompting, search access, etc.). It contains 450+ non-trivial questions requiring varying degrees of tool use and autonomy, with unambiguous answers. Divided into 3 difficulty levels. |
| 6-2 | `OSWorld/` | 📖 | Evaluates the ability of agents to perform complex tasks within a complete operating system environment, including file management, application operation, and system configuration. |
| 6-2, 6-12 | `android_world/` | 📖 | Evaluates agent performance in an Android mobile environment, including app navigation, UI interaction, and task completion capabilities (external benchmark repo). |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Uses synthetic DHIS2-style aggregate data to objectively evaluate a public-health reporting agent's tool calls, calculation accuracy, evidence citations, and unsupported claims. |
> Backtick-named external benchmarks must be cloned separately. [`android-world/`](android-world/) (hyphenated) is this repo's **T3A evaluation analysis notes** (see its [README](android-world/README.md)), not the same path as the external `android_world/` benchmark source.
## Project Types
| Icon | Type | Meaning |
| :--: | --- | --- |
| ✅ | **Standalone** | Full code in this repo, runs after configuring API Key |
| 📖 | **Reproduction Guide** | Detailed doc depending on **external repos** to `git clone` |
| 🚧 | **Design Doc** | Architecture/implementation plan only, runnable code still WIP |
+51
View File
@@ -0,0 +1,51 @@
# Capítulo 7 · Evaluación de Agentes
> Convertir el rendimiento en señales comparables: entornos, métricas, significación estadística, selección guiada por evaluación
← [Volver al README principal](../docs/es/README.md) · 📖 [Leer texto del capítulo](../book-es/chapter7.es.md)
Los requisitos, la evidencia directa y los límites de cada experimento se detallan en el [registro de aceptación](EXPERIMENT_LEDGER.md).
## Cómo leer los experimentos
El texto usa skeletons breves para explicar el flujo de control; el directorio de experimentos contiene adaptadores SDK completos, registros, pruebas y evidencias de aceptación. No hace falta leer cada archivo línea por línea.
- **Starter:** Empieza por el objetivo, el comando mínimo y la aceptación; comienza con [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Sigue el punto de entrada, el bucle central, el esquema de estado/mensajes, las herramientas y el verificador.
- **Maintainer:** Después revisa pruebas, manifiestos, fallos, rollback y adaptadores de proveedores.
En la primera pasada puedes omitir credenciales, presentación y compatibilidad de proveedores; vuelve al reproducir una cifra.
## Proyectos Complementarios
| Exp. | Proyecto | Tipo | Descripción |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | Ejecuta la evaluación multirronda con doble control de τ²-bench y la compara con las definiciones de tareas, condiciones de éxito y simulador de usuario de τ-bench |
| 6-2 | `tau2-bench/` | 📖 | Completa manualmente tareas graduadas de τ²-bench y registra sus trayectorias; es solo una de las seis clases de benchmarks que se muestrean en 6-2 |
| 6-2 | `terminal-bench/` | 📖 | Evalúa la capacidad integral del Agent en un entorno de terminal real (compilación, entrenamiento y despliegue), con unas 100 tareas y un marco de ejecución |
| 6-2 | `SWE-bench/` | 📖 | Evalúa la capacidad de los LLM para resolver incidencias reales de GitHub en las variantes SWE-bench, Lite, Verified y Multimodal |
| 6-2 | `GAIA/` | 📖 | Evalúa herramientas, búsqueda y autonomía mediante más de 450 preguntas no triviales con respuestas inequívocas y tres niveles de dificultad |
| 6-2 | `OSWorld/` | 📖 | Evalúa tareas complejas en un sistema operativo completo: gestión de archivos, uso de aplicaciones y configuración del sistema |
| 6-2, 6-12 | `android_world/` | 📖 | Evalúa navegación de aplicaciones, interacción con la IU y finalización de tareas en Android (repositorio de benchmark externo) |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | La rúbrica multidimensional de cuatro niveles se ejecutó sobre 180/180 evaluaciones reales (60 casos × 3 sistemas); el [índice independiente](user-memory-system-evaluation/results/full_6_3_structured_rubric_evidence.json) conserva razones, evidencia, casos límite y el veto por alucinación con estado `complete` |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 180/180 trayectorias reales (60 casos × 3 sistemas), sin errores y con precios completos en la moneda nativa; el [resultado de aceptación](user-memory-system-evaluation/results/full_6_4_60_cases_costed.json) tiene estado `complete` |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Ejecuta 11 casos problemáticos de prefijos de trayectoria en representaciones de memoria JSON, Markdown y similares a Python, con llamadas reales a OpenRouter y comprobaciones deterministas de políticas. |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | La [aceptación real](tts-quality-eval/validation/mistral_multimodal_20260730/manifest.json) completa 8/8 evaluaciones Voxtral de cuatro dimensiones sobre dos proveedores y cuatro clases de muestras; cada audio candidato y de referencia tiene hash |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Tabla de clasificación del rendimiento de Agentes basada en ELO y comparaciones directas |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | Compara GPT-5.6-sol y Claude Sonnet 5 en la transición de la exploración a la primera edición bajo el mismo Coding Harness neutral; se completaron 18/18 celdas sin errores de API y el [manifiesto](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) vincula trayectorias y resúmenes mediante hashes verificables |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Desglose integral de costos para una tarea multirronda de reembolso, con diseño compatible con caché KV y cuantificación A/B del ahorro por compresión de contexto |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | Están implementadas las campañas 8K/32K/128K × 512/2048, rampas por límites, costos del Agent y disponibilidad durante 168 horas; el [manifiesto](model-benchmark/results/manifest.json) actual solo contiene pruebas reales de humo y disponibilidad |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | La matriz completa 4×3×2×60 conserva 1.440/1.440 trayectorias reales sin errores ni uso sin precio, con métricas completas de recuperación y tareas, análisis de interacción y un verificador independiente aprobado. |
| 6-12 | [android-world](android-world/) | 📖 | Informe y notas de análisis de fallos de la evaluación de T3A Agent en AndroidWorld (punto de partida de 6-12, no código fuente del benchmark) |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | La campaña oficial con una GPU completó 256 episodios por brazo: chunk 1 obtuvo 0/256 y chunk 25 obtuvo 26/256, con hashes de los 512 rollouts. |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Evalúa objetivamente las llamadas a herramientas, la exactitud de los cálculos, las citas de evidencia y las afirmaciones sin fundamento sobre datos agregados sintéticos al estilo DHIS2 |
> Los benchmarks externos entre comillas invertidas deben clonarse por separado. [`android-world/`](android-world/) (con guion) contiene las notas internas sobre la evaluación de T3A; no es la misma ruta que el código externo `android_world/`.
## Tipos de Proyectos
| Icono | Tipo | Significado |
| :--: | --- | --- |
| ✅ | **Autónomo** | Código completo en este repositorio, se ejecuta tras configurar la Clave API |
| 📖 | **Guía de Reproducción** | Documento detallado que depende de **repositorios externos** para realizar `git clone` |
| 🚧 | **En curso** | Existe una implementación, pero el alcance del experimento o su evidencia de aceptación aún no satisface todos los requisitos del texto |
+49
View File
@@ -0,0 +1,49 @@
# 7. fejezet · Ügynökök kiértékelése
> A teljesítményt összehasonlítható jellé alakítja értékelési környezetekkel, adathalmazokkal, mérőszámokkal, megfigyelhetőséggel és értékelésvezérelt kiválasztással.
← [Vissza a magyar főoldalhoz](../docs/hu/README.md) · 📖 [A fejezet olvasása](../book-hu/chapter7.md)
## Hogyan olvassuk a kísérleteket?
A törzsszöveg rövid mechanizmus-skeletonokkal magyarázza a vezérlési folyamatot; a kísérleti könyvtárakban találhatók a teljes SDK-adapterek, naplók, tesztek és átvételi bizonyítékok. Nem kell minden fájlt sorról sorra elolvasni.
- **Starter:** Kezdje a céllal, a minimális paranccsal és az átvételi feltételekkel; induljon innen: [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Kövesse a belépési pontot, a fő ciklust, az állapot-/üzenetsémát, az eszközöket és az ellenőrzőt.
- **Maintainer:** Végül olvassa el a teszteket, a bizonyíték-manifeszteket, a hibakezelést, a visszaállítási útvonalakat és a provider-adaptereket.
Első olvasáskor átugorható a hitelesítő adatok betöltése, a megjelenítési réteg és a provider-kompatibilitás; a számok reprodukálásakor térjen vissza.
## Kapcsolódó projektek
| Kísérlet | Projekt | Típus | Leírás |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | Többkörös, kettős vezérlésű τ²-bench értékelést futtat, és összeveti a τ-bench-csel. |
| 6-2 | `tau2-bench/` | 📖 | Mintafeladatokat old meg kézzel a τ²-bench-ből, és rögzíti a végrehajtási nyomvonalakat. |
| 6-2 | `terminal-bench/` | 📖 | Valós terminálkörnyezetben tesztel teljes, végponttól végpontig tartó feladatokat. |
| 6-2 | `SWE-bench/` | 📖 | Valós GitHub Issue-k tesztelhető javítással történő megoldását értékeli. |
| 6-2 | `GAIA/` | 📖 | Többszintű feladatokon méri a keresést, eszközhasználatot és autonómiát. |
| 6-2 | `OSWorld/` | 📖 | Teljes operációsrendszer-környezetben értékeli a fájl-, alkalmazás- és konfigurációkezelést. |
| 6-2, 6-12 | `android_world/` | 📖 | Androidon méri az alkalmazásnavigációt és a felhasználói felület kezelését. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | Többdimenziós memóriaértékelési rubrikát futtat, minden ítélethez bizonyítékkal. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Azonos esetkészleten hasonlítja össze a JSON Cards, RAG és hibrid rendszereket. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Tizenegy hibás trajektória-előtag esetet futtat JSON, Markdown és Python-szerű memóriareprezentációkon, valós OpenRouter-hívásokkal és determinisztikus szabályzat-ellenőrzésekkel. |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | Rubrikaalapú multimodális LLM-bíróval hasonlít össze TTS-konfigurációkat. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Páronkénti összehasonlítások és ELO-pontszám alapján készít ágensranglistát. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | Azonos, semleges Coding Harness alatt hasonlítja össze a GPT-5.6-sol és a Claude Sonnet 5 átmenetét a feltárástól az első szerkesztésig; mind a 18/18 cella API-hiba nélkül lefutott, a [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) pedig ellenőrizhető hash-ekkel köti össze a nyomvonalakat és az összesítéseket. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Felbontja a teljes költséget, és méri a cache-barát tervezés és tömörítés megtakarítását. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | TTFT-t, késleltetést, áteresztőképességet, megbízhatóságot és költséget mér; a hosszú kampány még nem teljes. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | A teljes 4×3×2×60 mátrix 1 440/1 440 valós trajektóriát őriz meg hiba és árazatlan használat nélkül, teljes visszakeresési és feladatmetrikákkal, interakcióelemzéssel és sikeres független ellenőrzéssel. |
| 6-12 | [android-world](android-world/) | 📖 | Repository-n belüli T3A-értékelési jelentés és AndroidWorld-hibaelemzés. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | Az egy GPU-s hivatalos futás karonként 256 epizódot teljesített: chunk 1 0/256, chunk 25 26/256; mind az 512 rollout hash-e megmaradt. |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Közegészségügyi jelentések eszközhívásait, számításait, hivatkozásait és állításait értékeli. |
> A kódformázással jelölt benchmarkokat külön kell klónozni. Az `android-world/` helyi elemzési jegyzet, nem az `android_world/` benchmark forrása.
## Projekttípusok
| Ikon | Típus | Jelentés |
| :--: | --- | --- |
| ✅ | **Önálló** | A teljes kód a repository-ban található, és az API-kulcsok beállítása után futtatható. |
| 📖 | **Reprodukciós útmutató** | Külső repository szükséges, amelyet külön kell `git clone` paranccsal letölteni. |
| 🚧 | **Folyamatban** | Az implementáció vagy az elfogadási bizonyíték még nem teljes. |
+49
View File
@@ -0,0 +1,49 @@
# Bab 7 · Evaluasi Agent
> Mengubah performa menjadi sinyal yang dapat dibandingkan melalui lingkungan evaluasi, dataset, metrik, observabilitas, dan pemilihan berbasis evaluasi.
← [Kembali ke README utama](../docs/id/README.md) · 📖 [Baca bab](../book-id/chapter7.md)
## Cara Membaca Eksperimen
Teks utama memakai skeleton mekanisme singkat untuk menjelaskan alur kontrol; direktori eksperimen berisi adapter SDK lengkap, log, pengujian, dan bukti penerimaan. Anda tidak perlu membaca setiap berkas baris demi baris.
- **Starter:** Mulai dari tujuan, perintah minimum, dan syarat penerimaan; awali dengan [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Telusuri titik masuk, loop inti, skema status/pesan, alat, dan verifier.
- **Maintainer:** Terakhir, baca pengujian, manifest bukti, penanganan kegagalan, rollback, dan adapter provider.
Pada pembacaan pertama, lewati kredensial, presentasi, dan kompatibilitas provider; kembali saat mereproduksi angka.
## Proyek Pendamping
| Eksperimen | Proyek | Jenis | Deskripsi |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | Menjalankan evaluasi multi-putaran dual-control τ²-bench dan membandingkannya dengan τ-bench. |
| 6-2 | `tau2-bench/` | 📖 | Menyelesaikan sampel tugas τ²-bench secara manual dan mencatat trajectory. |
| 6-2 | `terminal-bench/` | 📖 | Menguji tugas end-to-end pada lingkungan terminal nyata. |
| 6-2 | `SWE-bench/` | 📖 | Mengevaluasi penyelesaian Issue GitHub nyata dengan patch yang dapat diuji. |
| 6-2 | `GAIA/` | 📖 | Mengevaluasi pencarian, penggunaan tool, dan otonomi pada soal bertingkat. |
| 6-2 | `OSWorld/` | 📖 | Mengevaluasi operasi file, aplikasi, dan konfigurasi pada lingkungan OS lengkap. |
| 6-2, 6-12 | `android_world/` | 📖 | Mengevaluasi navigasi aplikasi dan interaksi UI pada Android. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | Menjalankan Rubric memori multi-dimensi dengan bukti untuk setiap penilaian. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Membandingkan JSON Cards, RAG, dan sistem hibrida pada kumpulan kasus yang sama. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Menjalankan 11 kasus buruk awalan trajectory pada representasi memori JSON, Markdown, dan bergaya Python dengan panggilan OpenRouter nyata serta pemeriksaan kebijakan deterministik. |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | Membandingkan konfigurasi TTS menggunakan LLM multimodal sebagai juri berbasis Rubric. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Membuat papan peringkat Agent berdasarkan perbandingan berpasangan dan rating ELO. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | Membandingkan GPT-5.6-sol dan Claude Sonnet 5 saat beralih dari eksplorasi ke edit pertama di bawah Coding Harness netral yang sama; seluruh 18/18 sel selesai tanpa error API, dan [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) mengikat trajectory serta ringkasan dengan hash yang dapat diverifikasi. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Mengurai biaya end-to-end dan mengukur penghematan desain ramah cache serta kompresi. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | Mengukur TTFT, latensi, throughput, reliabilitas, dan biaya model; kampanye panjang belum selesai. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Matriks penuh 4×3×2×60 menyimpan 1.440/1.440 trajectory nyata tanpa error atau penggunaan tanpa harga, lengkap dengan metrik retrieval dan tugas, analisis interaksi, serta verifikator independen yang lulus. |
| 6-12 | [android-world](android-world/) | 📖 | Laporan evaluasi T3A dan analisis kegagalan AndroidWorld di dalam repositori. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | Kampanye resmi satu GPU menyelesaikan 256 episode per lengan; chunk 1 mendapat 0/256 dan chunk 25 mendapat 26/256, dengan hash untuk seluruh 512 rollout. |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Mengevaluasi panggilan tool, kalkulasi, sitasi, dan klaim laporan kesehatan publik. |
> Benchmark dengan nama berformat kode harus dikloning secara terpisah. `android-world/` adalah catatan analisis lokal, bukan sumber benchmark `android_world/`.
## Jenis Proyek
| Ikon | Jenis | Arti |
| :--: | --- | --- |
| ✅ | **Mandiri** | Kode lengkap tersedia di repositori dan dapat dijalankan setelah API Key dikonfigurasi. |
| 📖 | **Panduan Reproduksi** | Memerlukan repositori eksternal yang harus di-`git clone`. |
| 🚧 | **Dalam Proses** | Implementasi atau bukti penerimaan belum lengkap. |
+49
View File
@@ -0,0 +1,49 @@
# 第7章 · Agent の評価
> Agent の性能を比較可能なシグナルに変える。評価環境、データセット設計、指標体系、統計的有意性、可観測性、評価駆動の選定、そしてプロダクショングレードの内部評価とシミュレーション環境を扱う。
← [メイン README に戻る](../docs/ja/README.md) · 📖 [章の本文を読む](../book-ja/chapter7.ja.md)
## 実験の読み方
本文では短い mechanism skeleton で制御フローを説明し、実験ディレクトリには完全な SDK アダプター、ログ、テスト、受け入れ証拠を置きます。すべてのファイルを一行ずつ読む必要はありません。
- **Starter:** 目的・最小コマンド・受け入れ条件から始め、まず [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** エントリポイント、中心ループ、状態/メッセージ schema、ツール、検証器を追います。
- **Maintainer:** 最後にテスト、証拠 manifest、失敗処理、rollback 経路、provider adapter を読みます。
初読では認証情報、表示層、provider 互換層を飛ばし、数値を再現するときに戻ってください。
## 付随プロジェクト
| 実験 | プロジェクト | 種類 | 説明 |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | 計算、検索、データ処理などのシナリオを含め、複雑な推論のためにツールを使う Agent の能力の評価に焦点を当てる。 |
| 6-2 | `tau2-bench/` | 📖 | τ²-bench の段階別タスクを手動で完了し、軌跡を記録する。 |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | 4段階 Rubric を180件の構造化判定に適用し、根拠とハルシネーション拒否を記録する。 |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 60ケースを3システムで実行し、コストを完全に集計する。 |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | JSON、Markdown、Python 風のメモリ表現を対象に、実際の OpenRouter 呼び出しと決定論的なポリシーチェックで 11 件の trajectory-prefix 不良ケースを実行する。 |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 完全な 4×3×2×60 マトリクスで 1,440/1,440 件の実軌跡をエラーや未課金利用なしに保持し、検索・タスク指標と相互作用分析を完備、独立検証にも合格。 |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | 単一 GPU の正式実験で各 action-chunk 群 256 エピソードを完了。chunk 1 は 0/256、chunk 25 は 26/256 で、512 rollout のハッシュを保存。 |
| 6-2 | `terminal-bench/` | 📖 | Terminal-Bench は、実際のターミナル環境における AI Agent の性能をテストするためのベンチマークである。コードのコンパイルからモデルの訓練、サーバーのセットアップまで、Agent が実際のエンドツーエンドタスクをどう処理するかを評価する。約 100 タスクのデータセットと実行フレームワークを含み、さまざまな Agent 実装をサポートする。 |
| 6-2 | `SWE-bench/` | 📖 | SWE-bench は、大規模言語モデルが実際の GitHub issue を解決する能力を評価するためのベンチマークである。コードベースと issue の説明が与えられると、モデルは問題を解決するパッチを生成しなければならない。SWE-bench、SWE-bench Lite、SWE-bench Verified、SWE-bench Multimodal という複数のバージョンを含む。 |
| 6-2 | `GAIA/` | 📖 | GAIA は次世代の LLM(ツール拡張、効率的なプロンプティング、検索アクセスなどを備えたもの)を評価することを目的としている。さまざまな程度のツール利用と自律性を必要とし、曖昧さのない回答を持つ 450 以上の非自明な問題を含む。3 つの難易度レベルに分かれている。 |
| 6-2 | `OSWorld/` | 📖 | ファイル管理、アプリケーション操作、システム構成を含む、完全なオペレーティングシステム環境内で複雑なタスクを実行する Agent の能力を評価する。 |
| 6-2, 6-12 | `android_world/` | 📖 | アプリのナビゲーション、UI 操作、タスク完了能力を含む、Android モバイル環境における Agent の性能を評価する。 |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | さまざまな TTS 構成(異なるモデル/音声/速度)を用いて同じ難易度の高いテキストセットを合成し、次にマルチモーダルの LLM-as-a-Judge を用いて Rubric に従って各次元(明瞭さ、自然さなど)を採点し、結果を再現可能な構成比較表に集約する。 |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | ELO レーティングシステムに基づく Agent 性能リーダーボードを実装し、ペアワイズ比較を通じて異なる Agent の相対的な能力を評価する。 |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | 同一の中立的な Coding Harness の下で、GPT-5.6-sol と Claude Sonnet 5 が探索から最初の編集へ移るしきい値を比較する。18/18 セルが API エラーなしで完了し、[manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) が軌跡と集計を検証可能なハッシュで結び付ける。 |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | 典型的な複数ターンの Agent タスク(カスタマーサービスの返金)に対して全チェーンのコスト内訳を行う。カスタムの軽量トレーシングシステムを用いて各 LLM 呼び出しの入力/出力/キャッシュトークン、レイテンシ、コストを記録し、集計して「どのステップが最も高価か」を特定し、次に A/B テストを用いて KV Cache に優しい設計とコンテキスト圧縮による実際の節約を定量化する。 |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | 複数の OpenAI 互換 LLM API プロバイダーの横断的なベンチマークを実施する。ストリーミングインターフェースを用いて Time to First TokenTTFT)を正確に測定し、並行実行下でのエンドツーエンドレイテンシのパーセンタイル(p50/p95)、スループット、成功率を算出する。単一のコマンドで多次元の比較表を生成し、モデル選定がリーダーボードを見るだけではなく多面的なトレードオフであることを示す。 |
| 6-12 | [android-world](android-world/) | 📖 | 本書による AndroidWorld 上での T3A Agent の評価レポートと失敗分析ノート(実験 6-12 の起点。ベンチマークのソースコードではない) |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | 合成 DHIS2 スタイルの集計データに基づき、公衆衛生レポート Agent のツール呼び出し、計算精度、証拠引用、根拠のない主張を客観的に評価する。 |
> バッククォート表記の外部ベンチマークは別途 clone が必要です。[`android-world/`](android-world/)(ハイフン区切り)は本リポジトリ内の **T3A 評価分析ノート**(同ディレクトリの [README](android-world/README.md) を参照)であり、外部の `android_world/` ベンチマークソースとは別パスです。
## プロジェクトの種類
| アイコン | 種類 | 意味 |
| :--: | --- | --- |
| ✅ | **単独実行** | このリポジトリに完全なコードがあり、API キーを設定すれば実行できる |
| 📖 | **再現ガイド** | `git clone` が必要な**外部リポジトリ**に依存する詳細ドキュメント |
| 🚧 | **設計ドキュメント** | アーキテクチャ/実装計画のみで、実行可能なコードは未完成 |
+49
View File
@@ -0,0 +1,49 @@
# 제7장 · 에이전트 평가
> 에이전트 성능을 서로 비교할 수 있는 신호로 바꿉니다. 평가 환경, 데이터셋 설계, 지표 체계, 통계적 유의성, 관측 가능성, 평가 기반 선택, 프로덕션급 내부 평가·시뮬레이션 환경을 다룹니다.
← [한국어 메인 README로 돌아가기](../docs/ko/README.md) · 📖 [제7장 본문 읽기](../book-ko/chapter7.ko.md)
## 실험 읽는 방법
본문은 짧은 메커니즘 skeleton으로 제어 흐름을 설명하고, 실험 디렉터리에는 완전한 SDK 어댑터·로그·테스트·검수 증거를 둡니다. 모든 파일을 줄 단위로 읽을 필요는 없습니다.
- **Starter:** 목표, 최소 명령, 검수 조건부터 시작하고 다음에서 출발하세요: [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** 진입점, 핵심 루프, 상태/메시지 스키마, 도구와 verifier를 따라갑니다.
- **Maintainer:** 마지막으로 테스트, 증거 manifest, 실패 처리, rollback 경로와 provider adapter를 읽습니다.
첫 읽기에서는 credential, UI, provider 호환 계층을 건너뛰고 수치를 재현할 때 돌아오세요.
## 연계 프로젝트
| 실험 | 프로젝트 | 유형 | 설명 |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | τ²-bench의 이중 제어 멀티턴 평가를 실행하고, τ-bench와 작업 정의·성공 조건·사용자 시뮬레이터 설계를 비교합니다. |
| 6-2 | `tau2-bench/` | 📖 | τ²-bench의 난이도별 작업을 직접 수행하고 궤적을 기록합니다. 이는 실험 6-2에서 표본을 추출하는 여섯 가지 벤치마크 중 하나입니다. |
| 6-2 | `terminal-bench/` | 📖 | 실제 터미널 환경에서 AI 에이전트 성능을 시험하는 벤치마크입니다. 코드 컴파일부터 모델 학습, 서버 설정까지 실제 엔드투엔드 작업을 에이전트가 처리하는 방식을 평가합니다. 약 100개 작업으로 구성된 데이터셋과 실행 프레임워크를 포함하며 여러 에이전트 구현을 지원합니다. |
| 6-2 | `SWE-bench/` | 📖 | 대규모 언어 모델이 실제 GitHub Issue를 해결하는 능력을 평가하는 벤치마크입니다. 코드베이스와 Issue 설명을 받은 모델은 문제를 해결하는 패치를 생성해야 합니다. SWE-bench, SWE-bench Lite, SWE-bench Verified, SWE-bench Multimodal 등 여러 버전이 있습니다. |
| 6-2 | `GAIA/` | 📖 | 도구 확장, 효율적인 프롬프팅, 검색 접근 등을 갖춘 차세대 LLM을 평가합니다. 답이 명확하면서도 여러 수준의 도구 사용과 자율성이 필요한 450개 이상의 까다로운 질문을 담고 있으며, 세 가지 난이도로 나뉩니다. |
| 6-2 | `OSWorld/` | 📖 | 파일 관리, 애플리케이션 조작, 시스템 설정 등 완전한 운영체제 환경에서 복잡한 작업을 수행하는 에이전트의 능력을 평가합니다. |
| 6-2, 6-12 | `android_world/` | 📖 | 앱 탐색, UI 상호작용, 작업 완료 능력 등 Android 모바일 환경에서 에이전트 성능을 평가하는 외부 벤치마크 저장소입니다. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | 4단계 다차원 루브릭을 60개 사례 × 3개 시스템의 실제 판정 기록 180/180건에 모두 적용했습니다. 독립 검수 인덱스는 차원별 이유와 근거 또는 경계 사례, 환각 발생 시 즉시 탈락 조건을 검증하며 상태는 `complete`입니다. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 60개 사례 × 3개 시스템의 실제 궤적 180/180건을 오류 없이 수집했고, 원통화 기준 가격도 빠짐없이 반영했습니다. 검수 결과의 상태는 `complete`입니다. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | JSON, Markdown, Python 유사 메모리 표현에서 11개의 trajectory-prefix 불량 사례를 실제 OpenRouter 호출과 결정론적 정책 검사로 실행합니다. |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | 같은 고난도 텍스트 모음을 여러 TTS 설정(모델·음성·속도)으로 합성한 뒤, 멀티모달 LLM-as-a-Judge가 루브릭에 따라 명료도·자연스러움 등 각 항목을 채점합니다. 결과를 재현 가능한 설정 비교표로 집계합니다. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | [전체 정식 검증](elo-leaderboard/validation/runs/exp6-6-arena-20260731-v1/manifest.json)은 공개 Arena 레코드 1,799,991개(블라인드 투표 1,670,250개, 모델 129개)를 처리했습니다. 온라인 Elo와 Bradley-Terry 순위의 Spearman 상관은 0.787, Top-20 중복은 12/20이며, 승률 행렬·월별 스냅샷 17개·도표 3개·D3 애니메이션을 하나의 해시 manifest로 검증했습니다. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | 동일한 중립적 Coding Harness에서 GPT-5.6-sol과 Claude Sonnet 5가 탐색에서 첫 편집으로 전환하는 임계점을 비교합니다. 18/18 셀이 API 오류 없이 완료됐고, [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json)가 궤적과 요약을 검증 가능한 해시로 연결합니다. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | 전형적인 다중 턴 에이전트 작업(고객 서비스 환불)의 전체 비용을 단계별로 분석합니다. 맞춤형 경량 추적 시스템으로 LLM 호출마다 입력·출력·캐시 토큰, 지연 시간, 비용을 기록하고 집계해 가장 비싼 단계를 찾습니다. 이어 A/B 테스트로 KV Cache 친화적 설계와 컨텍스트 압축의 실제 절감 효과를 정량화합니다. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | 여러 OpenAI 호환 LLM API 제공자를 나란히 벤치마크합니다. 스트리밍 인터페이스로 첫 토큰까지 걸린 시간(TTFT)을 정밀 측정하고, 동시 실행 환경에서 엔드투엔드 지연 시간 백분위수(p50/p95), 처리량, 성공률을 계산합니다. 명령 하나로 다차원 비교표를 만들어 모델 선택이 단순한 순위표 이상의 복합적인 절충임을 보여 줍니다. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 전체 4×3×2×60 매트릭스에서 오류나 미가격 사용 없이 1,440/1,440개의 실제 궤적을 보존했으며, 검색·작업 지표와 상호작용 분석을 완비하고 독립 검증을 통과했습니다. |
| 6-12 | [android-world](android-world/) | 📖 | 이 저장소에 포함된 AndroidWorld T3A 평가 보고서와 실패 분석 노트입니다. 실험 6-12의 출발점이며 벤치마크 원본은 아닙니다. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | 단일 GPU 공식 실험에서 action-chunk 군별 256개 에피소드를 완료했습니다. chunk 1은 0/256, chunk 25는 26/256이며 512개 rollout 해시를 보존합니다. |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | 합성 DHIS2 형식 집계 데이터로 공중보건 보고 에이전트의 도구 호출, 계산 정확도, 근거 인용, 근거 없는 주장을 객관적으로 평가합니다. |
> 백틱으로 표기한 외부 벤치마크는 별도로 clone해야 합니다. 하이픈이 들어간 [`android-world/`](android-world/)는 이 저장소의 **T3A 평가 분석 노트**([README](android-world/README.md) 참고)이며, 외부 `android_world/` 벤치마크 원본과는 다른 경로입니다.
## 프로젝트 유형
| 아이콘 | 유형 | 의미 |
| :--: | --- | --- |
| ✅ | **독립 실행** | 전체 코드가 이 저장소에 있으며, API 키를 설정하면 실행할 수 있습니다. |
| 📖 | **재현 가이드** | **외부 저장소**를 `git clone`해야 하는 상세 안내 문서입니다. |
| 🚧 | **진행 중** | 구현은 있지만 실험 범위나 검수 증거가 아직 본문의 요구사항을 모두 충족하지 못했습니다. |
+84
View File
@@ -0,0 +1,84 @@
# 第 7 章 · Agent 的评估
> 把表现变成可比较信号:评估环境、指标、统计显著性、评估驱动选型
← [返回主目录](../README.md) · 📖 [读本章正文](../book/chapter7.md)
逐实验的正文要求、直接证据与未完成边界见 [验收台账](EXPERIMENT_LEDGER.md)。
## 如何阅读实验
正文伪代码先建立 reset → run → snapshot → verifier → record 的评估闭环;实验目录再展开统计与证据:
- **Starter**:从 [tau2-bench-eval](tau2-bench-eval/) 跑一个固定任务,先看环境 reset、轨迹保存和结果 verifier
- **Builder**:阅读 [user-memory-system-evaluation](user-memory-system-evaluation/) 的 Rubric/证据 schema,再看 [elo-leaderboard](elo-leaderboard/) 的配对统计;
- **Maintainer**:检查 veto 规则、seed/任务配对、bootstrap 或 McNemar 实现、manifest hash 和失败样本。
首次可跳过 provider 适配器、图表和长期开跑脚本;先确认“过程违规”和“最终失败”是两类独立信号。
## 配套项目
| 编号 | 项目 | 类型 | 一句话说明 |
| :--: | --- | :--: | --- |
| 7-1 | [tau2-bench-eval](tau2-bench-eval/) | ✅ | 已在固定上游提交上完成 5 个 telecom 双控任务:4/5 通过;保存原始轨迹、成本、内容哈希及错选线路导致漏做流量加油的失败分析 |
| 7-2 | [experiment-7-2-human-benchmark](experiment-7-2-human-benchmark/) | ✅ | Codex 作为人工操作员,预注册并完成 GAIA、AndroidWorld、SWE-bench Verified、τ²-bench、Terminal-Bench、OSWorld-Verified 各简单/中等/困难一题,共 18/18 个首轮正式结果:13 通过、5 失败;逐题保留任务、轨迹、官方评估及成败解释 |
| 7-2 | `terminal-bench/` | 📖 | Terminal-Bench 外部任务与执行框架;7-2 的三档人工操作结果与失败分析已收录于上行案例集 |
| 7-2 | `SWE-bench/` | 📖 | SWE-bench Verified 外部代码修复基准;7-2 的三档补丁轨迹与官方 harness 结果已收录于上行案例集 |
| 7-2 | `GAIA/` | 📖 | GAIA 外部数据集;7-2 的 Level 1/2/3 作答、核验与舍入失败边界已收录于上行案例集 |
| 7-2 | `OSWorld/` | 📖 | OSWorld-Verified 外部桌面环境;7-2 的三档 GUI 操作轨迹与官方结果已收录于上行案例集 |
| 7-2, 7-12 | `android_world/` | 📖 | 评估 Agent 在 Android 环境的应用导航、UI 交互与任务完成能力(外部基准仓库;7-2 的实际结果见上行) |
| 7-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | 四档多维 Rubric 已在 60 用例 × 3 系统的 180/180 条真实评判记录上完整执行;[独立验收索引](user-memory-system-evaluation/results/full_7_3_structured_rubric_evidence.json)验证逐维理由/证据或边界案例及幻觉一票否决,状态为 `complete` |
| 7-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 60 用例 × 3 系统共 180/180 条真实轨迹,零错误且原生币种定价完整;[验收结果](user-memory-system-evaluation/results/full_7_4_60_cases_costed.json)状态为 `complete` |
| 7-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | 已用真实 `openai/gpt-5.6-sol` 经 OpenRouter 完成 11 个 trajectory-prefix bad case × JSON/Markdown/Python-like 三种表示,共 33/33 个 API 单元、0 个 API 错误;三种表示均为 6/11 通过,结果和哈希保存在 `results/policy_prefix_live.json``results/manifest.json` |
| 7-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | [真实验收](tts-quality-eval/validation/mistral_multimodal_20260730/manifest.json)完成 OpenAI/Fish 两 provider × 四类语料的 8/8 双音频 Voxtral 四维评审;候选/参考音频逐项哈希,早期 Gemini/OpenRouter 失败证据仍保留 |
| 7-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | [正式全量验收](elo-leaderboard/validation/runs/exp7-7-arena-20260731-v1/manifest.json)处理 1,799,991 条公开 Arena 记录(1,670,250 条盲选票、129 个模型),在线 Elo 与 Bradley-Terry 排名 Spearman 0.787、Top-20 重合 12/20;胜率矩阵、17 个月度快照、三张图与 D3 动画均由同一 manifest 哈希绑定并复核通过 |
| 7-8 | [model-action-threshold](model-action-threshold/) | ✅ | 同一中性 Coding Harness 下完成 GPT-5.6-sol / Claude Sonnet 5 × 三任务 × 三次重复的 18/18 单元实测;[manifest](model-action-threshold/results/exp7-8-action-threshold-20260731-v1/manifest.json)零 API 错误并绑定完整轨迹与汇总哈希 |
| 7-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | 多轮 Agent 任务(客服退款)全链路成本拆解 + KV-cache 友好设计/上下文压缩的 A/B 节省量化 |
| 7-10 | [model-benchmark](model-benchmark/) | 🚧 | 完整 8K/32K/128K × 512/2048、限流爬坡、Agent 成本与 168 小时可用性 campaign 已实现;现有[验收清单](model-benchmark/results/manifest.json)只有真实 smoke/readiness,不能替代完整长期实验 |
| 7-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | [全矩阵验收](user-memory-system-evaluation/results/full_7_11_60_case_matrix.json)完成 60 用例 × 24 单元(4 嵌入 × 3 reranker × 2 主模型)共 1,440/1,440 条真实轨迹,零错误、零未定价用量,检索/任务指标与交互分析完整;[独立验证器](user-memory-system-evaluation/validation/verify_full_matrix_20260731.py)复核通过(ALL CHECKS PASSED),后端替代方案如实记录于 [readiness 证据](user-memory-system-evaluation/results/full_matrix_backend_readiness_20260731.json) |
| 7-12 | [android-world](android-world/) | ✅ | [完整候选实验证据](android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json)保留 116 任务 × 5 轮的 580/580 条唯一 episode(包括评估失败),运行时错误为零:严格 T3A 成功 26 条(4.4828%),平均 evaluator reward 0.133621,由 77 条满分状态与 1 条 `0.5` 部分 reward 组成。实验在完成官方初始化且配齐 24/24 应用的 Pixel 6/API-33 上执行,本地 Qwen2.5-7Brevision `a09a35458c702b33eeacc393d103063234e8bc28`)通过 vLLM 0.19.0 运行于 RTX PRO 6000 Blackwell 96 GB。执行与证据已完成,但未批准部署;候选 Qwen 与配对源 Doubao 不同,因而不支持同模型提升或非劣性结论 |
| 7-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | [正式单卡运行](openvla-robotwin2-eval/validation/runs/exp7-13-localgpu-20260803-v1/manifest.json)完成 chunk 1/25 各 128 IID + 128 OOD episodes,严格门禁及 512 个 rollout hash 全通过;chunk 1 为 0/256、chunk 25 为 26/256,低绝对成功率作为真实结果保留 |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | 基于合成 DHIS2 风格汇总数据,客观评估公共卫生报告 Agent 的工具调用、计算准确性、证据引用与无依据声明 |
> 📖 表中带反引号的外部基准需自行克隆。[`android-world/`](android-world/)(连字符)是本仓库内的 **T3A 评估分析笔记**(见该目录 [README](android-world/README.md)),与外部 `android_world/` 基准源码不是同一路径。
## 跨章 Bad Case 回归协议
正文新增的两类作用域/保真度 Bad Case 评估不把训练代码重复复制到第七章:第七章负责记录首个错误、片段作用域、逐层字符串哈希和轨迹前缀回归;第八章的 [`curly-quote-sft`](../chapter8/curly-quote-sft/) 与 [`exact-copy-sft`](../chapter8/exact-copy-sft/) 复用这些标签生成训练数据,并在独立边界集和保留集上回归。前者按中文自然语言、英文原文、代码和 JSON 作用域评分,后者按 byte/code-point/token exactness 和真实工具参数匹配评分。
## 实验 7-1 / 7-2 外部复现锚点
以下映射以[正文](../book/chapter7.md)为准。SHA 来自对应 checkout 的 `origin``HEAD`。7-1 已保留五任务正式运行的[验收证据](tau2-bench-eval/validation/runs/exp7-1-openrouter-gpt41mini-telecom-20260802-v1/manifest.json);7-2 的 18 个分级人工操作案例、正式结果与兼容边界见[独立报告](experiment-7-2-human-benchmark/README.md)。下表继续保留复现来源、路径和入口。
| 实验 | 上游与本地路径 | 固定提交 | 正文对应入口 |
| :--: | --- | --- | --- |
| 7-17-2 的 τ²-bench 样本 | [`sierra-research/tau2-bench`](https://github.com/sierra-research/tau2-bench) → `chapter7/tau2-bench` | `8d005b0e5b9e4af0bc055886fa7f95fc86d1710e` | 正文要求重点观察新增的双控 telecom 领域:`tau2 run --domain telecom --agent-llm <model> --user-llm <model> --num-trials 1 --num-tasks 5` |
| 7-1 原始 τ-bench 对照(仅溯源) | [论文](https://arxiv.org/abs/2406.12045) · [`sierra-research/tau-bench`](https://github.com/sierra-research/tau-bench/tree/59a200c6d575d595120f1cb70fea53cef0632f6b)**不承诺本地 checkout** | `59a200c6d575d595120f1cb70fea53cef0632f6b` | 该历史版本入口:`python run.py --agent-strategy tool-calling --env retail --model gpt-4o --model-provider openai --user-model gpt-4o --user-model-provider openai --user-strategy llm --max-concurrency 10` |
| 7-2 GAIA | [`gaia-benchmark/GAIA`](https://huggingface.co/datasets/gaia-benchmark/GAIA) → `chapter7/GAIA` | `682dd723ee1e1697e00360edccf2366dc8418dd9` | 从 `2023/validation/metadata.level1.parquet``metadata.level2.parquet``metadata.level3.parquet` 各选一题人工完成并核对答案 |
| 7-2 AndroidWorld | [`google-research/android_world`](https://github.com/google-research/android_world) → `chapter7/android_world` | `0e95d641e244504c22087cc29b013f3b2428a261` | `python minimal_task_runner.py --task=ContactsAddContact`(先按上游 README 配置 emulator |
| 7-2 SWE-Bench Verified | [`SWE-bench/SWE-bench`](https://github.com/SWE-bench/SWE-bench) → `chapter7/SWE-bench` | `5cd4be9fb23971679cbbafe5a0ecade27cef99be` | 安装后先用 `python -m swebench.harness.run_evaluation --predictions_path gold --max_workers 1 --instance_ids sympy__sympy-20590 --run_id validate-gold` 验证 harness,再人工处理选定 Verified issue |
| 7-2 Terminal-Bench | [`laude-institute/terminal-bench`](https://github.com/laude-institute/terminal-bench) → `chapter7/terminal-bench` | `8384a179b1b8688f6ea5233a4d9d51218df1ac96` | 任务定义在 `tasks/`;若要核对 harness 参数,运行 `tb run --help` |
| 7-2 OSWorld-Verified | [`xlang-ai/OSWorld`](https://github.com/xlang-ai/OSWorld) → `chapter7/OSWorld` | `8365edc975efd0477a0d62444a5beed562ab5a7b` | `python quickstart.py --provider_name vmware --path_to_vm "path/to/your/vm.vmx"`;再从 Verified 任务中抽样人工完成 |
从仓库根目录取得同一版本:
```bash
git clone https://github.com/sierra-research/tau2-bench.git chapter7/tau2-bench && git -C chapter7/tau2-bench checkout --detach 8d005b0e5b9e4af0bc055886fa7f95fc86d1710e
git clone https://huggingface.co/datasets/gaia-benchmark/GAIA chapter7/GAIA && git -C chapter7/GAIA checkout --detach 682dd723ee1e1697e00360edccf2366dc8418dd9
git clone https://github.com/google-research/android_world.git chapter7/android_world && git -C chapter7/android_world checkout --detach 0e95d641e244504c22087cc29b013f3b2428a261
git clone https://github.com/SWE-bench/SWE-bench.git chapter7/SWE-bench && git -C chapter7/SWE-bench checkout --detach 5cd4be9fb23971679cbbafe5a0ecade27cef99be
git clone https://github.com/laude-institute/terminal-bench.git chapter7/terminal-bench && git -C chapter7/terminal-bench checkout --detach 8384a179b1b8688f6ea5233a4d9d51218df1ac96
git clone https://github.com/xlang-ai/OSWorld.git chapter7/OSWorld && git -C chapter7/OSWorld checkout --detach 8365edc975efd0477a0d62444a5beed562ab5a7b
```
原始 τ-bench 行只用于复核 7-1 的历史设计差异,不在本仓库的 checkout 清单中。其当前 README 已明确警告:该仓库的 airline/retail 任务版本过时,应使用后续的 [`tau2-bench`](https://github.com/sierra-research/tau2-bench)(现已继续演进为 τ³-bench)获取修订任务与新领域。因此,不应把历史 τ-bench 的 retail 命令当成当前 τ²/τ³-bench 的推荐运行入口。
实验 7-2 是**操作员亲自执行并记录轨迹**,不是把六套 Agent harness 全跑一遍。本仓库的[已完成案例集](experiment-7-2-human-benchmark/)由 Codex 明确署名为人工操作员,并分别记录每个基准的简单、中等、困难任务 ID、环境版本、步骤、最终答案/状态与标准验证结果;失败案例未在评估后修改或重跑。
## 项目类型说明
| 图标 | 类型 | 含义 |
| :--: | --- | --- |
| ✅ | **可独立运行** | 本仓库自带完整代码,配置好 API Key 即可运行 |
| 📖 | **复现指南** | 依赖需自行 `git clone` 的**外部仓库**(训练框架、评测基准等) |
| 🚧 | **进行中** | 已有实现,但实验范围或验收证据尚未满足正文全部要求 |
+48
View File
@@ -0,0 +1,48 @@
# Глава 7 · Оценка агентов
> Превращает качество агента в сравнимые сигналы. Охватывает среды оценки, дизайн наборов данных, системы метрик, статистическую значимость, наблюдаемость, выбор на основе оценки, а также промышленные внутренние среды оценки и симуляции.
← [К оглавлению](../docs/ru/README.md) · 📖 [Читать главу](../book-ru/chapter7.md)
## Как читать эксперименты
В основном тексте короткие скелеты механизмов объясняют поток управления; в каталогах экспериментов находятся полные адаптеры SDK, журналы, тесты и приёмочные доказательства. Читать каждый файл построчно не требуется.
- **Starter:** Начните с цели, минимальной команды и условий приёмки; начните с [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Проследите точку входа, основной цикл, схему состояния/сообщений, инструменты и проверяющий модуль.
- **Maintainer:** Затем изучите тесты, манифесты доказательств, обработку сбоев, откат и адаптеры провайдеров.
При первом чтении можно пропустить ключи, слой представления и совместимость провайдеров; вернитесь при воспроизведении чисел.
## Сопутствующие проекты
| Эксп. | Проект | Тип | Описание |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | Фокусируется на оценке способности агента использовать инструменты для сложных рассуждений, включая сценарии вычислений, поиска и обработки данных. |
| 6-2 | `tau2-bench/` | 📖 | Ручное выполнение градуированных задач τ²-bench с записью траекторий. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | Применяет четырёхуровневую рубрику к 180 структурированным оценкам с доказательствами и запретом галлюцинаций. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Запускает 60 случаев на трёх системах с полным учётом стоимости. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Запускает 11 негативных случаев с префиксами траекторий для представлений памяти в JSON, Markdown и Python-подобном формате, с реальными вызовами OpenRouter и детерминированными проверками политик. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Полная матрица 4×3×2×60 сохранила 1 440/1 440 реальных траекторий без ошибок и неучтённого использования, с полными метриками поиска и задач, анализом взаимодействия и успешно пройденной независимой проверкой. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | Официальный запуск на одной GPU завершил по 256 эпизодов в каждой группе: chunk 1 — 0/256, chunk 25 — 26/256; сохранены хэши 512 rollout. |
| 6-2 | `terminal-bench/` | 📖 | Terminal-Bench — бенчмарк для проверки качества ИИ-агентов в реальных терминальных средах. От компиляции кода до обучения моделей и настройки серверов — оценивает, как агенты справляются с реальными сквозными задачами. Включает набор из ~100 задач и фреймворк исполнения, поддерживая разные реализации агентов. |
| 6-2 | `SWE-bench/` | 📖 | SWE-bench — бенчмарк для оценки способности больших языковых моделей решать реальные GitHub-issue. По кодовой базе и описанию проблемы модель должна сгенерировать патч, устраняющий проблему. Включает несколько версий: SWE-bench, SWE-bench Lite, SWE-bench Verified и SWE-bench Multimodal. |
| 6-2 | `GAIA/` | 📖 | GAIA нацелен на оценку LLM нового поколения (с дополнением инструментами, эффективным промптингом, доступом к поиску и т. д.). Содержит 450+ нетривиальных вопросов, требующих разной степени использования инструментов и автономности, с однозначными ответами. Разделён на 3 уровня сложности. |
| 6-2 | `OSWorld/` | 📖 | Оценивает способность агентов выполнять сложные задачи в полноценной среде операционной системы, включая управление файлами, работу с приложениями и настройку системы. |
| 6-2, 6-12 | `android_world/` | 📖 | Оценивает качество агента в мобильной среде Android: навигация по приложениям, взаимодействие с UI и выполнение задач (внешний репозиторий бенчмарка). |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | Синтезирует один и тот же набор сложных текстов разными конфигурациями TTS (модель/голос/скорость), затем через мультимодальный LLM-as-a-Judge оценивает каждое измерение (чёткость, естественность и т. д.) по рубрике, агрегируя результаты в воспроизводимую сравнительную таблицу конфигураций. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Реализует таблицу лидеров агентов на основе рейтинговой системы ELO, оценивая относительные способности разных агентов через попарные сравнения. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | Сравнивает GPT-5.6-sol и Claude Sonnet 5 в момент перехода от исследования к первой правке при одном и том же нейтральном Coding Harness; все 18/18 ячеек завершены без ошибок API, а [манифест](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) связывает траектории и сводки проверяемыми хешами. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Делает сквозную декомпозицию стоимости для типичной многораундовой задачи агента (возврат в поддержке): с помощью собственной лёгкой системы трассировки записывает входные/выходные/кэш-токены, задержку и стоимость каждого вызова LLM, агрегирует, чтобы найти «самый дорогой шаг», а затем A/B-тестами количественно оценивает реальную экономию от дружественного к KV-кэшу дизайна и сжатия контекста. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | Проводит горизонтальный бенчмарк нескольких OpenAI-совместимых провайдеров LLM API. Через потоковый интерфейс точно измеряет время до первого токена (TTFT), рассчитывает перцентили сквозной задержки (p50/p95), пропускную способность и долю успеха при конкурентности. Одной командой строит многомерную сравнительную таблицу, показывая, что выбор модели — это многогранный компромисс, а не просто взгляд на таблицу лидеров. |
| 6-12 | [android-world](android-world/) | 📖 | Внутрирепозиторные отчёт и разбор ошибок T3A на AndroidWorld (старт эксперимента 6-12; не исходники бенчмарка). |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | На синтетических агрегированных данных в стиле DHIS2 объективно оценивает вызовы инструментов, точность вычислений, цитирование доказательств и неподтверждённые утверждения у агента отчётности в общественном здравоохранении. |
> Внешние бенчмарки в обратных кавычках нужно клонировать отдельно. [`android-world/`](android-world/) (через дефис) — **аналитические заметки T3A** в этом репозитории (см. [README](android-world/README.md)), это не тот же путь, что внешний `android_world/`.
## Типы проектов
| Значок | Тип | Значение |
| :--: | --- | --- |
| ✅ | **Автономный** | Полный код в этом репозитории, запускается после настройки API-ключа |
| 📖 | **Гайд по воспроизведению** | Подробный документ, зависящий от **внешних репозиториев** через `git clone` |
| 🚧 | **Проектный документ** | Только архитектура/план реализации, рабочий код ещё в разработке |
+49
View File
@@ -0,0 +1,49 @@
# அத்தியாயம் 7 · ஏஜென்ட் மதிப்பீடு
> ஏஜெண்டின் செயல்திறனை ஒப்பிடக்கூடிய சமிக்ஞைகளாக மாற்றுகிறது. மதிப்பீட்டுச் சூழல்கள், தரவுத்தொகுப்பு வடிவமைப்பு, அளவுகோல் அமைப்பு முதல் புள்ளியியல் முக்கியத்துவம், கண்காணிப்புத்தன்மை, மதிப்பீடு-இயக்கப்படும் தேர்வு வரை, மேலும் உற்பத்தி-தர உள் மதிப்பீடு மற்றும் உருவகப்படுத்துதல் சூழல்கள் வரை விவரிக்கிறது.
← [முக்கிய README க்குத் திரும்பு](../docs/ta/README.md) · 📖 [அத்தியாய உரையைப் படி](../book-ta/chapter7.ta.md)
## சோதனைகளை எப்படிப் படிப்பது
முதன்மை உரை குறுகிய mechanism skeleton-களால் control flow-ஐ விளக்குகிறது; முழு SDK adapters, logs, tests, acceptance evidence ஆகியவை experiment கோப்பகத்தில் உள்ளன. ஒவ்வொரு கோப்பையும் வரி வரியாகப் படிக்க வேண்டியதில்லை.
- **Starter:** இலக்கு, குறைந்தபட்ச கட்டளை, ஏற்றுக்கொள்ளும் நிபந்தனைகளில் தொடங்குங்கள்; முதலில் [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** நுழைவுப் புள்ளி, மையச் சுழற்சி, state/message schema, கருவிகள், verifier ஆகியவற்றைப் பின்தொடருங்கள்.
- **Maintainer:** பின்னர் tests, evidence manifest, தோல்வி கையாளல், rollback பாதை, provider adapter ஆகியவற்றைப் படியுங்கள்.
முதல் வாசிப்பில் credentials, UI, provider-compatibility அடுக்குகளைத் தவிர்க்கலாம்; முடிவுகளை மீண்டும் உருவாக்கும்போது திரும்பிப் பாருங்கள்.
## துணை திட்டங்கள்
| சோதனை | Project | Type | Description |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | கணிப்பு, தேடல் மற்றும் தரவுச் செயலாக்கம் போன்ற சூழ்நிலைகள் உட்பட, கருவிகளைப் பயன்படுத்தி ஏஜென்ட் சிக்கலான பகுத்தறிவு மேற்கொள்ளும் திறனை மதிப்பீடு செய்வதில் கவனம் செலுத்துகிறது. |
| 6-2 | `tau2-bench/` | 📖 | தரப்படுத்தப்பட்ட τ²-bench பணிகளை கைமுறையாக முடித்து trajectory-களை பதிவு செய்கிறது. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | 180 கட்டமைக்கப்பட்ட மதிப்பீடுகளில் நான்கு-நிலை rubric-ஐ ஆதாரங்களுடனும் hallucination veto-வுடனும் இயக்குகிறது. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | முழுமையான செலவுக் கணக்கீட்டுடன் மூன்று அமைப்புகளில் 60 cases-ஐ இயக்குகிறது. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | JSON, Markdown மற்றும் Python போன்ற நினைவகப் பிரதிநிதித்துவங்களில் 11 trajectory-prefix தவறான cases-ஐ உண்மையான OpenRouter அழைப்புகளுடனும் நிர்ணயிக்கப்பட்ட policy சோதனைகளுடனும் இயக்குகிறது. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | முழு 4×3×2×60 matrix-ல் 1,440/1,440 உண்மையான trajectory-கள் பிழையோ விலையிடப்படாத பயன்பாடோ இன்றி சேமிக்கப்பட்டுள்ளன; retrieval/task அளவீடுகள், interaction analysis மற்றும் சுயாதீன சரிபார்ப்பு நிறைவேறியுள்ளன. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | ஒற்றை GPU அதிகாரப்பூர்வ இயக்கம் ஒவ்வொரு action-chunk குழுவிலும் 256 episode-களை முடித்தது; chunk 1: 0/256, chunk 25: 26/256, மேலும் 512 rollout hash-கள் சேமிக்கப்பட்டன. |
| 6-2 | `terminal-bench/` | 📖 | Terminal-Bench என்பது உண்மையான முனையச் சூழலில் AI ஏஜென்ட்டின் செயல்திறனைச் சோதிக்கும் ஒரு தரநிர்ணயம் ஆகும். குறியீட்டைத் தொகுப்பதில் இருந்து மாதிரிகளைப் பயிற்றுவித்தல், சேவையகங்களை அமைத்தல் வரை, ஏஜென்ட் உண்மையான end-to-end பணிகளை எவ்வாறு கையாள்கிறது என்பதை மதிப்பீடு செய்கிறது. சுமார் 100 பணிகள் கொண்ட தரவுத்தொகுப்பு மற்றும் செயல்படுத்தும் கட்டமைப்பை உள்ளடக்கியது, பல ஏஜென்ட் செயலாக்கங்களை ஆதரிக்கிறது. |
| 6-2 | `SWE-bench/` | 📖 | SWE-bench என்பது பெரிய மொழி மாதிரிகள் உண்மையான GitHub சிக்கல்களைத் தீர்க்கும் திறனை மதிப்பீடு செய்யும் ஒரு தரநிர்ணயம் ஆகும். குறியீட்டுத் தளமும் சிக்கல் விளக்கமும் கொடுக்கப்பட்டால், மாதிரி சிக்கலைத் தீர்க்கக்கூடிய ஒரு பேட்சை உருவாக்க வேண்டும். SWE-bench, SWE-bench Lite, SWE-bench Verified மற்றும் SWE-bench Multimodal உட்படப் பல பதிப்புகளைக் கொண்டுள்ளது. |
| 6-2 | `GAIA/` | 📖 | GAIA என்பது அடுத்த தலைமுறை LLM-களை (கருவி மேம்பாடு, திறமையான அறிவுறுத்தல், தேடல் அணுகல் போன்ற திறன்கள் கொண்ட LLM-கள்) மதிப்பீடு செய்வதை நோக்கமாகக் கொண்டது. வெவ்வேறு அளவிலான கருவிகள் மற்றும் தன்னாட்சி தேவைப்படும் 450+ எளிதல்லாத கேள்விகளைக் கொண்டுள்ளது, பதில்கள் தெளிவானவை மற்றும் இருபொருளற்றவை. 3 சிரம நிலைகளாகப் பிரிக்கப்பட்டுள்ளது. |
| 6-2 | `OSWorld/` | 📖 | முழுமையான இயக்க முறைமைச் சூழலில் ஏஜென்ட் சிக்கலான பணிகளைச் செயல்படுத்தும் திறனை மதிப்பீடு செய்கிறது, கோப்பு மேலாண்மை, பயன்பாட்டுச் செயல்பாடுகள் மற்றும் கணினி கட்டமைப்பு ஆகியவை உட்பட. |
| 6-2, 6-12 | `android_world/` | 📖 | Android மொபைல் சூழலில் ஏஜென்ட்டின் செயல்திறனை மதிப்பீடு செய்கிறது, பயன்பாட்டு வழிசெலுத்தல், UI தொடர்பு மற்றும் பணி நிறைவு திறன் ஆகியவை உட்பட (வெளிப்புற benchmark களஞ்சியம்). |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | பல TTS கட்டமைப்புகளை (வெவ்வேறு model/voice/speed) பயன்படுத்தி ஒரே சவாலான உரைத் தொகுப்பை ஒலியாக்கி, பின்னர் மல்டிமோடல் LLM-as-a-Judge மூலம் Rubric இன்படி ஒவ்வொரு பரிமாணத்திற்கும் (தெளிவுத்தன்மை/இயற்கைத்தன்மை போன்றவை) மதிப்பெண் வழங்கி, மீண்டும் உருவாக்கக்கூடிய கட்டமைப்பு ஒப்பீட்டு அட்டவணையாகத் தொகுக்கிறது. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | ELO மதிப்பீட்டு அமைப்பின் அடிப்படையில் ஏஜென்ட் செயல்திறன் தரவரிசைப் பட்டியலைச் செயல்படுத்துகிறது, ஜோடிவரிசைப் போட்டிகள் மூலம் வெவ்வேறு ஏஜென்ட்டுகளின் ஒப்பீட்டுத் திறன்களை மதிப்பீடு செய்கிறது. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | ஒரே நடுநிலையான Coding Harness-இல் GPT-5.6-sol மற்றும் Claude Sonnet 5 ஆகியவை ஆராய்ச்சியிலிருந்து முதல் திருத்தத்துக்கு நகரும் நுழைவுநிலையை ஒப்பிடுகிறது; 18/18 செல்களும் API பிழையின்றி நிறைவடைந்தன, மேலும் [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) செயல்தடங்களையும் சுருக்கங்களையும் சரிபார்க்கக்கூடிய hash-களால் இணைக்கிறது. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | ஒரு பொதுவான பல-சுற்று ஏஜென்ட் பணிக்கு (வாடிக்கையாளர் சேவை பணம் திரும்பப்பெறுதல்) முழு-சங்கிலி செலவுப் பிரிவை மேற்கொள்கிறது: சுயமாக உருவாக்கிய இலகுரக tracing மூலம் ஒவ்வொரு LLM அழைப்பின் உள்ளீடு/வெளியீடு/கேச் டோக்கன்கள், தாமதம் மற்றும் செலவைப் பதிவு செய்து, "எந்தப் படி அதிக செலவாகிறது" என்பதைத் தொகுத்தறிந்து, பின்னர் A/B சோதனைகள் மூலம் KV-cache நட்பு வடிவமைப்பு + சூழல் சுருக்கம் ஆகியவற்றால் கிடைக்கும் உண்மையான சேமிப்பை அளவிடுகிறது. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | பல OpenAI-இணக்கமான LLM API வழங்குநர்களுக்கு கிடைமட்ட ஒப்பீட்டுத் தரநிர்ணயம் நடத்துகிறது, நீரோட்ட இடைமுகம் மூலம் முதல் டோக்கன் தாமதத்தை (TTFT) துல்லியமாக அளவிடுகிறது, இணையான சுமையின் கீழ் end-to-end தாமத குவான்டைல்கள் (p50/p95), செயலாக்க விகிதம் (throughput) மற்றும் வெற்றி விகிதம் ஆகியவற்றை அளவிடுகிறது, ஒரே கட்டளையில் பல்-பரிமாண ஒப்பீட்டு அட்டவணையை உருவாக்குகிறது, மேலும் மாதிரி தேர்வு என்பது தரவரிசைப் பட்டியலை மட்டும் பார்ப்பதல்ல, பல்-பரிமாணச் சமநிலை என்பதை விளக்குகிறது. |
| 6-12 | [android-world](android-world/) | 📖 | இந்தக் களஞ்சியத்தில் உள்ள AndroidWorld மீதான T3A மதிப்பீட்டு அறிக்கை மற்றும் தோல்வி பகுப்பாய்வு குறிப்புகள் (சோதனை 6-12 தொடக்கம்; benchmark மூலக் குறியீடு அல்ல). |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | செயற்கையான DHIS2-பாணி ஒருங்கிணைந்த தரவைப் பயன்படுத்தி, பொது சுகாதார அறிக்கையிடல் ஏஜென்ட்டின் கருவி அழைப்புகள், கணக்கீட்டுத் துல்லியம், ஆதார மேற்கோள்கள் மற்றும் ஆதாரமற்ற கூற்றுகளைப் புறநிலையாக மதிப்பீடு செய்கிறது. |
> 📖 Backtick குறியில் உள்ள வெளிப்புற benchmark-களை தனியாக clone செய்ய வேண்டும். [`android-world/`](android-world/) (hyphen) என்பது இந்தக் களஞ்சியத்தின் **T3A மதிப்பீட்டு பகுப்பாய்வுக் குறிப்புகள்** (அதன் [README](android-world/README.md) பார்க்க); இது வெளிப்புற `android_world/` benchmark மூலக் குறியீட்டுப் பாதை அல்ல.
## திட்ட வகைகள்
| சின்னம் | வகை | பொருள் |
| :--: | --- | --- |
| ✅ | **தனித்து இயங்கும்** | முழு குறியீடு இந்த களஞ்சியத்தில், API Key உள்ளமைத்தவுடன் இயங்கும் |
| 📖 | **மறு உருவாக்க வழிகாட்டி** | **வெளிப்புற களஞ்சியங்களை** `git clone` செய்ய வேண்டிய விரிவான ஆவணம் |
| 🚧 | **வடிவமைப்பு ஆவணம்** | கட்டமைப்பு/செயலாக்கத் திட்டம் மட்டும், இயங்கும் குறியீடு இன்னும் WIP |
+49
View File
@@ -0,0 +1,49 @@
# Bölüm 7 · Agent Değerlendirmesi
> Agent performansını karşılaştırılabilir sinyallere dönüştürür. Değerlendirme ortamlarını, veri kümesi tasarımını, metrik sistemlerini, istatistiksel anlamlılığı, gözlemlenebilirliği, değerlendirme odaklı seçimi ve üretim seviyesinde dahili değerlendirme ile simülasyon ortamlarını kapsar.
← [Ana README'ye dön](../README.tr.md) · 📖 [Bölüm metnini oku](../book-tr/chapter7.tr.md)
## Deneyler nasıl okunur
Metin, kontrol akışını açıklamak için kısa mekanizma skeleton'ları kullanır; deney dizininde tam SDK adaptörleri, günlükler, testler ve kabul kanıtı bulunur. Her dosyayı satır satır okumanız gerekmez.
- **Starter:** Hedef, en kısa komut ve kabul koşullarıyla başlayın; önce [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Giriş noktasını, ana döngüyü, durum/mesaj şemasını, araçları ve doğrulayıcıyı izleyin.
- **Maintainer:** Son olarak testleri, kanıt manifestlerini, hata işlemeyi, rollback yollarını ve sağlayıcı adaptörlerini okuyun.
İlk okumada kimlik bilgisi yükleme, sunum katmanı ve sağlayıcı uyumluluğunu atlayıp sayıları yeniden üretirken dönün.
## Eşlik Eden Projeler
| Proje | Tür | Açıklama |
| --- | :--: | --- |
| `terminal-bench/` | 📖 | Terminal-Bench, gerçek terminal ortamlarında AI Agent performansını test etmek için bir kıstastır. Kod derlemekten model eğitmeye ve sunucu kurmaya kadar, Agent'ların gerçek uçtan uca görevleri nasıl ele aldığını değerlendirir. ~100 görevlik bir veri kümesi ve çeşitli Agent uygulamalarını destekleyen bir yürütme çerçevesi içerir. |
| `SWE-bench/` | 📖 | SWE-bench, büyük dil modellerinin gerçek GitHub issue'larını çözme yeteneğini değerlendirmek için bir kıstastır. Bir kod tabanı ve issue açıklaması verildiğinde model, sorunu çözen bir yama üretmelidir. SWE-bench, SWE-bench Lite, SWE-bench Verified ve SWE-bench Multimodal dahil birden çok sürüm içerir. |
| `GAIA/` | 📖 | GAIA, yeni nesil LLM'leri (araç genişletmeli, verimli promptlamalı, arama erişimli vb.) değerlendirmeyi amaçlar. Farklı derecelerde araç kullanımı ve özerklik gerektiren, belirsiz olmayan yanıtlara sahip 450'den fazla önemsiz olmayan soru içerir. 3 zorluk seviyesine ayrılmıştır. |
| `OSWorld/` | 📖 | Ajanların dosya yönetimi, uygulama işletme ve sistem yapılandırması dahil, eksiksiz bir işletim sistemi ortamı içinde karmaşık görevleri yerine getirme yeteneğini değerlendirir. |
| `android_world/` (6-2, 6-12) | 📖 | Ajan performansını bir Android mobil ortamında değerlendirir; uygulama gezinme, UI etkileşimi ve görev tamamlama yeteneklerini kapsar. |
| `tau2-bench/` (6-1) | 📖 | Bir ajanın hesaplama, arama ve veri işleme gibi senaryolar dahil, karmaşık muhakeme için araç kullanma yeteneğini değerlendirmeye odaklanır. |
| `tau2-bench/` (6-2) | 📖 | Derecelendirilmiş τ²-bench görevlerini elle tamamlar ve yörüngeleri kaydeder. |
| [user-memory-evaluation](../chapter3/user-memory-evaluation/) (6-3) | ✅ | Dört seviyeli rubric'i kanıt ve halüsinasyon vetosuyla 180 yapılandırılmış değerlendirmede çalıştırır. |
| [user-memory-system-evaluation](user-memory-system-evaluation/) (6-4) | ✅ | Tam maliyet muhasebesiyle üç sistem üzerinde 60 vakayı çalıştırır. |
| [user-memory-policy-eval](user-memory-policy-eval/) (6-5) | ✅ | JSON, Markdown ve Python benzeri bellek gösterimlerinde 11 hatalı yörünge öneki vakasını gerçek OpenRouter çağrıları ve belirlenimci politika kontrolleriyle çalıştırır. |
| [user-memory-system-evaluation](user-memory-system-evaluation/) (6-11) | ✅ | Tam 4×3×2×60 matrisinde 1.440/1.440 gerçek yörüngeyi hata veya fiyatlandırılmamış kullanım olmadan korur; erişim/görev metrikleri, etkileşim analizi ve bağımsız doğrulama tamamlanmıştır. |
| [openvla-robotwin2-eval](openvla-robotwin2-eval/) (6-13) | ✅ | Tek GPU'lu resmi çalışma kol başına 256 episode tamamladı; chunk 1 0/256, chunk 25 26/256 aldı ve 512 rollout hash'i saklandı. |
| [elo-leaderboard](elo-leaderboard/) (6-7) | ✅ | ELO derecelendirme sistemine dayalı bir ajan performansı liderlik tablosu uygular; ikili karşılaştırmalarla farklı ajanların göreli yeteneklerini değerlendirir. |
| [model-action-threshold](model-action-threshold/) (6-8) | ✅ | Aynı tarafsız Coding Harness altında GPT-5.6-sol ile Claude Sonnet 5'in keşiften ilk düzenlemeye geçiş eşiğini karşılaştırır; 18/18 hücre API hatası olmadan tamamlanmış, [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) ise yürütme izleriyle özetleri doğrulanabilir hash'lerle bağlamıştır. |
| [model-benchmark](model-benchmark/) (6-10) | 🚧 | Birden çok OpenAI uyumlu LLM API sağlayıcısının yatay bir kıstasını yapar. İlk Token Süresini (TTFT) hassas biçimde ölçmek için bir akış arayüzü kullanır, eşzamanlılık altında uçtan uca gecikme yüzdeliklerini (p50/p95), verimi ve başarı oranını hesaplar. Tek bir komut, model seçiminin yalnızca bir liderlik tablosuna bakmaktan ibaret olmadığını, çok yönlü bir ödünleşim olduğunu gösteren çok boyutlu bir karşılaştırma tablosu üretir. |
| [agent-cost-analysis](agent-cost-analysis/) (6-9) | ✅ | Tipik çok turlu bir ajan görevi (müşteri hizmetleri iadesi) için tam zincir maliyet analizi yapar: her LLM çağrısı için girdi/çıktı/önbellek token'larını, gecikmeyi ve maliyeti kaydetmek için özel hafif bir izleme sistemi kullanır, "hangi adımın en pahalı olduğunu" belirlemek için toplar, ardından KV-cache dostu tasarım ve context sıkıştırmadan elde edilen gerçek tasarrufları nicelleştirmek için A/B testi kullanır. |
| [tts-quality-eval](tts-quality-eval/) (6-6) | ✅ | Aynı zorlu metin kümesini çeşitli TTS yapılandırmalarıyla (farklı model/ses/hız) sentezler, ardından her boyutu (netlik, doğallık vb.) bir Rubric'e göre puanlamak için çok modlu bir LLM-as-a-Judge kullanır, sonuçları yeniden üretilebilir bir yapılandırma karşılaştırma tablosunda toplar. |
| [android-world](android-world/) (6-12) | 📖 | AndroidWorld üzerinde T3A Agent değerlendirmesi ve başarısızlık analizi için başlangıç raporu; kıstasın kaynak kodu yerine Deney 6-12'nin uygulama notlarını içerir. |
| [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Sentetik DHIS2 tarzı özet veriler üzerinde bir halk sağlığı raporlama Agent'ının araç çağrılarını, hesaplama doğruluğunu, kanıt kullanımını ve dayanaksız iddialarını nesnel olarak değerlendirir. |
> `chapter7/android-world/` (tire ile yazılan) kıstas kodu değil, bilakis kitabın android_world üzerindeki T3A Agent başarısızlık vakaları hakkındaki analiz notlarıdır (`t3a*.md`); referans okuma materyali olarak kullanılabilir.
## Proje Türleri
| İkon | Tür | Anlamı |
| :--: | --- | --- |
| ✅ | **Bağımsız** | Bu depoda tam kod, API Key yapılandırıldıktan sonra çalışır |
| 📖 | **Yeniden Üretim Rehberi** | `git clone` ile **harici depolara** bağımlı ayrıntılı belge |
| 🚧 | **Tasarım Belgesi** | Yalnızca mimari/uygulama planı, çalıştırılabilir kod henüz hazır değil |
+49
View File
@@ -0,0 +1,49 @@
# Chương 7 · Đánh giá Agent
> biến biểu hiện của Agent thành tín hiệu có thể so sánh. Từ môi trường đánh giá, thiết kế bộ dữ liệu, hệ thống chỉ số, đến ý nghĩa thống kê, observability, chọn mô hình dựa trên đánh giá, cho tới đánh giá nội bộ và môi trường mô phỏng cấp sản xuất.
← [Về README chính](../docs/vi/README.md) · 📖 [Đọc nội dung chương](../book-vi/chapter7.vi.md)
## Cách đọc các thí nghiệm
Phần văn bản dùng skeleton cơ chế ngắn để giải thích luồng điều khiển; thư mục thí nghiệm chứa adapter SDK đầy đủ, log, kiểm thử và bằng chứng nghiệm thu. Không cần đọc từng tệp theo từng dòng.
- **Starter:** Bắt đầu từ mục tiêu, lệnh tối thiểu và điều kiện nghiệm thu; hãy bắt đầu với [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** Lần theo điểm vào, vòng lặp lõi, schema trạng thái/tin nhắn, công cụ và verifier.
- **Maintainer:** Sau đó đọc test, manifest bằng chứng, xử lý lỗi, đường rollback và adapter nhà cung cấp.
Lần đầu có thể bỏ qua credential, lớp trình bày và tương thích provider; quay lại khi cần tái tạo số liệu.
## Dự án đi kèm
| Thí nghiệm | Project | Type | Description |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | Tập trung đánh giá năng lực Agent dùng công cụ để suy luận phức tạp, bao gồm tính toán, tìm kiếm, xử lý dữ liệu và các ngữ cảnh khác. |
| 6-2 | `tau2-bench/` | 📖 | Hoàn thành thủ công các nhiệm vụ phân cấp của τ²-bench và ghi lại quỹ đạo. |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | Chạy rubric bốn mức trên 180 đánh giá có cấu trúc, kèm bằng chứng và quyền phủ quyết hallucination. |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Chạy 60 trường hợp trên ba hệ thống với hạch toán chi phí đầy đủ. |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | Chạy 11 trường hợp lỗi tiền tố quỹ đạo trên các biểu diễn bộ nhớ dạng JSON, Markdown và tương tự Python bằng lời gọi OpenRouter thực cùng các kiểm tra chính sách tất định. |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | Ma trận đầy đủ 4×3×2×60 giữ lại 1.440/1.440 quỹ đạo thực, không có lỗi hay lượt dùng chưa tính giá, kèm đủ chỉ số truy hồi/tác vụ, phân tích tương tác và trình xác minh độc lập đạt. |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | Chiến dịch chính thức trên một GPU đã hoàn thành 256 episode mỗi nhánh; chunk 1 đạt 0/256, chunk 25 đạt 26/256 và lưu hash của 512 rollout. |
| 6-2 | `terminal-bench/` | 📖 | Terminal-Bench là benchmark kiểm thử biểu hiện của AI Agent trong môi trường terminal thực. Từ biên dịch mã đến huấn luyện mô hình, thiết lập server, benchmark đánh giá cách Agent xử lý các nhiệm vụ đầu-cuối thực tế. Bao gồm bộ dữ liệu khoảng 100 nhiệm vụ và framework thực thi, hỗ trợ nhiều triển khai Agent. |
| 6-2 | `SWE-bench/` | 📖 | SWE-bench là benchmark đánh giá khả năng của mô hình ngôn ngữ lớn trong việc giải quyết các vấn đề GitHub thật. Với một codebase và mô tả issue, mô hình cần sinh patch có thể giải quyết vấn đề. Bao gồm nhiều phiên bản: SWE-bench, SWE-bench Lite, SWE-bench Verified và SWE-bench Multimodal. |
| 6-2 | `GAIA/` | 📖 | GAIA nhằm đánh giá thế hệ LLM tiếp theo (LLM có năng lực tăng cường bằng công cụ, prompt hiệu quả, truy cập tìm kiếm, v.v.). Bao gồm hơn 450 câu hỏi phi tầm thường cần mức độ công cụ và tự chủ khác nhau, với đáp án rõ ràng không mơ hồ. Chia thành 3 cấp độ khó. |
| 6-2 | `OSWorld/` | 📖 | Đánh giá năng lực của Agent khi thực thi nhiệm vụ phức tạp trong môi trường hệ điều hành đầy đủ, bao gồm quản lý file, thao tác ứng dụng và cấu hình hệ thống. |
| 6-2, 6-12 | `android_world/` | 📖 | Đánh giá biểu hiện của Agent trong môi trường di động Android, bao gồm điều hướng ứng dụng, tương tác UI và khả năng hoàn thành nhiệm vụ (repo benchmark ngoài). |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | Dùng nhiều cấu hình TTS (model/voice/speed khác nhau) để tổng hợp cùng một nhóm văn bản thử thách, sau đó dùng LLM-as-a-Judge đa phương thức chấm điểm từng chiều theo Rubric (độ rõ/naturalness, v.v.), tổng hợp thành bảng so sánh cấu hình có thể tái hiện. |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | Triển khai bảng xếp hạng hiệu năng Agent dựa trên hệ thống điểm ELO, đánh giá năng lực tương đối của các Agent khác nhau thông qua so sánh đối đầu. |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | So sánh GPT-5.6-sol và Claude Sonnet 5 tại thời điểm chuyển từ khám phá sang lần chỉnh sửa đầu tiên dưới cùng một Coding Harness trung lập; cả 18/18 ô đều hoàn tất không có lỗi API, và [manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) liên kết trajectory cùng bản tổng hợp bằng các hash có thể kiểm chứng. |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | Phân rã toàn tuyến chi phí của nhiệm vụ Agent nhiều vòng điển hình (hoàn tiền chăm sóc khách hàng): dùng tracing nhẹ tự xây để ghi lại token input/output/cache, độ trễ và chi phí của từng lần gọi LLM; tổng hợp “bước nào đắt nhất”, rồi dùng A/B để định lượng mức tiết kiệm thực của thiết kế thân thiện KV-cache + nén ngữ cảnh. |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | Benchmark ngang nhiều nhà cung cấp LLM API tương thích OpenAI; dùng giao diện streaming để đo chính xác độ trễ token đầu tiên (TTFT), đo các phân vị độ trễ đầu-cuối (p50/p95), throughput và tỷ lệ thành công dưới tải đồng thời. Một lệnh tạo bảng so sánh đa chiều, cho thấy chọn mô hình là đánh đổi nhiều chiều chứ không chỉ nhìn bảng xếp hạng. |
| 6-12 | [android-world](android-world/) | 📖 | Ghi chú phân tích báo cáo đánh giá T3A trên AndroidWorld trong repo này (điểm bắt đầu Thí nghiệm 6-12; không phải mã nguồn benchmark). |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | Sử dụng dữ liệu tổng hợp nhân tạo theo phong cách DHIS2 để đánh giá khách quan lời gọi công cụ, độ chính xác tính toán, trích dẫn bằng chứng và các tuyên bố không có căn cứ của Agent báo cáo y tế công cộng. |
> 📖 Các benchmark bên ngoài (tên đặt trong dấu backtick) cần tự clone riêng. [`android-world/`](android-world/) (có gạch nối) là **ghi chú phân tích đánh giá T3A** trong repo này (xem [README](android-world/README.md) của nó), không phải cùng đường dẫn với mã nguồn benchmark `android_world/` bên ngoài.
## Phân loại dự án
| Biểu tượng | Loại | Ý nghĩa |
| :--: | --- | --- |
| ✅ | **Chạy độc lập** | Có mã đầy đủ trong kho, chạy được sau khi cấu hình API Key |
| 📖 | **Hướng dẫn tái hiện** | Tài liệu chi tiết, cần `git clone` **kho ngoài** |
| 🚧 | **Tài liệu thiết kế** | Chỉ có kiến trúc/phương án, mã chạy được đang hoàn thiện |
+49
View File
@@ -0,0 +1,49 @@
# 第 7 章 · Agent 的評估
> 把表現變成可比較訊號:評估環境、指標、統計顯著性、評估驅動選型
← [返回主目錄](../docs/zh-TW/README.md) · 📖 [讀本章正文](../book/chapter7.md)
## 如何閱讀實驗
正文用短小的機制 skeleton 說明控制流;實驗目錄放完整的 SDK 適配、日誌、測試與驗收證據,不需要逐行讀完每個檔案。
- **Starter:** 先讀目標、最小指令與驗收條件;可從 [tau2-bench-eval](tau2-bench-eval/);
- **Builder:** 沿著入口、核心迴圈、狀態/訊息 schema、工具與驗證器閱讀。
- **Maintainer:** 最後再看測試、證據 manifest、失敗處理、回滾路徑與 provider adapter。
第一次閱讀可先跳過憑證載入、展示層和 provider 相容層;要重現數字時再回來查看。
## 配套專案
| 編號 | 專案 | 型別 | 一句話說明 |
| :--: | --- | :--: | --- |
| 6-1 | `tau2-bench/` | 📖 | 專注評估 Agent 使用工具進行複雜推理(計算、搜尋、資料處理)的能力 |
| 6-2 | `tau2-bench/` | 📖 | 人工完成 τ²-bench 分級任務並記錄軌跡。 |
| 6-3 | [user-memory-evaluation](../chapter3/user-memory-evaluation/) | ✅ | 四級 Rubric 已在 180 條結構化評判上執行,保留證據並設置幻覺一票否決。 |
| 6-4 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 在三個系統上執行 60 個案例,並完成成本核算。 |
| 6-5 | [user-memory-policy-eval](user-memory-policy-eval/) | ✅ | 以真實 OpenRouter 呼叫與確定性政策檢查,在 JSON、Markdown 與類 Python 記憶表示上執行 11 個 trajectory-prefix 錯誤案例。 |
| 6-11 | [user-memory-system-evaluation](user-memory-system-evaluation/) | ✅ | 完整 4×3×2×60 矩陣保留 1,440/1,440 條真實軌跡,無錯誤或未計價使用,並具備完整檢索/任務指標、互動分析與通過的獨立驗證。 |
| 6-13 | [openvla-robotwin2-eval](openvla-robotwin2-eval/) | ✅ | 單 GPU 正式實驗完成每個 action-chunk 組 256 回合;chunk 1 為 0/256、chunk 25 為 26/256,並保留 512 個 rollout 雜湊。 |
| 6-2 | `terminal-bench/` | 📖 | 測試 Agent 在真實終端機環境的端到端能力(編譯/訓練/部署),約 100 任務 + 執行框架 |
| 6-2 | `SWE-bench/` | 📖 | 評估 LLM 解決真實 GitHub 問題的能力,含 SWE-bench/Lite/Verified/Multimodal 多個版本 |
| 6-2 | `GAIA/` | 📖 | 評估下一代 LLM 的工具/搜尋/自主能力,450+ 個答案明確的非平凡問題,分 3 級難度 |
| 6-2 | `OSWorld/` | 📖 | 評估 Agent 在完整 OS 環境執行複雜任務的能力:檔案管理、應用操作、系統設定 |
| 6-2, 6-12 | `android_world/` | 📖 | 評估 Agent 在 Android 環境的應用導覽、UI 互動與任務完成能力(外部基準倉庫) |
| 6-6 | [tts-quality-eval](tts-quality-eval/) | ✅ | 多種 TTS 設定合成挑戰文字,LLM-as-a-Judge 按 Rubric 逐維度打分,輸出可復現對比表 |
| 6-7 | [elo-leaderboard](elo-leaderboard/) | ✅ | 基於 ELO 評分的 Agent 效能排行榜,透過對戰比較相對能力 |
| 6-8 | [model-action-threshold](model-action-threshold/) | ✅ | 在同一個中性的 Coding Harness 下,比較 GPT-5.6-sol 與 Claude Sonnet 5 從探索轉入首次修改的門檻;18/18 個單元均無 API 錯誤完成,[manifest](model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json) 以可驗證雜湊綁定軌跡與彙總 |
| 6-9 | [agent-cost-analysis](agent-cost-analysis/) | ✅ | 多輪 Agent 任務(客服退款)全鏈路成本拆解 + KV-cache 友善設計/上下文壓縮的 A/B 節省量化 |
| 6-10 | [model-benchmark](model-benchmark/) | 🚧 | 對多家 OpenAI 相容 API 橫向壓測 TTFT、p50/p95 延遲、吞吐與成功率,一條命令出對比表 |
| 6-12 | [android-world](android-world/) | 📖 | 本書對 T3A Agent 在 AndroidWorld 上的評估報告與失敗分析筆記(實驗 6-12 起點;非基準原始碼) |
| — | [public-health-reporting-eval](public-health-reporting-eval/) | ✅ | 基於合成 DHIS2 風格彙總資料,客觀評估公共衛生報告 Agent 的工具呼叫、計算準確性、證據引用與無依據聲明 |
> 📖 表中帶反引號的外部基準需自行克隆。[`android-world/`](android-world/)(連字號)是本倉庫內的 **T3A 評估分析筆記**(見該目錄 [README](android-world/README.md)),與外部 `android_world/` 基準原始碼不是同一路徑。
## 專案型別說明
| 圖示 | 型別 | 含義 |
| :--: | --- | --- |
| ✅ | **可獨立執行** | 本倉庫自帶完整程式碼,設定好 API Key 即可執行 |
| 📖 | **復現指南** | 依賴需自行 `git clone` 的**外部倉庫**(訓練框架、評測基準等) |
| 🚧 | **設計文件** | 僅包含架構與實現方案,可執行程式碼仍在完善中 |
+325
View File
@@ -0,0 +1,325 @@
# Agent End-to-End Cost Analysis / Agent 任务端到端成本分析(实验 7-9)
## English
This project performs a full cost decomposition for a multi-turn agent workflow (refund handling), including input/cache/output tokens, latency, and cost distribution. It enables practical measurement of where costs come from and how optimization strategies affect total spend.
### What it does
The benchmark runs a fixed 8-turn customer refund scenario and records every LLM call with a lightweight tracing layer:
- token usage (prompt, cached prompt, output)
- latency
- model cost by pricing table
It then reports:
- per-step cost breakdown
- cost component breakdown (non-cached input / cached input / output)
- p50/p95/p99 for per-step cost
- full 2×2 A/B comparison between optimization levers
The two levers are:
- **KV-cache friendliness**: keep a stable prefix to maximize cache hits
- **Context compression**: summarize long tool outputs for earlier turns
### Default model and API behavior
- Default model: `gpt-5.6-luna`.
- Preferred credentials: `OPENAI_API_KEY`, fallback to `OPENROUTER_API_KEY` (`gpt-*` remapped to `openai/*`).
- If OpenRouter keys exist for `gpt-5.x`, it is preferred due to authentication requirements.
- Offline mode is supported via `sample_trace.json` so all tables can be recomputed without API calls.
### Files
| File | Purpose |
|---|---|
| `config.py` | pricing model definitions and pricing presets |
| `tracer.py` | tracing helper and cost decomposition/aggregation |
| `agent.py` | 8-turn refund agent with `run_scenario(kv_cache, compress)` |
| `demo.py` | CLI entry for online/offline runs |
| `sample_trace.json` | captured 2×2 scenario token records for offline recomputation |
| `tests/` | offline pytest regressions for trace parsing and usage accounting |
| `requirements.txt` / `env.example` | dependencies and environment templates |
### Run
```bash
# From the repository root: use the shared Chapter 6 environment
uv sync --locked --python 3.12 --extra ch6
# Activate it before changing directories:
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# Windows cmd: .venv\Scripts\activate.bat
# pip fallback when uv is not installed:
# python -m pip install -e ".[ch6]"
cd chapter7/agent-cost-analysis
# Single-project compatibility path, still supported during migration:
# python -m pip install -r requirements.txt
export OPENAI_API_KEY=your-openai-api-key # or OPENROUTER_API_KEY=your-openrouter-api-key
python demo.py
python demo.py --offline --scenario all
```
### Tests
Automated tests are offline and do not require API keys.
```bash
# From the repository root, include the dev extra for pytest:
uv sync --locked --python 3.12 --extra ch6 --extra dev
# pip testing fallback:
# python -m pip install -e ".[ch6,dev]"
cd chapter7/agent-cost-analysis
python -m pytest tests
```
### CLI options
| Argument | Meaning |
|---|---|
| `--live` / `--offline` | call real model (default) / recompute from trace |
| `--scenario` | `ab` (naive+both), `all` (four scenarios), or subset list |
| `--trace` | trace file for offline mode |
| `--save-trace` | persist observed token usage from online runs |
| `--model` | model name for price preset |
| `--price-input` / `--price-cached` / `--price-output` | override per-million-token prices |
| `--no-warmup` | disable prefix warmup for KV-cache scenario |
| `--output` | export full result JSON |
### A/B scenarios (2×2)
| Scenario | KV-cache | Compression | Context design |
|---|---|---|---|
| `naive` | no | no | random session header + full tool returns |
| `kv` | yes | no | stable long prefix |
| `compress` | no | yes | only keep last 2 turns full; older turns summarized |
| `both` | yes | yes | stable prefix + compressed history |
The task logic is identical across scenarios so differences isolate optimization effects.
### Interpretation
Empirical results show:
- KV-cache can produce large improvements when prefixes are stable.
- Compression lowers prompt growth while keeping functional behavior.
- Joint optimization usually gives best total cost, though cache gains and compression gains are not simply additive.
### Offline recomputation
Offline mode reads `sample_trace.json` and re-runs only cost arithmetic, enabling:
- quick replication without keys
- quick “what-if” with different model prices
### Notes
- Observed numbers can vary due to real API behavior and cache timing.
- Prompt cache is best-effort and may miss in some turns.
- Tool-return token estimate uses tokenizer counts against current model encoder.
- Key precedence: prefer `OPENAI_API_KEY`; fallback is automatic via OpenRouter for supported paths.
---
## 中文
# 实验 7-9:Agent 任务的端到端成本分析
配套《深入理解 AI Agent》第 6 章「实验 7-9 ★:Agent 任务的端到端成本分析」。
对一个典型的多轮 Agent 任务(客服退款)做**全链路成本拆解**,用**自建的轻量 tracing / 可观测系统**记录每次 LLM 调用的输入/输出/缓存 token、时延与成本:按步骤聚合出「哪一步最贵」,按**成本构成**拆出「未缓存输入 / 缓存输入 / 输出各占多少、工具返回注入了多少 token」,并给出**单步成本分布(p50/p95/p99**;再做完整 **2×2 A/B 对比**,量化 **KV-cache 复用****上下文压缩** 两个杠杆各自及叠加后的真实成本差异。
- 默认模型 **`gpt-5.6-luna`**(当前廉价旗舰),通过 openai Python SDK 调用。首选 `OPENAI_API_KEY`;未设置时**自动回退到 `OPENROUTER_API_KEY`**(走 OpenRouter 兼容端点,`gpt-*` 映射为 `openai/*`)。由于 `gpt-5.x` 直连 OpenAI 需组织实名认证,只要存在 `OPENROUTER_API_KEY` 就优先走 OpenRouter。
- KV-cache 的节省是**真实**的:利用 OpenAI 的自动 prompt caching(前缀 ≥ 1024 token 且命中近期相同前缀时,`usage.prompt_tokens_details.cached_tokens > 0`,这部分输入按缓存价 5 折计费)。
- 提供**离线模式**:不打模型,读入一份此前真实运行录下的 token 用量(`sample_trace.json`canned token counts),用可配置单价**重新计算**成本/成本构成/A/B 对比表——无需 API key 即可复现全部表格,也可一键换算到其它模型定价。
## 文件
| 文件 | 说明 |
|------|------|
| `config.py` | 模型与价格:`Pricing` 单价对象 + 常见 OpenAI 模型单价预设,token→成本换算 |
| `tracer.py` | 自建轻量 tracing:包裹每次 LLM 调用记录 token/缓存/时延/成本;成本构成拆解、单步成本分布、按步骤拆解表;支持从录制用量离线复算(`from_records`|
| `agent.py` | 多轮客服退款 Agent 任务;`run_scenario(kv_cache, compress)` 把两个开关正交组合成 2×2 场景,并用 tiktoken 估算「工具返回注入」token |
| `demo.py` | 命令行入口(argparse):在线跑真实模型 / 离线复算;选择 A/B 场景、模型单价、输出文件 |
| `sample_trace.json` | 一次真实运行录下的四个场景逐步 token 用量(离线模式的输入,成本按当前单价重算)|
| `tests/` | 离线 pytest 回归测试:trace 解析与 usage 计费容错 |
| `requirements.txt` / `env.example` | 依赖与环境变量示例 |
## 运行
```bash
# 在仓库根目录使用统一的第 6 章环境
uv sync --locked --python 3.12 --extra ch6
# 切换目录前先激活环境:
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell.\.venv\Scripts\Activate.ps1
# Windows cmd.venv\Scripts\activate.bat
# 未安装 uv 时可用 pip 兜底:
# python -m pip install -e ".[ch6]"
cd chapter7/agent-cost-analysis
# 迁移期间仍支持单项目兼容路径:
# python -m pip install -r requirements.txt
# 在线(真实调用模型,需要 key):默认跑 A(朴素)+B(优化) 两组
export OPENAI_API_KEY=your-openai-api-key # 或 export OPENROUTER_API_KEY=your-openrouter-api-key(自动回退)
python demo.py
# 离线(无需 key):用内置 canned trace 复算全部表格
python demo.py --offline --scenario all
```
在线模式会真实调用 OpenAI`--scenario all` 约几十次 chat completion,运行一两分钟。
## 测试
自动化测试均为离线回归测试,不需要 API Key。
```bash
# 在仓库根目录安装 pytest 所需的 dev extra
uv sync --locked --python 3.12 --extra ch6 --extra dev
# pip 测试兜底路径:
# python -m pip install -e ".[ch6,dev]"
cd chapter7/agent-cost-analysis
python -m pytest tests
```
### 命令行参数(`python demo.py --help`
| 参数 | 说明 |
|------|------|
| `--live` / `--offline` | 在线真实调用(默认)/ 离线从 trace 文件复算(无需 key)|
| `--scenario NAME` | `ab`(默认=naive+both) / `all`(2×2 四组) / 逗号分隔子集 `naive,kv,compress,both` |
| `--trace FILE` | 离线读取的 canned trace,默认 `sample_trace.json` |
| `--save-trace FILE` | 在线跑时把真实 token 用量落盘,供之后 `--offline` 复算 |
| `--model NAME` | 模型名(决定默认单价预设:`gpt-4o-mini`/`gpt-4o`/`gpt-4.1-mini`/`gpt-4.1`|
| `--price-input/-cached/-output` | 覆盖三档单价(每百万 token 美元)|
| `--no-warmup` | 关闭 KV-cache 组的前缀预热(默认预热以稳定命中缓存)|
| `--output FILE` | 把成本拆解结果(含成本构成/分布/逐步用量)写成 JSON |
> **不改任何参数直接 `python demo.py`,行为与之前一致**:在线跑 A(朴素) 与 B(优化) 两组并打印拆解 + A/B 对比表。
## A/B 四种策略(完整 2×2)
同一个 8 轮客服退款任务(查订单 → 查物流 → 查退款政策 → 查知识库 → 风控 → 发起退款 → 通知 → 关单),四组做的是**同样的逻辑工作**,只在上下文构造上不同——因此成本差异纯粹来自「是否 KV-cache 友好」与「是否压缩上下文」两个正交开关:
| 场景 | KV-cache | 压缩 | 上下文构造 |
|------|:--:|:--:|------|
| `naive` A 朴素 | ✗ | ✗ | 每轮 system 前塞随机 session 头(破坏前缀)+ 历史工具返回原样带全 |
| `kv` 仅缓存 | ✓ | ✗ | 稳定长前缀(system 逐字节不变)+ 历史不压缩 |
| `compress` 仅压缩 | ✗ | ✓ | 前缀不稳定 + 仅最近 2 轮保留完整工具返回、更早压成一句话摘要 |
| `both` B 优化 | ✓ | ✓ | 稳定长前缀 + 上下文压缩(两个杠杆叠加)|
> 为聚焦「输入侧」两个杠杆,四组都用 `temperature=0` 且限制输出长度(`max_tokens=160`),让输出 token 成本近似为四组相等的固定项,避免模型生成长度的随机波动干扰对比。
>
> 工具环境是「受控」的(工具返回内容预设,真实系统里来自订单/物流/知识库后端),但**每一次 LLM 调用、每一份 token 用量、每一分成本都是真实打到 OpenAI 得到的**,保证可复现。
## 真实运行输出(gpt-4o-mini
以下为一次真实运行(`python demo.py --scenario all`)的输出,`sample_trace.json` 即由该次运行落盘、供 `--offline` 复现。
### (a) 单次任务成本拆解:按步骤 + 按成本构成 + 分布
```
===== 成本拆解: A 朴素(无缓存/无压缩)(单次任务全链路拆解) =====
步骤 工具/动作 输入tok 缓存tok 工具tok 输出tok 时延(s) 成本($)
---------------------------------------------------------------------------------------
turn-1 query_order 1113 0 276 104 3.15 0.000229
turn-2 query_logistics 1807 0 829 99 2.09 0.000330
turn-3 check_refund_policy 2154 0 1046 139 2.69 0.000406
turn-4 query_knowledge_base 2564 0 1287 160 2.92 0.000481
turn-5 query_user_history 2863 0 1389 136 2.69 0.000511
turn-6 issue_refund 3123 0 1490 160 3.07 0.000564
turn-7 send_notification 3408 0 1579 160 3.09 0.000607
turn-8 close_ticket 3668 0 1648 160 2.50 0.000646
---------------------------------------------------------------------------------------
合计 20700 0 9544 1118 22.20 0.003776
最贵的一步 → turn-8 / close_ticket: $0.000646(占总成本 17.1%
成本构成:
未缓存输入 20700 tok $0.003105 (82.2%)
缓存输入 0 tok $0.000000 (0.0%)
输出 1118 tok $0.000671 (17.8%)
其中「工具返回注入」累计输入 9544 tok (同一份工具返回在后续每轮被反复计费)
单步成本分布(n=8): 均值 $0.000472 p50 $0.000481 p95 $0.000646 p99 $0.000646
===== 成本拆解: B 优化(KV缓存+压缩)(单次任务全链路拆解) =====
步骤 工具/动作 输入tok 缓存tok 工具tok 输出tok 时延(s) 成本($)
---------------------------------------------------------------------------------------
turn-1 query_order 1056 1024 276 139 2.41 0.000165
turn-2 query_logistics 1781 0 829 112 2.18 0.000334
turn-3 check_refund_policy 2143 1024 1046 151 2.56 0.000335
turn-4 query_knowledge_base 2310 0 1052 160 3.03 0.000442
turn-5 query_user_history 2060 1024 635 160 2.48 0.000328
turn-6 issue_refund 2143 1024 551 160 2.71 0.000341
turn-7 send_notification 2188 1024 430 160 2.79 0.000347
turn-8 close_ticket 2354 1024 429 122 2.38 0.000349
---------------------------------------------------------------------------------------
合计 16035 6144 5248 1164 20.54 0.002643
最贵的一步 → turn-4 / query_knowledge_base: $0.000442(占总成本 16.7%
成本构成:
未缓存输入 9891 tok $0.001484 (56.1%)
缓存输入 6144 tok $0.000461 (17.4%)
输出 1164 tok $0.000698 (26.4%)
其中「工具返回注入」累计输入 5248 tok (同一份工具返回在后续每轮被反复计费)
单步成本分布(n=8): 均值 $0.000330 p50 $0.000335 p95 $0.000442 p99 $0.000442
```
可以看到:朴素组 A 的输入 token 随轮次从 1113 一路涨到 3668(上下文累积效应,最后一步最贵),「工具返回注入」累计吃掉 9544 输入 token;优化组 B 的输入 token 被压缩策略压住(末轮 2354 而非 3668,工具注入累计降到 5248),且多数轮次持续命中 1024 缓存 token,缓存输入把整块费用打了 5 折。
### (b) 完整 2×2 A/B 对比
```
===== A/B 成本对比(同一个 8 轮客服退款任务)=====
方案 总输入tok 缓存tok 缓存率 输出tok 总成本($) vs基线
------------------------------------------------------------------------------------------
A 朴素(无缓存/无压缩) 20700 0 0.0% 1118 0.003776 基线
KV 仅缓存(稳定前缀/不压缩) 20386 13568 66.6% 1112 0.002707 -28.3%
仅压缩(前缀不稳定/摘要) 16177 0 0.0% 1147 0.003115 -17.5%
B 优化(KV缓存+压缩) 16035 6144 38.3% 1164 0.002643 -30.0%
------------------------------------------------------------------------------------------
重点对比:A 朴素(无缓存/无压缩) → B 优化(KV缓存+压缩)
总 token: A=21818 → B=17199 减少 4619 (21.2%)
缓存 token: A=0 → B=6144 B 靠稳定前缀命中缓存)
总成本: A=$0.003776 → B=$0.002643 降低 $0.001133 (30.0%)
成本倍率: A 是 B 的 1.43 倍
```
结论(读出两个杠杆各自与叠加的贡献):
- **仅 KV-cache**:不改上下文长度,只靠稳定前缀让重复的系统提示/工具定义/历史轮次按缓存价计费,缓存率冲到 66.6%,端到端成本就降了 **28.3%**——本例中它是单个最有效的杠杆。
- **仅压缩**:把旧轮次工具返回压成摘要,总输入 token 从 20700 降到 16177(约 −22%),端到端成本降 **17.5%**;但因为前缀不稳定,缓存率为 0。
- **两者叠加(B 优化)**:输入 token 最少、又能命中缓存,端到端总成本降 **30.0%**(A 是 B 的 1.43 倍)。单看输入侧成本,降幅落在书中「KV Cache 可降低 30%-60% 输入 token 成本」的经验区间内。
> 注意 KV 与压缩并非简单相加:压缩缩短了历史,可被缓存的「历史轮次」也随之变少,所以 B 的缓存 token(6144)反而少于「仅 KV」(13568)。这正是评估的价值——两个优化叠加时要实测协同效应,而不是把各自的收益直接相加。
## 离线复算(无需 API key
`sample_trace.json` 存的是上面这次真实运行的**逐步 token 用量(实测值)**;离线模式只做「按单价重算成本」这一步纯离线数学,因此无需 key 即可复现全部表格,还能一键换算到其它模型定价:
```bash
python demo.py --offline --scenario all # 用 gpt-4o-mini 单价复算
python demo.py --offline --scenario all --model gpt-4o # 同一份 token 用量,换 gpt-4o 单价重算
python demo.py --offline --price-input 0.20 --price-cached 0.10 --price-output 0.80
```
`gpt-4o` 单价后,同一批 token 的四组总成本等比放大(结论与占比不变,A=$0.062930 → B=$0.044047,仍是 30.0%)——这说明**成本优化的相对收益由 token 结构决定,与绝对单价无关**。
## 注意事项
- 具体数字每次在线运行会有小幅波动:OpenAI 的 prompt cache 是**尽力而为**的(按约 128 token 的块缓存、约 5–10 分钟过期、偶尔未命中),因此某些轮次可能出现 `cached_tokens=0`(如上面 B 组的 turn-2/turn-4)。`demo.py` 在正式计量前会先对 KV-cache 组跑一次「预热」把稳定前缀写入缓存,让命中更稳定;`--no-warmup` 可关闭。
- 价格写在 `config.py``PRICING_PRESETS`),默认是 gpt-4o-mini 的公开单价(输入 \$0.15 / 缓存 \$0.075 / 输出 \$0.60 每百万 token)。**换模型**:用 `--model`(如 `--model gpt-4o`)或 `--price-*` 直接覆盖单价即可;缓存命中要求稳定前缀 ≥ 1024 token,换更强模型不影响该机制。
- 「工具返回注入」token 用 tiktoken 按当前模型的编码器离线估算(统计每轮输入里工具返回文本占多少 token),用于回答书中「一次工具返回可能占 2000-5000 token,且在后续每轮被反复计费」这一放大因素。
- 凭据:首选 `OPENAI_API_KEY`;未设置时自动回退到 `OPENROUTER_API_KEY`(走 OpenRouter`gpt-*` 映射为 `openai/*`)。`gpt-5.x` 直连需组织实名认证,故有 `OPENROUTER_API_KEY` 时优先走 OpenRouter。离线复算(`--offline`)无需任何 key。
+273
View File
@@ -0,0 +1,273 @@
"""
一个多轮「客服退款 Agent」任务,用于成本分析(对应书 6.x 表6-4 的客服退款示例)。
为了让实验可复现、不依赖模型工具调用的随机性,这里用「受控工具环境」:
每一轮我们把上一步工具的返回结果喂给模型,由模型(真实 LLM 调用)决定下一步怎么做。
工具返回内容是预设好的(真实 API 里会是订单系统/物流系统的返回),
但每一次 LLM 调用、每一份 token 用量、每一分成本都是真实的。
本文件把「是否 KV-cache 友好」和「是否压缩上下文」两个开关正交拆开,
可组合出完整的 2×2 A/B(对应书中「对比启用/禁用 KV Cache、启用/禁用上下文压缩」):
run_scenario(kv_cache=False, compress=False) —— A 朴素(前缀不稳定 + 不压缩)
run_scenario(kv_cache=True, compress=False) —— 仅 KV-cache(稳定长前缀,历史不压缩)
run_scenario(kv_cache=False, compress=True) —— 仅压缩(前缀不稳定,旧轮次摘要)
run_scenario(kv_cache=True, compress=True) —— B 优化(稳定前缀 + 压缩,两个杠杆叠加)
兼容旧接口:run_naive == (False, False)run_optimized == (True, True)。
"""
import uuid
from functools import lru_cache
from config import MODEL, Pricing
from tracer import Tracer
# 最近保留几轮完整工具返回(更早的压成一句话摘要)。压缩关闭时视为无穷大。
KEEP_VERBOSE = 2
# 限制每轮输出长度:本实验聚焦「输入侧」的 KV-cache 与压缩两个杠杆,
# 把输出 token 控制在相近水平,可避免模型生成长度的随机波动干扰 A/B 成本对比。
MAX_OUTPUT_TOKENS = 160
# ---------------------------------------------------------------------------
# 一个「足够长且稳定」的系统提示 + 工具定义(> 1024 token),
# 这是 KV-cache 命中的关键:稳定的长前缀才会被 OpenAI 自动缓存。
# 内容是一个真实感的客服退款 Agent 的系统规范与工具手册。
# ---------------------------------------------------------------------------
STABLE_SYSTEM_PROMPT = """你是「云购商城」的高级客服 Agent,专门处理售后与退款事务。你必须严格遵循以下工作规范。
# 角色与目标
你的目标是在保障平台规则的前提下,高效、礼貌地帮助用户完成退款、退货、换货、物流查询等售后诉求。
你要主动澄清诉求、核对订单状态、判断是否符合退款政策,并在权限范围内执行操作。
# 可用工具手册(tool manual
1. query_order(order_id): 查询订单详情。返回字段包括:order_id, status, item_name, sku, price,
quantity, pay_time, pay_channel, buyer_note, seller_note, is_prepaid, warehouse, promotion_tags。
2. query_logistics(order_id): 查询物流轨迹。返回字段:carrier, tracking_no, current_status,
last_scan_time, last_scan_location, estimated_delivery, full_trace(数组,含每个扫描节点)。
3. check_refund_policy(sku, reason): 查询该 SKU 在给定退款原因下的退款政策。返回字段:
refundable, need_return, restocking_fee_rate, refund_window_days, special_notes, approval_required。
4. query_user_history(user_id): 查询用户历史行为,用于风控。返回:total_orders, refund_count_90d,
dispute_count, risk_level, vip_tier, register_days。
5. issue_refund(order_id, amount, reason): 发起退款。返回:refund_id, status, expected_arrival,
channel, operator。仅当政策允许且金额不超过订单实付金额时才可调用。
6. send_notification(user_id, channel, template, params): 给用户发通知(sms/app/email)。
# 决策规范
- 先核对订单是否存在、状态是否允许退款(已发货未签收、已签收 7 天内、未发货均有不同处理路径)。
- 未发货:可直接全额退款,无需退货。
- 已发货未签收:需拦截物流或等待退回,退款在退货签收后发起。
- 已签收 7 天内且商品无质量问题:适用 7 天无理由,可能收取一定比例的手续费(restocking_fee_rate)。
- 商品质量问题:全额退款且不收手续费,需用户提供凭证。
- 涉及大额退款(> 500 元)或高风险用户(risk_level=high)需要人工审批(approval_required=true)。
- 每一步都要给出简短的中文推理,说明你「基于什么信息、决定下一步调用哪个工具或给出什么结论」。
# 输出要求
- 保持专业、简洁、有同理心。
- 每轮只推进一步,不要臆造工具尚未返回的数据。
- 最终解决时,明确告知用户退款金额、到账时间与后续动作。
# 合规与风控
- 不得泄露其它用户信息;不得承诺超出政策的赔付;金额与政策以工具返回为准。
- 对疑似欺诈(短期高频退款、异常物流轨迹)要保持谨慎并触发人工审批。
请始终遵守以上全部规范。"""
# ---------------------------------------------------------------------------
# 预设的多轮剧本:用户诉求 + 每一步工具返回(真实 API 里来自后端系统)。
# 工具返回故意写得比较「啰嗦」(大 JSON),以体现工具结果注入上下文的 token 成本。
# ---------------------------------------------------------------------------
USER_REQUEST = (
"你好,我上周买的蓝牙耳机(订单号 ORD20240517001)到货后一直连不上,"
"试了各种办法都没用,我想退货退款,怎么处理?"
)
# 每一轮:(逻辑步骤名, 关联工具, 该工具的"啰嗦"返回文本)
# 工具返回都写得比较大(真实的订单/物流/知识库返回往往几百到上千 token),
# 以体现「工具结果注入上下文后在后续每一轮被反复计费」这一放大因素。
_LOGISTICS_TRACE = ",".join(
'{"time":"2024-05-%02dT%02d:%02d","loc":"%s","desc":"%s","operator":"SF%04d","scan_type":"auto"}'
% (17 + i // 6, 6 + i, (i * 7) % 60, loc, desc, 1000 + i)
for i, (loc, desc) in enumerate([
("华东1仓", "包裹已揽收,称重0.42kg"), ("华东1仓分拣中心", "已分拣,发往上海转运"),
("上海转运中心", "到达转运中心"), ("上海转运中心", "已发出,运输中"),
("苏州中转场", "途经中转"), ("上海浦东集散点", "到达派送网点"),
("上海浦东集散点", "安排派送"), ("浦东xx营业点", "派送中,联系收件人"),
("浦东xx营业点", "首次派送未接通"), ("浦东xx营业点", "二次派送"),
("浦东xx营业点", "已签收,签收人:本人"),
])
)
TOOL_RESULTS = [
("turn-1", "query_order",
'{"order_id":"ORD20240517001","status":"SIGNED","item_name":"Acme 主动降噪蓝牙耳机 Pro",'
'"sku":"SKU-BT-9981","price":499.00,"quantity":1,"pay_time":"2024-05-17T10:22:31",'
'"pay_channel":"wechat_pay","buyer_note":"希望尽快发货,送人用","seller_note":"已核对库存",'
'"is_prepaid":true,"warehouse":"华东1仓","promotion_tags":["满300减30","会员日","新客礼"],'
'"actual_paid":469.00,"coupon_used":"CPN-30","points_earned":469,"invoice_requested":false,'
'"sign_time":"2024-05-19T14:03:11","after_sale_window_end":"2024-05-26T23:59:59",'
'"sub_items":[{"sku":"SKU-BT-9981","name":"耳机主体","qty":1},{"sku":"SKU-BT-9981-CASE","name":"充电盒","qty":1},{"sku":"SKU-BT-9981-TIP","name":"耳塞套装","qty":1}],'
'"address_hash":"a1b2c3d4","channel":"app","device":"iOS","order_source":"首页推荐位"}'),
("turn-2", "query_logistics",
'{"carrier":"顺丰速运","tracking_no":"SF1234567890123","current_status":"已签收",'
'"last_scan_time":"2024-05-19T14:03:11","last_scan_location":"上海市浦东新区xx营业点",'
'"estimated_delivery":"2024-05-19","weight_kg":0.42,"volume":"20x15x8cm","insured":true,'
'"full_trace":[' + _LOGISTICS_TRACE + ']}'),
("turn-3", "check_refund_policy",
'{"sku":"SKU-BT-9981","reason":"quality_issue_cannot_connect","refundable":true,'
'"need_return":true,"restocking_fee_rate":0.0,"refund_window_days":7,'
'"special_notes":"质量问题类退款免手续费;需用户回寄并由质检确认是否为质量问题;'
'若质检判定非质量问题(如人为损坏、私自拆修),将按原路退回商品且不予退款;'
'回寄运费由平台承担,用户需在系统中申请电子面单;退款在质检通过后 1 个工作日内发起;'
'3C 电子类目已激活/绑定账号的商品,需先解绑再回寄,否则质检不予通过。",'
'"approval_required":false,"category":"3C-电子","quality_claim_supported":true,'
'"return_label_provided":true,"qc_sla_days":2,"related_policy_ids":["P-3C-01","P-3C-07","P-QC-12"]}'),
("turn-4", "query_knowledge_base",
'{"query":"蓝牙耳机无法连接 排查","hits":['
'{"kb_id":"KB-1001","title":"蓝牙耳机无法连接的常见原因","content":"1.未进入配对模式;'
'2.手机蓝牙缓存异常需忘记设备重连;3.固件版本过低;4.电量过低;5.多设备抢占连接。"},'
'{"kb_id":"KB-1002","title":"Acme Pro 系列重置方法","content":"长按充电盒按键15秒至指示灯红白交替闪烁即完成重置,'
'随后在手机端删除旧配对记录重新搜索。若重置后仍无法搜索到设备,多为硬件故障,建议走质量问题退换。"},'
'{"kb_id":"KB-1003","title":"质量问题判定标准","content":"重置无效 + 换设备仍无法连接 + 无进液/外观损伤,'
'通常判定为质量问题,支持免费退换。"}],"suggested_action":"引导用户重置;若无效则判定质量问题走退款流程"}'),
("turn-5", "query_user_history",
'{"user_id":"U-88123","total_orders":37,"refund_count_90d":1,"dispute_count":0,'
'"risk_level":"low","vip_tier":"gold","register_days":1180,"payment_disputes":0,'
'"avg_order_value":312.5,"last_refund_reason":"尺码不合适","chargeback_count":0,'
'"complaint_count":0,"account_status":"normal","fraud_flags":[],"lifetime_value":11562.5}'),
("turn-6", "issue_refund",
'{"refund_id":"RF20240520777","status":"APPROVED","amount":469.00,'
'"expected_arrival":"1-3 个工作日","channel":"原路退回-微信","operator":"agent-bot",'
'"return_shipping":"平台承担","return_address":"华东1仓退货组","return_label":"SF-RET-998877",'
'"qc_required":true,"qc_deadline":"2024-05-27","refund_flow":"pending_return->qc->refund"}'),
("turn-7", "send_notification",
'{"user_id":"U-88123","channel":"app","template":"refund_approved",'
'"delivered":true,"message_id":"MSG-556677","sent_time":"2024-05-20T15:20:03",'
'"params":{"refund_id":"RF20240520777","amount":469.00,"return_label":"SF-RET-998877"},'
'"read_receipt":false,"fallback_sms_scheduled":true}'),
("turn-8", "close_ticket",
'{"ticket_id":"TK-20240520-3345","status":"resolved","resolution":"refund_after_return",'
'"csat_survey_sent":true,"handle_time_s":184,"escalated":false,"agent":"agent-bot",'
'"summary_logged":true,"tags":["退款","质量问题","3C","已闭环"]}'),
]
# 供「上下文压缩」策略使用的旧轮次一句话摘要(把啰嗦的工具返回压成要点)
TOOL_SUMMARIES = {
"turn-1": "[摘要] 订单 ORD20240517001Acme降噪耳机Pro,实付469元,已于5/19签收,售后窗口至5/26。",
"turn-2": "[摘要] 物流:顺丰已签收(5/19 14:03,本人签收),11 个轨迹节点均正常无异常。",
"turn-3": "[摘要] 退款政策:质量问题可退、免手续费,需回寄质检,回寄运费平台承担,无需人工审批。",
"turn-4": "[摘要] 知识库:先引导重置耳机;重置无效即判定质量问题,支持免费退换。",
"turn-5": "[摘要] 用户风控:37单/90天仅1次退款/low风险/gold会员,信誉良好,无欺诈标记。",
"turn-6": "[摘要] 已发起退款 RF20240520777:469元原路退微信,需回寄质检,平台承担回寄运费。",
"turn-7": "[摘要] 已通过 app 通知用户退款已批准,附回寄面单。",
}
def _next_user_msg(tool_name: str, tool_result: str) -> str:
"""把工具返回包装成喂给模型的下一条 user 消息。"""
return (
f"[工具 {tool_name} 返回结果]\n{tool_result}\n\n"
f"请基于以上结果给出你的推理,并决定下一步动作。"
)
@lru_cache(maxsize=1)
def _encoder():
"""按当前模型取 tiktoken 编码器(离线可用),用于估算「工具返回注入」的 token。"""
import tiktoken
try:
return tiktoken.encoding_for_model(MODEL)
except Exception:
return tiktoken.get_encoding("cl100k_base")
def _ntok(text: str) -> int:
return len(_encoder().encode(text))
# ---------------------------------------------------------------------------
# 四种 A/B 场景的登记表:名字 + 两个开关。
# ---------------------------------------------------------------------------
SCENARIOS = {
"naive": ("A 朴素(无缓存/无压缩)", False, False),
"kv": ("KV 仅缓存(稳定前缀/不压缩)", True, False),
"compress": ("仅压缩(前缀不稳定/摘要)", False, True),
"both": ("B 优化(KV缓存+压缩)", True, True),
}
def build_messages(idx, step, tool, result, turns, kv_cache, compress):
"""构造第 idx 轮要发给模型的 messages,并返回本轮输入里「工具返回注入」的累计 token。
kv_cache=True → system 用逐字节稳定的长前缀(可被 OpenAI 自动缓存);
kv_cache=False → 每轮在 system 最前面塞随机 session 头,破坏前缀一致性。
compress=True → 仅最近 KEEP_VERBOSE 轮保留完整工具返回,更早轮次压成一句话摘要。
"""
if kv_cache:
system = {"role": "system", "content": STABLE_SYSTEM_PROMPT}
else:
volatile_head = f"[会话追踪] session={uuid.uuid4()} 请求序号={uuid.uuid4()}\n\n"
system = {"role": "system", "content": volatile_head + STABLE_SYSTEM_PROMPT}
history = [{"role": "user", "content": USER_REQUEST}]
tool_ctx_tokens = 0
for j, (p_step, p_assistant, p_tool, p_result) in enumerate(turns):
history.append({"role": "assistant", "content": p_assistant})
if compress and idx - j > KEEP_VERBOSE:
compact = TOOL_SUMMARIES.get(p_step, f"[摘要] {p_tool} 已完成。")
history.append({"role": "user", "content": compact})
tool_ctx_tokens += _ntok(compact)
else:
history.append({"role": "user", "content": _next_user_msg(p_tool, p_result)})
tool_ctx_tokens += _ntok(p_result)
messages = [system] + history + [
{"role": "user", "content": _next_user_msg(tool, result)}
]
tool_ctx_tokens += _ntok(result) # 本轮新注入的工具返回
return messages, tool_ctx_tokens
def run_scenario(client, kv_cache: bool, compress: bool, name: str = None,
pricing: Pricing = None) -> Tracer:
"""跑一遍 8 轮客服退款任务,两个开关正交组合出 2×2 中的一格。
两组做的是同样的逻辑工作,只在上下文构造上不同,因此成本差异纯粹来自
KV-cache 复用与上下文压缩这两个输入侧杠杆。
"""
tracer = Tracer(client, name=name or f"kv={kv_cache},compress={compress}",
pricing=pricing)
turns = []
for idx, (step, tool, result) in enumerate(TOOL_RESULTS):
messages, tool_ctx = build_messages(
idx, step, tool, result, turns, kv_cache, compress)
model_name = MODEL.lower()
if model_name.startswith("kimi-k2.5"):
temperature = 0.6
elif any(tag in model_name for tag in ("kimi-k3", "gpt-5")):
temperature = 1
else:
temperature = 0
request = {
"model": MODEL,
"messages": messages,
"temperature": temperature,
"max_tokens": MAX_OUTPUT_TOKENS,
}
if MODEL.lower().startswith("kimi-k2.5"):
request["extra_body"] = {"thinking": {"type": "disabled"}}
resp = tracer.chat(step=step, tool=tool, tool_ctx_tokens=tool_ctx, **request)
assistant_text = resp.choices[0].message.content or ""
turns.append((step, assistant_text, tool, result))
return tracer
def run_naive(client, pricing: Pricing = None) -> Tracer:
"""(a) 朴素做法:前缀不稳定 + 不压缩历史(KV-cache 命中不了、上下文疯长)。"""
return run_scenario(client, kv_cache=False, compress=False,
name=SCENARIOS["naive"][0], pricing=pricing)
def run_optimized(client, pricing: Pricing = None) -> Tracer:
"""(b) KV-cache 友好 + 上下文压缩:稳定长前缀命中缓存 + 旧轮次摘要。"""
return run_scenario(client, kv_cache=True, compress=True,
name=SCENARIOS["both"][0], pricing=pricing)
+129
View File
@@ -0,0 +1,129 @@
"""
全局配置:模型与价格。
价格换算成本时使用「每百万 token 单价(美元)」。
默认值取自 OpenAI gpt-4o-mini 的公开定价(2024-2025):
- 输入 : $0.15 / 1M tokens
- 缓存命中输入 : $0.075 / 1M tokens (命中 prompt cache 的输入按 5 折计费)
- 输出 : $0.60 / 1M tokens
注意:
1. 默认模型为 gpt-5.6-luna(当前廉价旗舰)。首选凭据是 OPENAI_API_KEY;若未设置,
自动回退到 OPENROUTER_API_KEY 并把模型名映射成 OpenRouter idgpt-* -> openai/*)。
由于 gpt-5.x 直连 OpenAI 需要组织实名认证,只要 OPENROUTER_API_KEY 存在就优先走
OpenRouter(见 make_client_and_model)。仍可用 COST_DEMO_MODEL / --model 切换任意模型。
2. OpenAI 的 prompt caching 是「自动」的:当请求前缀 >= 1024 token 且与近期请求
命中相同前缀时,usage.prompt_tokens_details.cached_tokens 会大于 0
这部分 token 按缓存价(更便宜)计费。本项目正是用它来真实体现 KV-cache 的节省。
OpenRouter 转发 OpenAI 时同样在 prompt_tokens_details.cached_tokens 回传缓存命中。)
"""
import os
from dataclasses import dataclass
from dotenv import load_dotenv
load_dotenv()
# 使用的模型(默认当前廉价旗舰 gpt-5.6-luna;可用 COST_DEMO_MODEL / --model 覆盖)
MODEL = os.environ.get("COST_DEMO_MODEL", "gpt-5.6-luna")
# OpenRouter 回退:无 OPENAI_API_KEY 时用 OPENROUTER_API_KEY 走 OpenAI 兼容端点。
OPENROUTER_BASE_URL = "https://openrouter.ai/api/v1"
def _to_openrouter_model(model: str) -> str:
"""把模型名映射成 OpenRouter id:含 '/' 视为原生 idgpt-* -> openai/*
claude-* -> anthropic/claude-opus-4.8;其余回退到 openai/gpt-5.6-luna。"""
if "/" in model:
return model
if model.startswith("gpt-"):
return "openai/" + model
if model.startswith("claude-"):
return "anthropic/claude-opus-4.8"
return "openai/gpt-5.6-luna"
def make_client_and_model(model: str):
"""构造 OpenAI 兼容 client 并返回 (client, 实际调用的模型名)。
回退策略(universal OpenRouter fallback):
- gpt-5.x 且存在 OPENROUTER_API_KEY -> 优先走 OpenRouter(直连需组织实名认证);
- 否则有 OPENAI_API_KEY -> 直连 OpenAI,模型名不变;
- 否则有 OPENROUTER_API_KEY -> 走 OpenRouter,模型名按 _to_openrouter_model 映射;
- 两者皆无 -> 抛出清晰错误。
"""
from openai import OpenAI
primary = os.environ.get("OPENAI_API_KEY", "").strip()
orkey = os.environ.get("OPENROUTER_API_KEY", "").strip()
prefer_openrouter = bool(orkey) and model.startswith("gpt-5")
if not prefer_openrouter and primary:
return OpenAI(timeout=60.0, max_retries=2), model
if orkey:
return (
OpenAI(base_url=OPENROUTER_BASE_URL, api_key=orkey,
timeout=60.0, max_retries=2),
_to_openrouter_model(model),
)
if primary:
return OpenAI(timeout=60.0, max_retries=2), model
raise RuntimeError(
"缺少可用凭据:请设置 OPENAI_API_KEY(直连 OpenAI),或设置 "
"OPENROUTER_API_KEY(自动回退到 OpenRouter);或改用 --offline 离线复算(无需 key)。"
)
# 每百万 token 的美元单价(默认 gpt-4o-mini
PRICE_INPUT_PER_M = 0.15 # 普通输入
PRICE_CACHED_PER_M = 0.075 # 命中缓存的输入(gpt-4o-mini 缓存读取为输入价的 50%)
PRICE_OUTPUT_PER_M = 0.60 # 输出
@dataclass(frozen=True)
class Pricing:
"""一组每百万 token 的美元单价。"""
input_per_m: float
cached_per_m: float
output_per_m: float
def cost_usd(self, prompt_tokens: int, cached_tokens: int,
completion_tokens: int) -> float:
"""按 token 用量换算成本(美元)。
prompt_tokens : usage.prompt_tokens,包含了缓存命中的部分
cached_tokens : usage.prompt_tokens_details.cached_tokens,命中缓存的输入 token
completion_tokens: usage.completion_tokens
未命中缓存的输入 = prompt_tokens - cached_tokens,按普通输入价计费;
命中缓存的输入按缓存价计费。
"""
uncached_input = max(prompt_tokens - cached_tokens, 0)
return (
uncached_input / 1_000_000 * self.input_per_m
+ cached_tokens / 1_000_000 * self.cached_per_m
+ completion_tokens / 1_000_000 * self.output_per_m
)
# 常见 OpenAI 模型的公开单价预设(每百万 token,美元),方便 CLI 用 --model 一键切换。
# 换更强的模型不影响 KV-cache 机制(仍要求稳定前缀 >= 1024 token)。
PRICING_PRESETS = {
"gpt-4o-mini": Pricing(0.15, 0.075, 0.60),
"gpt-4o": Pricing(2.50, 1.25, 10.00),
"gpt-4.1-mini": Pricing(0.40, 0.10, 1.60),
"gpt-4.1": Pricing(2.00, 0.50, 8.00),
}
def default_pricing() -> Pricing:
"""返回默认模型(config 中 MODEL)的单价;未知模型回退到模块级 PRICE_* 默认值。"""
return PRICING_PRESETS.get(
MODEL, Pricing(PRICE_INPUT_PER_M, PRICE_CACHED_PER_M, PRICE_OUTPUT_PER_M)
)
def cost_usd(prompt_tokens: int, cached_tokens: int, completion_tokens: int,
pricing: "Pricing | None" = None) -> float:
"""按 token 用量换算成本(美元)。默认用模块级单价,可传入自定义 Pricing。"""
p = pricing or Pricing(PRICE_INPUT_PER_M, PRICE_CACHED_PER_M, PRICE_OUTPUT_PER_M)
return p.cost_usd(prompt_tokens, cached_tokens, completion_tokens)
@@ -0,0 +1,452 @@
"""
Agent trajectory cost-efficiency analyzer (实验 7-9 成本效率分析).
Builds on the span/trace model from ``tracer.py``: an agent task is a sequence
of turns, each turn carrying token usage (prompt / cached / completion), tool
context tokens, and latency. This module turns a recorded trajectory into an
:class:`EfficiencyReport` — per-turn metrics, a single efficiency score, and
actionable recommendations (wasteful turns, compression opportunities, cache
miss patterns).
It is fully offline: it never calls a model. Pricing is configured per million
tokens (same convention as ``config.Pricing``) and defaults to gpt-4o-mini.
Two trajectory shapes are accepted:
1. A bare list of turn dicts (the spans of one scenario).
2. A trace dict as written by the tracer — ``{"turns": [...]}``,
``{"spans": [...]}``, or ``{"scenarios": [{"spans": [...]}, ...]}`` (the
first scenario with spans is analyzed). A top-level ``"pricing"`` key is
honoured when no explicit pricing was given to the constructor.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from typing import Any
# --------------------------------------------------------------------------- #
# Data shapes
# --------------------------------------------------------------------------- #
@dataclass
class TurnMetrics:
"""Per-turn cost-efficiency metrics."""
turn_id: int
input_tokens: int
output_tokens: int
cache_hit_ratio: float
cost_usd: float
latency_ms: float
tool_calls: int
classification: str # productive / wasteful / cached / expensive
@property
def total_tokens(self) -> int:
return self.input_tokens + self.output_tokens
@dataclass
class EfficiencyReport:
"""Aggregate cost-efficiency report for a whole trajectory."""
total_turns: int
total_cost_usd: float
total_tokens: int
efficiency_score: float
turn_metrics: list[TurnMetrics]
recommendations: list[str]
# Derived aggregate metrics (computed by analyze_trajectory).
cumulative_costs: list[float] = field(default_factory=list)
tokens_per_tool_call: float = 0.0
latency_per_turn: float = 0.0
# --------------------------------------------------------------------------- #
# Analyzer
# --------------------------------------------------------------------------- #
_STEP_RE = re.compile(r"turn[-_ ]]?(\d+)", re.IGNORECASE)
class CostEfficiencyAnalyzer:
"""Analyze the cost-efficiency of a recorded agent trajectory.
Parameters
----------
pricing:
Per-million-token USD prices with keys ``input``, ``output`` and
``cached``. ``None`` falls back to :meth:`default_pricing` (and to a
``pricing`` block embedded in the trajectory, if present).
wasteful_token_threshold:
A turn with no tool calls and at least this many total tokens is
classified ``wasteful``.
expensive_cost_threshold:
Per-turn cost (USD) above which a turn is ``expensive``. ``None`` means
relative: a turn is expensive when its cost exceeds 1.5x the mean
per-turn cost of the trajectory (computed in :meth:`analyze_trajectory`;
:meth:`analyze_turn` alone treats ``None`` as "never expensive").
cached_ratio_threshold:
Cache hit ratio at or above which a turn is ``cached``.
"""
def __init__(
self,
pricing: dict[str, float] | None = None,
*,
wasteful_token_threshold: int = 1000,
expensive_cost_threshold: float | None = None,
cached_ratio_threshold: float = 0.5,
) -> None:
self._pricing_explicit = pricing is not None
self.pricing: dict[str, float] = pricing or self.default_pricing()
self.wasteful_token_threshold = wasteful_token_threshold
self.expensive_cost_threshold = expensive_cost_threshold
self.cached_ratio_threshold = cached_ratio_threshold
# ---------- pricing ---------- #
@staticmethod
def default_pricing() -> dict[str, float]:
"""Default per-million-token USD prices (gpt-4o-mini)."""
return {"input": 0.15, "cached": 0.075, "output": 0.60}
def _cost_usd(
self, input_tokens: int, cached_tokens: int, output_tokens: int
) -> float:
"""USD cost for one turn given per-million-token pricing."""
uncached = max(input_tokens - cached_tokens, 0)
per_m = 1_000_000.0
return (
uncached / per_m * self.pricing.get("input", 0.0)
+ cached_tokens / per_m * self.pricing.get("cached", 0.0)
+ output_tokens / per_m * self.pricing.get("output", 0.0)
)
# ---------- turn normalization ---------- #
@staticmethod
def _parse_turn_id(turn: dict[str, Any], index: int) -> int:
raw = turn.get("turn_id")
if isinstance(raw, (int, float)):
return int(raw)
step = turn.get("step") or turn.get("turn") or ""
if isinstance(step, str):
m = _STEP_RE.search(step)
if m:
return int(m.group(1))
return index + 1
@staticmethod
def _coerce_int(value: Any) -> int:
"""Coerce nullable/numeric JSON values to int (None -> 0)."""
if value is None:
return 0
try:
return int(value)
except (TypeError, ValueError):
return 0
@staticmethod
def _coerce_float(value: Any) -> float:
if value is None:
return 0.0
try:
return float(value)
except (TypeError, ValueError):
return 0.0
def _normalize_turn(self, turn: dict[str, Any]) -> dict[str, Any]:
"""Map a raw turn/span dict onto the analyzer's canonical fields."""
input_tokens = self._coerce_int(
turn.get("prompt_tokens", turn.get("input_tokens"))
)
output_tokens = self._coerce_int(
turn.get("completion_tokens", turn.get("output_tokens"))
)
cached_tokens = self._coerce_int(turn.get("cached_tokens"))
explicit_ratio = turn.get("cache_hit_ratio")
if cached_tokens == 0 and explicit_ratio is not None:
cached_tokens = round(self._coerce_float(explicit_ratio) * input_tokens)
if input_tokens > 0:
cache_hit_ratio = cached_tokens / input_tokens
elif explicit_ratio is not None:
cache_hit_ratio = self._coerce_float(explicit_ratio)
else:
cache_hit_ratio = 0.0
cache_hit_ratio = max(0.0, min(1.0, cache_hit_ratio))
latency_ms: float
if turn.get("latency_ms") is not None:
latency_ms = self._coerce_float(turn.get("latency_ms"))
elif turn.get("latency_s") is not None:
latency_ms = self._coerce_float(turn.get("latency_s")) * 1000.0
else:
latency_ms = 0.0
tool_calls = turn.get("tool_calls")
if tool_calls is None:
tool = turn.get("tool")
tool_calls = 1 if (isinstance(tool, str) and tool) else 0
else:
tool_calls = self._coerce_int(tool_calls)
return {
"turn_id": self._parse_turn_id(turn, -1),
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"cached_tokens": cached_tokens,
"cache_hit_ratio": cache_hit_ratio,
"latency_ms": latency_ms,
"tool_calls": tool_calls,
"tool_ctx_tokens": self._coerce_int(turn.get("tool_ctx_tokens", -1)),
}
# ---------- classification ---------- #
def _classify(
self,
total_tokens: int,
tool_calls: int,
cache_hit_ratio: float,
cost_usd: float,
expensive_threshold: float,
) -> str:
if tool_calls == 0 and total_tokens >= self.wasteful_token_threshold:
return "wasteful"
if cost_usd >= expensive_threshold and expensive_threshold > 0:
return "expensive"
if cache_hit_ratio >= self.cached_ratio_threshold:
return "cached"
return "productive"
# ---------- public API ---------- #
def analyze_turn(self, turn: dict[str, Any]) -> TurnMetrics:
"""Analyze a single turn dict into :class:`TurnMetrics`.
Uses the absolute ``expensive_cost_threshold`` configured on the
analyzer; when it is ``None`` the turn is never classified expensive
here (a relative threshold is only available to :meth:`analyze_trajectory`,
which sees the whole distribution).
"""
n = self._normalize_turn(turn)
cost = self._cost_usd(n["input_tokens"], n["cached_tokens"], n["output_tokens"])
threshold = self.expensive_cost_threshold
if threshold is None:
threshold = float("inf")
classification = self._classify(
n["input_tokens"] + n["output_tokens"],
n["tool_calls"],
n["cache_hit_ratio"],
cost,
threshold,
)
return TurnMetrics(
turn_id=n["turn_id"],
input_tokens=n["input_tokens"],
output_tokens=n["output_tokens"],
cache_hit_ratio=n["cache_hit_ratio"],
cost_usd=cost,
latency_ms=n["latency_ms"],
tool_calls=n["tool_calls"],
classification=classification,
)
def _extract_turns(self, trajectory: dict[str, Any] | list[dict]) -> list[dict]:
"""Pull the list of turn dicts out of any supported trajectory shape."""
if isinstance(trajectory, list):
return list(trajectory)
if not isinstance(trajectory, dict):
raise TypeError(
"trajectory must be a list of turn dicts or a trace dict, "
f"got {type(trajectory).__name__}"
)
if "turns" in trajectory:
return list(trajectory["turns"] or [])
if "spans" in trajectory:
return list(trajectory["spans"] or [])
if "scenarios" in trajectory:
for scenario in trajectory["scenarios"] or []:
spans = scenario.get("spans") or []
if spans:
return list(spans)
return []
# A bare single-turn dict is treated as one turn.
if {"prompt_tokens", "input_tokens", "step", "tool"} & trajectory.keys():
return [trajectory]
return []
def analyze_trajectory(
self, trajectory: dict[str, Any] | list[dict]
) -> EfficiencyReport:
"""Analyze a full trajectory into an :class:`EfficiencyReport`."""
# Honour embedded pricing when no explicit pricing was configured.
if (
not self._pricing_explicit
and isinstance(trajectory, dict)
and isinstance(trajectory.get("pricing"), dict)
):
self.pricing = {**self.pricing, **trajectory["pricing"]}
turns = self._extract_turns(trajectory)
metrics = [self.analyze_turn(t) for t in turns]
total_turns = len(metrics)
total_cost = sum(m.cost_usd for m in metrics)
total_tokens = sum(m.total_tokens for m in metrics)
# Relative expensive threshold: 1.5x mean per-turn cost.
# When mean cost is zero (e.g. a fully cached or zero-token
# trajectory), every turn costs $0 and none should be flagged
# expensive — a zero threshold would mark all of them. Skip the
# relative reclassification in that case.
if self.expensive_cost_threshold is None and total_turns > 0:
mean_cost = total_cost / total_turns
rel_threshold = mean_cost * 1.5
if rel_threshold > 0:
for m in metrics:
if m.classification == "productive" and m.cost_usd >= rel_threshold:
m.classification = "expensive"
elif self.expensive_cost_threshold is None:
rel_threshold = float("inf")
else:
rel_threshold = self.expensive_cost_threshold
# Cumulative cost per turn (running sum).
cumulative: list[float] = []
running = 0.0
for m in metrics:
running += m.cost_usd
cumulative.append(running)
total_tool_calls = sum(m.tool_calls for m in metrics)
tokens_per_tool_call = (
total_tokens / total_tool_calls if total_tool_calls > 0 else 0.0
)
latency_per_turn = (
sum(m.latency_ms for m in metrics) / total_turns if total_turns > 0 else 0.0
)
# Efficiency score: productive-turn ratio weighted by token efficiency
# (fraction of tokens NOT spent on wasteful turns).
productive_turns = sum(1 for m in metrics if m.classification == "productive")
wasteful_tokens = sum(
m.total_tokens for m in metrics if m.classification == "wasteful"
)
if total_turns == 0:
efficiency_score = 0.0
else:
productive_ratio = productive_turns / total_turns
token_efficiency = (
1.0 - wasteful_tokens / total_tokens if total_tokens > 0 else 1.0
)
efficiency_score = max(0.0, min(1.0, productive_ratio * token_efficiency))
recommendations = self._recommendations(
metrics, efficiency_score, total_cost, total_tokens, rel_threshold
)
return EfficiencyReport(
total_turns=total_turns,
total_cost_usd=total_cost,
total_tokens=total_tokens,
efficiency_score=efficiency_score,
turn_metrics=metrics,
recommendations=recommendations,
cumulative_costs=cumulative,
tokens_per_tool_call=tokens_per_tool_call,
latency_per_turn=latency_per_turn,
)
# ---------- recommendations ---------- #
def _recommendations(
self,
metrics: list[TurnMetrics],
efficiency_score: float,
total_cost: float,
total_tokens: int,
expensive_threshold: float,
) -> list[str]:
recs: list[str] = []
# Wasteful turns: high tokens, no tool calls.
for m in metrics:
if m.classification == "wasteful":
recs.append(
f"Turn {m.turn_id} is wasteful: {m.total_tokens} tokens with "
f"no tool calls — consider context compression or early stopping."
)
# Expensive turns.
for m in metrics:
if m.classification == "expensive":
recs.append(
f"Turn {m.turn_id} is expensive: ${m.cost_usd:.6f} exceeds the "
f"${expensive_threshold:.6f}/turn threshold — review its prompt size."
)
# Cache miss pattern: high input tokens but low cache hit ratio overall.
if metrics:
high_input_turns = [m for m in metrics if m.input_tokens >= 1024]
if high_input_turns:
mean_ratio = sum(m.cache_hit_ratio for m in high_input_turns) / len(
high_input_turns
)
if mean_ratio < self.cached_ratio_threshold:
recs.append(
f"Cache miss pattern: mean cache hit ratio is {mean_ratio:.2%} "
f"across {len(high_input_turns)} turns with >=1024 input tokens "
f"— stabilize the prompt prefix to benefit from KV-cache."
)
# Context compression opportunity: tool_ctx_tokens growing across turns.
ctx_growth = self._max_tool_ctx_growth(metrics)
if ctx_growth > 0:
recs.append(
f"Context compression opportunity: tool context tokens grow by "
f"{ctx_growth} across the trajectory — summarize prior tool results "
f"to avoid re-billing them every turn."
)
# Overall efficiency verdict.
if metrics:
if efficiency_score < 0.5:
recs.append(
f"Low efficiency score ({efficiency_score:.2f}): fewer than half "
f"of turns are productive — review the trajectory structure."
)
elif efficiency_score >= 0.8:
recs.append(
f"High efficiency score ({efficiency_score:.2f}): trajectory is "
f"cost-efficient."
)
return recs
@staticmethod
def _max_tool_ctx_growth(metrics: list[TurnMetrics]) -> int:
"""Largest per-step increase in tool context tokens (0 if unknown)."""
# tool_ctx_tokens is not stored on TurnMetrics; recompute from the
# fact that input_tokens tend to grow as context accumulates. We use
# the raw input-token delta as a proxy when tool_ctx is unavailable.
if len(metrics) < 2:
return 0
growth = 0
prev = metrics[0].input_tokens
for m in metrics[1:]:
delta = m.input_tokens - prev
if delta > growth:
growth = delta
prev = m.input_tokens
return growth
if __name__ == "__main__": # pragma: no cover - manual smoke
import json
import pathlib
here = pathlib.Path(__file__).resolve().parent
trace = json.loads((here / "sample_trace.json").read_text(encoding="utf-8"))
analyzer = CostEfficiencyAnalyzer()
report = analyzer.analyze_trajectory(trace)
print(f"turns={report.total_turns} cost=${report.total_cost_usd:.6f} "
f"score={report.efficiency_score:.3f}")
for r in report.recommendations:
print(" -", r)
+291
View File
@@ -0,0 +1,291 @@
"""
实验 7-9:Agent 任务的端到端成本分析(可运行 demo + CLI)。
两种运行方式:
1) 在线(--live,默认):真实调用模型(默认 gpt-5.6-luna),token 与 cached_tokens
取自 API 返回的 usage,成本按单价换算。需要 OPENAI_API_KEY 或 OPENROUTER_API_KEY
(无 OpenAI key 时自动回退到 OpenRoutergpt-5.x 只要有 OpenRouter key 就优先走它)。
2) 离线(--offline):不打模型,读入一份此前真实运行录下的 tracecanned token
counts),用可配置的单价重新计算成本、成本构成与 A/B 对比表。无需 API key。
无论哪种方式,都会产出两份交付:
(a) 单次任务的「按步骤 + 按成本构成」拆解(哪一步最贵、输入/缓存/输出各占多少)。
(b) A/B 对比表:朴素 vs 仅 KV-cache vs 仅压缩 vs 两者叠加(完整 2×2),
量化 总 token / 缓存 token / 缓存率 / 成本 / 相对基线的节省。
示例:
python demo.py # 在线,默认跑 A(朴素)+B(优化) 两组
python demo.py --scenario all # 在线,跑完整 2×2 四组
python demo.py --live --save-trace out.json # 在线跑并把真实用量落盘
python demo.py --offline # 离线,用内置 sample_trace.json 重算
python demo.py --offline --model gpt-4o # 离线,换 gpt-4o 单价重算同一份用量
python demo.py --offline --price-input 0.20 --price-cached 0.10 --price-output 0.80
"""
import argparse
import json
import os
import sys
import config
from config import PRICING_PRESETS, Pricing
DEFAULT_TRACE = os.path.join(os.path.dirname(__file__), "sample_trace.json")
SCENARIO_KEYS = ["naive", "kv", "compress", "both"]
def _pct(saved: float, base: float) -> str:
if base == 0:
return "0.0%"
return f"{saved / base * 100:.1f}%"
def build_pricing(args) -> Pricing:
"""根据 --model 预设 + --price-* 覆盖,构造本次计费用的单价。"""
base = PRICING_PRESETS.get(args.model)
if base is None:
base = config.default_pricing()
return Pricing(
input_per_m=args.price_input if args.price_input is not None else base.input_per_m,
cached_per_m=args.price_cached if args.price_cached is not None else base.cached_per_m,
output_per_m=args.price_output if args.price_output is not None else base.output_per_m,
)
def resolve_scenarios(arg: str):
"""把 --scenario 解析成有序去重的场景 key 列表。"""
if arg == "all":
return list(SCENARIO_KEYS)
if arg == "ab":
return ["naive", "both"]
keys, seen = [], set()
for k in arg.split(","):
k = k.strip()
if k and k not in seen:
keys.append(k)
seen.add(k)
return keys
# ---------------------------------------------------------------------------
# 采集:在线跑真实模型,或离线从 trace 文件读回
# ---------------------------------------------------------------------------
def collect_live(keys, pricing, warmup: bool):
import agent
if not (os.environ.get("OPENAI_API_KEY") or os.environ.get("OPENROUTER_API_KEY")):
print("未检测到 OPENAI_API_KEY 或 OPENROUTER_API_KEY,请先 export 其一 "
"(无 OpenAI key 时会自动回退到 OpenRouter),或改用 --offline(离线复算,无需 key)。",
file=sys.stderr)
sys.exit(1)
# 构造 client 并解析实际模型名(可能被回退映射成 OpenRouter id)。
client, resolved = config.make_client_and_model(config.MODEL)
if resolved != config.MODEL:
print(f">>> 已回退到 OpenRouter:模型 {config.MODEL} -> {resolved}")
config.MODEL = resolved
agent.MODEL = resolved
try:
agent._encoder.cache_clear()
except Exception:
pass
tracers = []
for k in keys:
name, kv, compress = agent.SCENARIOS[k]
# KV-cache 组先跑一次「预热」,把稳定前缀写入 OpenAI 的 prompt cache
# 让正式计量时更稳定地命中 cached_tokens(真实系统里前缀早已是热的)。
if kv and warmup:
print(f">>> 预热 [{name}] 的稳定前缀(写入 prompt cache...")
agent.run_scenario(client, kv, compress, name=name, pricing=pricing)
print(f">>> 正在运行 [{name}] {'(在线计量)' if kv else ''}...")
tr = agent.run_scenario(client, kv, compress, name=name, pricing=pricing)
tracers.append((k, tr))
return tracers
def collect_offline(keys, pricing, trace_path):
from tracer import Tracer
if not os.path.exists(trace_path):
print(f"找不到 trace 文件:{trace_path}", file=sys.stderr)
sys.exit(1)
with open(trace_path, "r", encoding="utf-8") as f:
data = json.load(f)
by_key = {s.get("key", s.get("name")): s for s in data.get("scenarios", [])}
print(f"离线模式:读入 {trace_path}")
print(f" 该 trace 采集自模型 = {data.get('model', '?')}"
f"{len(by_key)} 个场景的真实录制用量(token 数为实测,成本按当前单价重算)。")
tracers = []
for k in keys:
sc = by_key.get(k)
if sc is None:
print(f" [跳过] trace 中没有场景 '{k}'(可用在线模式 --save-trace 补录)",
file=sys.stderr)
continue
spans = sc.get("spans")
if not spans:
print(f" [跳过] trace 中场景 '{k}' 缺少 spans 数据"
f"(可用在线模式 --save-trace 补录)", file=sys.stderr)
continue
tr = Tracer.from_records(spans, name=sc.get("name", k), pricing=pricing)
tracers.append((k, tr))
if not tracers:
print("trace 里没有任何被选中的场景,退出。", file=sys.stderr)
sys.exit(1)
return tracers
# ---------------------------------------------------------------------------
# 交付 (b)A/B 对比表
# ---------------------------------------------------------------------------
def print_ab_table(tracers):
print("\n\n===== A/B 成本对比(同一个 8 轮客服退款任务)=====")
header = (f"{'方案':<26} {'总输入tok':>10} {'缓存tok':>10} {'缓存率':>8} "
f"{'输出tok':>8} {'总成本($)':>12} {'vs基线':>10}")
print(header)
print("-" * len(header))
base_cost = tracers[0][1].total_cost()
for _, tr in tracers:
pin = tr.total_prompt_tokens()
cac = tr.total_cached_tokens()
rate = f"{(cac / pin * 100):.1f}%" if pin else "0.0%"
cost = tr.total_cost()
vs = "基线" if abs(cost - base_cost) < 1e-12 else f"-{_pct(base_cost - cost, base_cost)}"
print(f"{tr.name:<26} {pin:>10} {cac:>10} {rate:>8} "
f"{tr.total_completion_tokens():>8} {cost:>12.6f} {vs:>10}")
print("-" * len(header))
# 用第一个(基线)和最后一个(通常是 both 优化)做重点量化
base_k, base = tracers[0]
best_k, best = tracers[-1]
if base_k != best_k:
tok_a = base.total_prompt_tokens() + base.total_completion_tokens()
tok_b = best.total_prompt_tokens() + best.total_completion_tokens()
cost_a, cost_b = base.total_cost(), best.total_cost()
print(f"\n重点对比:{base.name}{best.name}")
print(f" 总 token: A={tok_a} → B={tok_b} "
f"减少 {tok_a - tok_b} ({_pct(tok_a - tok_b, tok_a)})")
print(f" 缓存 token: A={base.total_cached_tokens()}"
f"B={best.total_cached_tokens()} B 靠稳定前缀命中缓存)")
print(f" 总成本: A=${cost_a:.6f} → B=${cost_b:.6f} "
f"降低 ${cost_a - cost_b:.6f} ({_pct(cost_a - cost_b, cost_a)})")
if cost_b > 0:
print(f" 成本倍率: A 是 B 的 {cost_a / cost_b:.2f}")
print("\n结论: 稳定长前缀让重复的系统提示/工具定义/历史轮次按缓存价计费,")
print(" 叠加上下文压缩控制上下文增长,二者共同显著降低了端到端成本。")
def dump_output(path, tracers, pricing, model):
out = {
"model": model,
"pricing": {"input": pricing.input_per_m, "cached": pricing.cached_per_m,
"output": pricing.output_per_m},
"scenarios": [],
}
import agent
for k, tr in tracers:
name = agent.SCENARIOS[k][0] if k in agent.SCENARIOS else tr.name
out["scenarios"].append({
"key": k, "name": name,
"total_cost": tr.total_cost(),
"component_costs": tr.component_costs(),
"cost_distribution": tr.cost_distribution(),
"spans": tr.to_records(),
})
with open(path, "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
print(f"\n已写出结果到 {path}")
def build_parser():
p = argparse.ArgumentParser(
prog="demo.py",
description="实验 7-9:Agent 任务端到端成本分析——对客服退款 Agent 做全链路成本拆解,"
"并对比 KV-cache / 上下文压缩两个杠杆的成本差异(完整 2×2 A/B)。",
epilog="示例:\n"
" python demo.py # 在线,默认跑 A(朴素)+B(优化)\n"
" python demo.py --scenario all # 在线,跑完整 2×2 四组\n"
" python demo.py --offline # 离线,用内置 canned trace 重算(无需 key\n"
" python demo.py --offline --model gpt-4o # 换单价离线重算\n"
" python demo.py --live --save-trace out.json # 在线跑并落盘真实用量\n",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
mode = p.add_mutually_exclusive_group()
mode.add_argument("--live", action="store_true",
help="在线模式(默认):真实调用 OpenAI,需要 OPENAI_API_KEY。")
mode.add_argument("--offline", action="store_true",
help="离线模式:不打模型,从 trace 文件读真实录制的 token 用量并按单价重算成本。")
p.add_argument("--trace", metavar="FILE", default=DEFAULT_TRACE,
help=f"离线模式读取的 tracecanned token counts)文件,默认 {os.path.basename(DEFAULT_TRACE)}")
p.add_argument("--save-trace", metavar="FILE", default=None,
help="在线模式下把本次真实 token 用量落盘为 trace 文件(供之后 --offline 复算)。")
p.add_argument("--scenario", metavar="NAME", default="ab",
help="选择要跑的 A/B 场景:ab(默认,=naive+both) / all(2×2 四组) / "
"或逗号分隔的子集 naive,kv,compress,both。")
p.add_argument("--model", metavar="NAME", default=config.MODEL,
help=f"模型名(决定默认单价预设,可选 {', '.join(PRICING_PRESETS)}),"
f"默认 {config.MODEL}")
p.add_argument("--price-input", type=float, default=None,
help="覆盖输入单价(每百万 token 美元)。")
p.add_argument("--price-cached", type=float, default=None,
help="覆盖缓存命中输入单价(每百万 token 美元)。")
p.add_argument("--price-output", type=float, default=None,
help="覆盖输出单价(每百万 token 美元)。")
p.add_argument("--no-warmup", action="store_true",
help="在线模式下关闭 KV-cache 组的前缀预热(默认预热以稳定命中缓存)。")
p.add_argument("--output", metavar="FILE", default=None,
help="把成本拆解结果(含成本构成/分布/逐步用量)写成 JSON 文件。")
return p
def main():
args = build_parser().parse_args()
# 让 agent / tracer 使用选定模型
config.MODEL = args.model
try:
import agent
agent.MODEL = args.model
agent._encoder.cache_clear()
except Exception:
pass
pricing = build_pricing(args)
keys = resolve_scenarios(args.scenario)
bad = [k for k in keys if k not in SCENARIO_KEYS]
if bad:
print(f"未知场景 {bad},可选:{SCENARIO_KEYS} / all / ab", file=sys.stderr)
sys.exit(2)
print(f"模型: {args.model}")
print(f"单价(每百万token): 输入 ${pricing.input_per_m} / 缓存输入 ${pricing.cached_per_m} "
f"/ 输出 ${pricing.output_per_m}")
if args.offline:
tracers = collect_offline(keys, pricing, args.trace)
else:
print("说明: OpenAI prompt caching 自动生效(前缀>=1024token 且近期命中相同前缀),")
print(" 命中的输入 token 出现在 usage.prompt_tokens_details.cached_tokens。")
tracers = collect_live(keys, pricing, warmup=not args.no_warmup)
if args.save_trace:
dump_output(args.save_trace, tracers, pricing, args.model)
# 交付 (a):逐场景成本拆解
for _, tr in tracers:
tr.print_breakdown(title=f"{tr.name}(单次任务全链路拆解)")
# 交付 (b)A/B 对比表
if len(tracers) >= 2:
print_ab_table(tracers)
if args.output:
dump_output(args.output, tracers, pricing, args.model)
if __name__ == "__main__":
main()
+10
View File
@@ -0,0 +1,10 @@
# 复制为 .env 或直接 export。默认模型为 gpt-5.6-luna(当前廉价旗舰)。
# 首选 OPENAI_API_KEY;若未设置则自动回退到 OPENROUTER_API_KEY(走 OpenRouter 兼容端点)。
# 提示:gpt-5.x 直连 OpenAI 需组织实名认证,只要设置了 OPENROUTER_API_KEY 就会优先走它。
OPENAI_API_KEY=your-openai-api-key
# OpenRouter 回退 key(无 OPENAI_API_KEY 时使用;模型名自动映射 gpt-* -> openai/*
# OPENROUTER_API_KEY=your-openrouter-api-key
# 可选:覆盖默认模型(默认 gpt-5.6-luna
# COST_DEMO_MODEL=gpt-5.6-luna
@@ -0,0 +1,3 @@
openai>=1.40.0
python-dotenv>=1.0.0
tiktoken>=0.7.0
@@ -0,0 +1,426 @@
{
"model": "kimi-k2.5",
"pricing": {
"input": 0.5911688312,
"cached": 0.1034545455,
"output": 3.1036363636
},
"scenarios": [
{
"key": "naive",
"name": "A 朴素(无缓存/无压缩)",
"total_cost": 0.0152468353252392,
"component_costs": {
"uncached_input_cost": 0.0115224716889192,
"cached_input_cost": 0.0,
"output_cost": 0.00372436363632,
"uncached_input_tokens": 19491,
"cached_input_tokens": 0,
"output_tokens": 1200,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0019058544156549,
"p50": 0.0019094753247440002,
"p95": 0.0025698109091944,
"p99": 0.0025698109091944,
"max": 0.0025698109091944
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1013,
"cached_tokens": 0,
"completion_tokens": 120,
"tool_ctx_tokens": 300,
"latency_s": 2.287337064743042
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1670,
"cached_tokens": 0,
"completion_tokens": 120,
"tool_ctx_tokens": 899,
"latency_s": 3.4959402084350586
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2008,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 5.752177000045776
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2390,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1482,
"latency_s": 3.8040597438812256
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2682,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1585,
"latency_s": 6.000391006469727
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2970,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1693,
"latency_s": 5.413041353225708
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3251,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1779,
"latency_s": 5.660830974578857
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3507,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1847,
"latency_s": 5.281260013580322
}
]
},
{
"key": "kv",
"name": "KV 仅缓存(稳定前缀/不压缩)",
"total_cost": 0.0079609602864777,
"component_costs": {
"uncached_input_cost": 0.0029623470131432,
"cached_input_cost": 0.0014759860006485,
"output_cost": 0.003522627272686,
"uncached_input_tokens": 5011,
"cached_input_tokens": 14267,
"output_tokens": 1135,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0009951200358097126,
"p50": 0.001072025558548,
"p95": 0.0011735883637072,
"p99": 0.0011735883637072,
"max": 0.0011735883637072
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 129,
"tool_ctx_tokens": 300,
"latency_s": 5.202546119689941
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1620,
"cached_tokens": 512,
"completion_tokens": 145,
"tool_ctx_tokens": 899,
"latency_s": 4.699679136276245
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 1990,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 4.436970949172974
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2375,
"cached_tokens": 1536,
"completion_tokens": 160,
"tool_ctx_tokens": 1482,
"latency_s": 5.637408018112183
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2663,
"cached_tokens": 2048,
"completion_tokens": 160,
"tool_ctx_tokens": 1585,
"latency_s": 4.351202964782715
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2953,
"cached_tokens": 2304,
"completion_tokens": 160,
"tool_ctx_tokens": 1693,
"latency_s": 5.36053204536438
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3233,
"cached_tokens": 2816,
"completion_tokens": 160,
"tool_ctx_tokens": 1779,
"latency_s": 4.889960765838623
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3489,
"cached_tokens": 3072,
"completion_tokens": 61,
"tool_ctx_tokens": 1847,
"latency_s": 1.9250061511993408
}
]
},
{
"key": "compress",
"name": "仅压缩(前缀不稳定/摘要)",
"total_cost": 0.0132817901303144,
"component_costs": {
"uncached_input_cost": 0.009309135584906399,
"cached_input_cost": 0.0,
"output_cost": 0.003972654545408001,
"uncached_input_tokens": 15747,
"cached_input_tokens": 0,
"output_tokens": 1280,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0016602237662893,
"p50": 0.0017338981818776,
"p95": 0.00189765194812,
"p99": 0.00189765194812,
"max": 0.00189765194812
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1012,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 300,
"latency_s": 5.470736980438232
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1710,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 899,
"latency_s": 4.1132972240448
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2093,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 3.5907950401306152
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2227,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1231,
"latency_s": 3.2491540908813477
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2021,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 780,
"latency_s": 7.461982250213623
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2114,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 677,
"latency_s": 4.925306081771851
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2200,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 486,
"latency_s": 5.416359186172485
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2370,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 499,
"latency_s": 4.96558403968811
}
]
},
{
"key": "both",
"name": "B 优化(KV缓存+压缩)",
"total_cost": 0.0086632688576729,
"component_costs": {
"uncached_input_cost": 0.0041033028573592,
"cached_input_cost": 0.0008138769094485001,
"output_cost": 0.0037460890908652,
"uncached_input_tokens": 6941,
"cached_input_tokens": 7867,
"output_tokens": 1207,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0010829086072091125,
"p50": 0.0011068749610916,
"p95": 0.0013616982857848,
"p99": 0.0013616982857848,
"max": 0.0013616982857848
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 116,
"tool_ctx_tokens": 300,
"latency_s": 2.2258410453796387
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1607,
"cached_tokens": 512,
"completion_tokens": 131,
"tool_ctx_tokens": 899,
"latency_s": 4.5919647216796875
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 1963,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 7.6671669483184814
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2097,
"cached_tokens": 768,
"completion_tokens": 160,
"tool_ctx_tokens": 1231,
"latency_s": 5.431225061416626
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1887,
"cached_tokens": 768,
"completion_tokens": 160,
"tool_ctx_tokens": 780,
"latency_s": 5.2519941329956055
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 1988,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 677,
"latency_s": 4.7071919441223145
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2073,
"cached_tokens": 1280,
"completion_tokens": 160,
"tool_ctx_tokens": 486,
"latency_s": 5.377019882202148
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2238,
"cached_tokens": 1536,
"completion_tokens": 160,
"tool_ctx_tokens": 499,
"latency_s": 4.964300155639648
}
]
}
]
}
@@ -0,0 +1,426 @@
{
"model": "kimi-k2.5",
"pricing": {
"input": 0.5911688312,
"cached": 0.1034545455,
"output": 3.1036363636
},
"scenarios": [
{
"key": "naive",
"name": "A 朴素(无缓存/无压缩)",
"total_cost": 0.0152468353252392,
"component_costs": {
"uncached_input_cost": 0.0115224716889192,
"cached_input_cost": 0.0,
"output_cost": 0.00372436363632,
"uncached_input_tokens": 19491,
"cached_input_tokens": 0,
"output_tokens": 1200,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0019058544156549,
"p50": 0.0019094753247440002,
"p95": 0.0025698109091944,
"p99": 0.0025698109091944,
"max": 0.0025698109091944
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1013,
"cached_tokens": 0,
"completion_tokens": 120,
"tool_ctx_tokens": 300,
"latency_s": 2.287337064743042
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1670,
"cached_tokens": 0,
"completion_tokens": 120,
"tool_ctx_tokens": 899,
"latency_s": 3.4959402084350586
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2008,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 5.752177000045776
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2390,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1482,
"latency_s": 3.8040597438812256
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2682,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1585,
"latency_s": 6.000391006469727
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2970,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1693,
"latency_s": 5.413041353225708
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3251,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1779,
"latency_s": 5.660830974578857
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3507,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1847,
"latency_s": 5.281260013580322
}
]
},
{
"key": "kv",
"name": "KV 仅缓存(稳定前缀/不压缩)",
"total_cost": 0.0079609602864777,
"component_costs": {
"uncached_input_cost": 0.0029623470131432,
"cached_input_cost": 0.0014759860006485,
"output_cost": 0.003522627272686,
"uncached_input_tokens": 5011,
"cached_input_tokens": 14267,
"output_tokens": 1135,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0009951200358097126,
"p50": 0.001072025558548,
"p95": 0.0011735883637072,
"p99": 0.0011735883637072,
"max": 0.0011735883637072
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 129,
"tool_ctx_tokens": 300,
"latency_s": 5.202546119689941
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1620,
"cached_tokens": 512,
"completion_tokens": 145,
"tool_ctx_tokens": 899,
"latency_s": 4.699679136276245
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 1990,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 4.436970949172974
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2375,
"cached_tokens": 1536,
"completion_tokens": 160,
"tool_ctx_tokens": 1482,
"latency_s": 5.637408018112183
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2663,
"cached_tokens": 2048,
"completion_tokens": 160,
"tool_ctx_tokens": 1585,
"latency_s": 4.351202964782715
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2953,
"cached_tokens": 2304,
"completion_tokens": 160,
"tool_ctx_tokens": 1693,
"latency_s": 5.36053204536438
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3233,
"cached_tokens": 2816,
"completion_tokens": 160,
"tool_ctx_tokens": 1779,
"latency_s": 4.889960765838623
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3489,
"cached_tokens": 3072,
"completion_tokens": 61,
"tool_ctx_tokens": 1847,
"latency_s": 1.9250061511993408
}
]
},
{
"key": "compress",
"name": "仅压缩(前缀不稳定/摘要)",
"total_cost": 0.0132817901303144,
"component_costs": {
"uncached_input_cost": 0.009309135584906399,
"cached_input_cost": 0.0,
"output_cost": 0.003972654545408001,
"uncached_input_tokens": 15747,
"cached_input_tokens": 0,
"output_tokens": 1280,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0016602237662893,
"p50": 0.0017338981818776,
"p95": 0.00189765194812,
"p99": 0.00189765194812,
"max": 0.00189765194812
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1012,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 300,
"latency_s": 5.470736980438232
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1710,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 899,
"latency_s": 4.1132972240448
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2093,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 3.5907950401306152
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2227,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1231,
"latency_s": 3.2491540908813477
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2021,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 780,
"latency_s": 7.461982250213623
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2114,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 677,
"latency_s": 4.925306081771851
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2200,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 486,
"latency_s": 5.416359186172485
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2370,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 499,
"latency_s": 4.96558403968811
}
]
},
{
"key": "both",
"name": "B 优化(KV缓存+压缩)",
"total_cost": 0.0086632688576729,
"component_costs": {
"uncached_input_cost": 0.0041033028573592,
"cached_input_cost": 0.0008138769094485001,
"output_cost": 0.0037460890908652,
"uncached_input_tokens": 6941,
"cached_input_tokens": 7867,
"output_tokens": 1207,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0010829086072091125,
"p50": 0.0011068749610916,
"p95": 0.0013616982857848,
"p99": 0.0013616982857848,
"max": 0.0013616982857848
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 116,
"tool_ctx_tokens": 300,
"latency_s": 2.2258410453796387
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1607,
"cached_tokens": 512,
"completion_tokens": 131,
"tool_ctx_tokens": 899,
"latency_s": 4.5919647216796875
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 1963,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 1164,
"latency_s": 7.6671669483184814
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2097,
"cached_tokens": 768,
"completion_tokens": 160,
"tool_ctx_tokens": 1231,
"latency_s": 5.431225061416626
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1887,
"cached_tokens": 768,
"completion_tokens": 160,
"tool_ctx_tokens": 780,
"latency_s": 5.2519941329956055
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 1988,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 677,
"latency_s": 4.7071919441223145
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2073,
"cached_tokens": 1280,
"completion_tokens": 160,
"tool_ctx_tokens": 486,
"latency_s": 5.377019882202148
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2238,
"cached_tokens": 1536,
"completion_tokens": 160,
"tool_ctx_tokens": 499,
"latency_s": 4.964300155639648
}
]
}
]
}
@@ -0,0 +1,54 @@
{
"schema_version": 1,
"experiment": "7-9",
"status": "incomplete",
"generated_at_utc": "2026-07-29T22:31:57.870199+00:00",
"run_dir": "chapter7/agent-cost-analysis/runs/exp7-9-kimi25-20260730-v2",
"git_commit": "4a7f37cf278bd15948c409f14533017c4c7fbc29",
"command": "python demo.py --live --scenario all --model kimi-k2.5 --save-trace TRACE --output REPORT",
"provider_receipt_count": 32,
"status_reasons": [
"The real 2x2 run covers one eight-turn refund workflow; the manuscript asks for several representative task types and aggregate task-level p50/p95/p99.",
"Reasoning-token usage and provider response IDs are recorded, but success-quality equivalence across optimizations is not independently judged.",
"The 32 measured calls have unique provider response IDs; 16 cache-warmup calls were intentionally excluded from measured traces and do not have retained receipts."
],
"inputs": [
{
"path": "chapter7/agent-cost-analysis/agent.py",
"bytes": 17109,
"sha256": "30866b34ead7ff2897964aaf1e07867691102982971903f8a2a4172c48bfd718"
},
{
"path": "chapter7/agent-cost-analysis/tracer.py",
"bytes": 12309,
"sha256": "fb673ac8f10ca037b6df3822e00e6cbb21b8381b731736fb1b2a30748e7226fe"
},
{
"path": "chapter7/agent-cost-analysis/demo.py",
"bytes": 13477,
"sha256": "75dc27633829fe2a730c5654a38f6af90fd58eceee630dfa985e386fa536c104"
},
{
"path": "chapter7/agent-cost-analysis/config.py",
"bytes": 5701,
"sha256": "0a92de4849994e74b28e3b5ad6c308cc154a2b984d8827f859b7f15d9fbbb463"
},
{
"path": "chapter7/model-benchmark/campaign_config.json",
"bytes": 8103,
"sha256": "9a839ad907852798b35af5b4623814c343943ae1596a78d650d1a1affc93b17a"
}
],
"artifacts": [
{
"path": "chapter7/agent-cost-analysis/runs/exp7-9-kimi25-20260730-v2/report.json",
"bytes": 17810,
"sha256": "92b0ee29c7df3dd28778d625fe699662dee56a5401141631ead450cd2ca37ace"
},
{
"path": "chapter7/agent-cost-analysis/runs/exp7-9-kimi25-20260730-v2/trace.json",
"bytes": 17810,
"sha256": "92b0ee29c7df3dd28778d625fe699662dee56a5401141631ead450cd2ca37ace"
}
]
}
@@ -0,0 +1,554 @@
{
"model": "kimi-k2.5",
"pricing": {
"input": 0.5911688312,
"cached": 0.1034545455,
"output": 3.1036363636
},
"scenarios": [
{
"key": "naive",
"name": "A 朴素(无缓存/无压缩)",
"total_cost": 0.01574031350707,
"component_costs": {
"uncached_input_cost": 0.0118080062343888,
"cached_input_cost": 0.0,
"output_cost": 0.0039323072726812,
"uncached_input_tokens": 19974,
"cached_input_tokens": 0,
"output_tokens": 1267,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.00196753918838375,
"p50": 0.0019544041559152,
"p95": 0.0026135574027031996,
"p99": 0.0026135574027031996,
"max": 0.0026135574027031996
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1015,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 5.299084663391113,
"response_id": "chatcmpl-6a6a7cb80695ae5455eab84a",
"response_model": "kimi-k2.5",
"response_created": 1785363640
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1706,
"cached_tokens": 0,
"completion_tokens": 147,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 5.079176902770996,
"response_id": "chatcmpl-6a6a7cbd80925af8edb8bb5e",
"response_model": "kimi-k2.5",
"response_created": 1785363646
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2081,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 3.0797362327575684,
"response_id": "chatcmpl-6a6a7cc24511d264b10f37dd",
"response_model": "kimi-k2.5",
"response_created": 1785363651
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2466,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1482,
"latency_s": 3.4901974201202393,
"response_id": "chatcmpl-6a6a7cc5edbcdbb3613f3353",
"response_model": "kimi-k2.5",
"response_created": 1785363654
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2753,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1585,
"latency_s": 6.93122410774231,
"response_id": "chatcmpl-6a6a7cc97f6a7380b55ebedb",
"response_model": "kimi-k2.5",
"response_created": 1785363658
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 3048,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1693,
"latency_s": 5.43744421005249,
"response_id": "chatcmpl-6a6a7cd0d7374244555b9001",
"response_model": "kimi-k2.5",
"response_created": 1785363665
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3324,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1779,
"latency_s": 5.579896926879883,
"response_id": "chatcmpl-6a6a7cd56f58f5f62fa2981a",
"response_model": "kimi-k2.5",
"response_created": 1785363670
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3581,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1847,
"latency_s": 5.334557056427002,
"response_id": "chatcmpl-6a6a7cdb2e6a18d37452de21",
"response_model": "kimi-k2.5",
"response_created": 1785363676
}
]
},
{
"key": "kv",
"name": "KV 仅缓存(稳定前缀/不压缩)",
"total_cost": 0.0082537957670069,
"component_costs": {
"uncached_input_cost": 0.0027956374027448,
"cached_input_cost": 0.0015289547279445,
"output_cost": 0.0039292036363176,
"uncached_input_tokens": 4729,
"cached_input_tokens": 14779,
"output_tokens": 1266,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0010317244708758625,
"p50": 0.001053935792344,
"p95": 0.0011930969351368,
"p99": 0.0011930969351368,
"max": 0.0011930969351368
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 2.981578826904297,
"response_id": "chatcmpl-6a6a7d043905097c262cd2d8",
"response_model": "kimi-k2.5",
"response_created": 1785363716
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1652,
"cached_tokens": 512,
"completion_tokens": 146,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 8.04120397567749,
"response_id": "chatcmpl-6a6a7d07d5eb4229617017c7",
"response_model": "kimi-k2.5",
"response_created": 1785363719
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2023,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 6.327085256576538,
"response_id": "chatcmpl-6a6a7d0f3905097c262cd2f5",
"response_model": "kimi-k2.5",
"response_created": 1785363727
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2408,
"cached_tokens": 1792,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1482,
"latency_s": 4.883957862854004,
"response_id": "chatcmpl-6a6a7d15bb40cba1f65cc486",
"response_model": "kimi-k2.5",
"response_created": 1785363733
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2696,
"cached_tokens": 2048,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1585,
"latency_s": 5.0474817752838135,
"response_id": "chatcmpl-6a6a7d1adc87cefde7ebcca2",
"response_model": "kimi-k2.5",
"response_created": 1785363738
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2986,
"cached_tokens": 2560,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1693,
"latency_s": 3.6142687797546387,
"response_id": "chatcmpl-6a6a7d1f3905097c262cd332",
"response_model": "kimi-k2.5",
"response_created": 1785363743
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3266,
"cached_tokens": 2816,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1779,
"latency_s": 3.981088161468506,
"response_id": "chatcmpl-6a6a7d233905097c262cd33e",
"response_model": "kimi-k2.5",
"response_created": 1785363747
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3522,
"cached_tokens": 3072,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1847,
"latency_s": 5.233341217041016,
"response_id": "chatcmpl-6a6a7d27d5eb42296170182d",
"response_model": "kimi-k2.5",
"response_created": 1785363751
}
]
},
{
"key": "compress",
"name": "仅压缩(前缀不稳定/摘要)",
"total_cost": 0.0125982511692596,
"component_costs": {
"uncached_input_cost": 0.0089390638965752,
"cached_input_cost": 0.0,
"output_cost": 0.0036591872726844,
"uncached_input_tokens": 15121,
"cached_input_tokens": 0,
"output_tokens": 1179,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.00157478139615745,
"p50": 0.0016801018182384,
"p95": 0.0018338057143504,
"p99": 0.0018338057143504,
"max": 0.0018338057143504
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1010,
"cached_tokens": 0,
"completion_tokens": 121,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 3.843111276626587,
"response_id": "chatcmpl-6a6a7d2c96a5dde81c3844dd",
"response_model": "kimi-k2.5",
"response_created": 1785363756
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1671,
"cached_tokens": 0,
"completion_tokens": 110,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 3.103135108947754,
"response_id": "chatcmpl-6a6a7d3006fe5c9e72518409",
"response_model": "kimi-k2.5",
"response_created": 1785363760
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2002,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 5.465490102767944,
"response_id": "chatcmpl-6a6a7d3384ead76d6951b7bb",
"response_model": "kimi-k2.5",
"response_created": 1785363763
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2136,
"cached_tokens": 0,
"completion_tokens": 149,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1231,
"latency_s": 3.5718801021575928,
"response_id": "chatcmpl-6a6a7d3880925af8edb8bcdd",
"response_model": "kimi-k2.5",
"response_created": 1785363769
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1916,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 780,
"latency_s": 3.5103132724761963,
"response_id": "chatcmpl-6a6a7d3cfa52c1b646ff684d",
"response_model": "kimi-k2.5",
"response_created": 1785363772
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2022,
"cached_tokens": 0,
"completion_tokens": 159,
"reasoning_tokens": 1,
"tool_ctx_tokens": 677,
"latency_s": 3.597494125366211,
"response_id": "chatcmpl-6a6a7d3f8a024a47783f03bf",
"response_model": "kimi-k2.5",
"response_created": 1785363776
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2102,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 486,
"latency_s": 9.166883945465088,
"response_id": "chatcmpl-6a6a7d43503e5013c8fe6ef8",
"response_model": "kimi-k2.5",
"response_created": 1785363779
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2262,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 499,
"latency_s": 4.3432776927948,
"response_id": "chatcmpl-6a6a7d4c687ea3db1a5a4d6c",
"response_model": "kimi-k2.5",
"response_created": 1785363789
}
]
},
{
"key": "both",
"name": "B 优化(KV缓存+压缩)",
"total_cost": 0.009057608026520501,
"component_costs": {
"uncached_input_cost": 0.0042445922080159995,
"cached_input_cost": 0.0008403612730964999,
"output_cost": 0.003972654545408001,
"uncached_input_tokens": 7180,
"cached_input_tokens": 8123,
"output_tokens": 1280,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0011322010033150626,
"p50": 0.0011570356364336,
"p95": 0.0014060359481248,
"p99": 0.0014060359481248,
"max": 0.0014060359481248
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 3.2073097229003906,
"response_id": "chatcmpl-6a6a7d7658a6e8702ecebaee",
"response_model": "kimi-k2.5",
"response_created": 1785363830
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1652,
"cached_tokens": 512,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 7.249981880187988,
"response_id": "chatcmpl-6a6a7d792d20c6aa1053113f",
"response_model": "kimi-k2.5",
"response_created": 1785363833
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2038,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 4.833154916763306,
"response_id": "chatcmpl-6a6a7d80f9aa90a3ebfd7fe5",
"response_model": "kimi-k2.5",
"response_created": 1785363841
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2172,
"cached_tokens": 768,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1231,
"latency_s": 5.391065835952759,
"response_id": "chatcmpl-6a6a7d852d20c6aa10531154",
"response_model": "kimi-k2.5",
"response_created": 1785363846
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1962,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 780,
"latency_s": 5.175268888473511,
"response_id": "chatcmpl-6a6a7d8b2d20c6aa10531162",
"response_model": "kimi-k2.5",
"response_created": 1785363851
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2063,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 677,
"latency_s": 5.356127977371216,
"response_id": "chatcmpl-6a6a7d90a183ecb32232c19d",
"response_model": "kimi-k2.5",
"response_created": 1785363856
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2148,
"cached_tokens": 1280,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 486,
"latency_s": 7.17538595199585,
"response_id": "chatcmpl-6a6a7d953905097c262cd43c",
"response_model": "kimi-k2.5",
"response_created": 1785363862
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2313,
"cached_tokens": 1536,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 499,
"latency_s": 5.300028085708618,
"response_id": "chatcmpl-6a6a7d9d56e2191f9ae786d6",
"response_model": "kimi-k2.5",
"response_created": 1785363869
}
]
}
]
}
@@ -0,0 +1,554 @@
{
"model": "kimi-k2.5",
"pricing": {
"input": 0.5911688312,
"cached": 0.1034545455,
"output": 3.1036363636
},
"scenarios": [
{
"key": "naive",
"name": "A 朴素(无缓存/无压缩)",
"total_cost": 0.01574031350707,
"component_costs": {
"uncached_input_cost": 0.0118080062343888,
"cached_input_cost": 0.0,
"output_cost": 0.0039323072726812,
"uncached_input_tokens": 19974,
"cached_input_tokens": 0,
"output_tokens": 1267,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.00196753918838375,
"p50": 0.0019544041559152,
"p95": 0.0026135574027031996,
"p99": 0.0026135574027031996,
"max": 0.0026135574027031996
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1015,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 5.299084663391113,
"response_id": "chatcmpl-6a6a7cb80695ae5455eab84a",
"response_model": "kimi-k2.5",
"response_created": 1785363640
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1706,
"cached_tokens": 0,
"completion_tokens": 147,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 5.079176902770996,
"response_id": "chatcmpl-6a6a7cbd80925af8edb8bb5e",
"response_model": "kimi-k2.5",
"response_created": 1785363646
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2081,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 3.0797362327575684,
"response_id": "chatcmpl-6a6a7cc24511d264b10f37dd",
"response_model": "kimi-k2.5",
"response_created": 1785363651
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2466,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1482,
"latency_s": 3.4901974201202393,
"response_id": "chatcmpl-6a6a7cc5edbcdbb3613f3353",
"response_model": "kimi-k2.5",
"response_created": 1785363654
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2753,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1585,
"latency_s": 6.93122410774231,
"response_id": "chatcmpl-6a6a7cc97f6a7380b55ebedb",
"response_model": "kimi-k2.5",
"response_created": 1785363658
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 3048,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1693,
"latency_s": 5.43744421005249,
"response_id": "chatcmpl-6a6a7cd0d7374244555b9001",
"response_model": "kimi-k2.5",
"response_created": 1785363665
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3324,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1779,
"latency_s": 5.579896926879883,
"response_id": "chatcmpl-6a6a7cd56f58f5f62fa2981a",
"response_model": "kimi-k2.5",
"response_created": 1785363670
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3581,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1847,
"latency_s": 5.334557056427002,
"response_id": "chatcmpl-6a6a7cdb2e6a18d37452de21",
"response_model": "kimi-k2.5",
"response_created": 1785363676
}
]
},
{
"key": "kv",
"name": "KV 仅缓存(稳定前缀/不压缩)",
"total_cost": 0.0082537957670069,
"component_costs": {
"uncached_input_cost": 0.0027956374027448,
"cached_input_cost": 0.0015289547279445,
"output_cost": 0.0039292036363176,
"uncached_input_tokens": 4729,
"cached_input_tokens": 14779,
"output_tokens": 1266,
"tool_ctx_tokens": 10749
},
"cost_distribution": {
"n": 8,
"mean": 0.0010317244708758625,
"p50": 0.001053935792344,
"p95": 0.0011930969351368,
"p99": 0.0011930969351368,
"max": 0.0011930969351368
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 2.981578826904297,
"response_id": "chatcmpl-6a6a7d043905097c262cd2d8",
"response_model": "kimi-k2.5",
"response_created": 1785363716
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1652,
"cached_tokens": 512,
"completion_tokens": 146,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 8.04120397567749,
"response_id": "chatcmpl-6a6a7d07d5eb4229617017c7",
"response_model": "kimi-k2.5",
"response_created": 1785363719
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2023,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 6.327085256576538,
"response_id": "chatcmpl-6a6a7d0f3905097c262cd2f5",
"response_model": "kimi-k2.5",
"response_created": 1785363727
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2408,
"cached_tokens": 1792,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1482,
"latency_s": 4.883957862854004,
"response_id": "chatcmpl-6a6a7d15bb40cba1f65cc486",
"response_model": "kimi-k2.5",
"response_created": 1785363733
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2696,
"cached_tokens": 2048,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1585,
"latency_s": 5.0474817752838135,
"response_id": "chatcmpl-6a6a7d1adc87cefde7ebcca2",
"response_model": "kimi-k2.5",
"response_created": 1785363738
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2986,
"cached_tokens": 2560,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1693,
"latency_s": 3.6142687797546387,
"response_id": "chatcmpl-6a6a7d1f3905097c262cd332",
"response_model": "kimi-k2.5",
"response_created": 1785363743
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3266,
"cached_tokens": 2816,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1779,
"latency_s": 3.981088161468506,
"response_id": "chatcmpl-6a6a7d233905097c262cd33e",
"response_model": "kimi-k2.5",
"response_created": 1785363747
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3522,
"cached_tokens": 3072,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1847,
"latency_s": 5.233341217041016,
"response_id": "chatcmpl-6a6a7d27d5eb42296170182d",
"response_model": "kimi-k2.5",
"response_created": 1785363751
}
]
},
{
"key": "compress",
"name": "仅压缩(前缀不稳定/摘要)",
"total_cost": 0.0125982511692596,
"component_costs": {
"uncached_input_cost": 0.0089390638965752,
"cached_input_cost": 0.0,
"output_cost": 0.0036591872726844,
"uncached_input_tokens": 15121,
"cached_input_tokens": 0,
"output_tokens": 1179,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.00157478139615745,
"p50": 0.0016801018182384,
"p95": 0.0018338057143504,
"p99": 0.0018338057143504,
"max": 0.0018338057143504
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1010,
"cached_tokens": 0,
"completion_tokens": 121,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 3.843111276626587,
"response_id": "chatcmpl-6a6a7d2c96a5dde81c3844dd",
"response_model": "kimi-k2.5",
"response_created": 1785363756
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1671,
"cached_tokens": 0,
"completion_tokens": 110,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 3.103135108947754,
"response_id": "chatcmpl-6a6a7d3006fe5c9e72518409",
"response_model": "kimi-k2.5",
"response_created": 1785363760
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2002,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 5.465490102767944,
"response_id": "chatcmpl-6a6a7d3384ead76d6951b7bb",
"response_model": "kimi-k2.5",
"response_created": 1785363763
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2136,
"cached_tokens": 0,
"completion_tokens": 149,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1231,
"latency_s": 3.5718801021575928,
"response_id": "chatcmpl-6a6a7d3880925af8edb8bcdd",
"response_model": "kimi-k2.5",
"response_created": 1785363769
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1916,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 780,
"latency_s": 3.5103132724761963,
"response_id": "chatcmpl-6a6a7d3cfa52c1b646ff684d",
"response_model": "kimi-k2.5",
"response_created": 1785363772
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2022,
"cached_tokens": 0,
"completion_tokens": 159,
"reasoning_tokens": 1,
"tool_ctx_tokens": 677,
"latency_s": 3.597494125366211,
"response_id": "chatcmpl-6a6a7d3f8a024a47783f03bf",
"response_model": "kimi-k2.5",
"response_created": 1785363776
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2102,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 486,
"latency_s": 9.166883945465088,
"response_id": "chatcmpl-6a6a7d43503e5013c8fe6ef8",
"response_model": "kimi-k2.5",
"response_created": 1785363779
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2262,
"cached_tokens": 0,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 499,
"latency_s": 4.3432776927948,
"response_id": "chatcmpl-6a6a7d4c687ea3db1a5a4d6c",
"response_model": "kimi-k2.5",
"response_created": 1785363789
}
]
},
{
"key": "both",
"name": "B 优化(KV缓存+压缩)",
"total_cost": 0.009057608026520501,
"component_costs": {
"uncached_input_cost": 0.0042445922080159995,
"cached_input_cost": 0.0008403612730964999,
"output_cost": 0.003972654545408001,
"uncached_input_tokens": 7180,
"cached_input_tokens": 8123,
"output_tokens": 1280,
"tool_ctx_tokens": 6036
},
"cost_distribution": {
"n": 8,
"mean": 0.0011322010033150626,
"p50": 0.0011570356364336,
"p95": 0.0014060359481248,
"p99": 0.0014060359481248,
"max": 0.0014060359481248
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 955,
"cached_tokens": 955,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 300,
"latency_s": 3.2073097229003906,
"response_id": "chatcmpl-6a6a7d7658a6e8702ecebaee",
"response_model": "kimi-k2.5",
"response_created": 1785363830
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1652,
"cached_tokens": 512,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 899,
"latency_s": 7.249981880187988,
"response_id": "chatcmpl-6a6a7d792d20c6aa1053113f",
"response_model": "kimi-k2.5",
"response_created": 1785363833
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2038,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1164,
"latency_s": 4.833154916763306,
"response_id": "chatcmpl-6a6a7d80f9aa90a3ebfd7fe5",
"response_model": "kimi-k2.5",
"response_created": 1785363841
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2172,
"cached_tokens": 768,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 1231,
"latency_s": 5.391065835952759,
"response_id": "chatcmpl-6a6a7d852d20c6aa10531154",
"response_model": "kimi-k2.5",
"response_created": 1785363846
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 1962,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 780,
"latency_s": 5.175268888473511,
"response_id": "chatcmpl-6a6a7d8b2d20c6aa10531162",
"response_model": "kimi-k2.5",
"response_created": 1785363851
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2063,
"cached_tokens": 1024,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 677,
"latency_s": 5.356127977371216,
"response_id": "chatcmpl-6a6a7d90a183ecb32232c19d",
"response_model": "kimi-k2.5",
"response_created": 1785363856
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2148,
"cached_tokens": 1280,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 486,
"latency_s": 7.17538595199585,
"response_id": "chatcmpl-6a6a7d953905097c262cd43c",
"response_model": "kimi-k2.5",
"response_created": 1785363862
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2313,
"cached_tokens": 1536,
"completion_tokens": 160,
"reasoning_tokens": 1,
"tool_ctx_tokens": 499,
"latency_s": 5.300028085708618,
"response_id": "chatcmpl-6a6a7d9d56e2191f9ae786d6",
"response_model": "kimi-k2.5",
"response_created": 1785363869
}
]
}
]
}
@@ -0,0 +1,426 @@
{
"model": "gpt-4o-mini",
"pricing": {
"input": 0.15,
"cached": 0.075,
"output": 0.6
},
"scenarios": [
{
"key": "naive",
"name": "A 朴素(无缓存/无压缩)",
"total_cost": 0.0037758,
"component_costs": {
"uncached_input_cost": 0.0031049999999999997,
"cached_input_cost": 0.0,
"output_cost": 0.0006708,
"uncached_input_tokens": 20700,
"cached_input_tokens": 0,
"output_tokens": 1118,
"tool_ctx_tokens": 9544
},
"cost_distribution": {
"n": 8,
"mean": 0.000471975,
"p50": 0.00048059999999999997,
"p95": 0.0006462,
"p99": 0.0006462,
"max": 0.0006462
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1113,
"cached_tokens": 0,
"completion_tokens": 104,
"tool_ctx_tokens": 276,
"latency_s": 3.1535720825195312
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1807,
"cached_tokens": 0,
"completion_tokens": 99,
"tool_ctx_tokens": 829,
"latency_s": 2.0884652137756348
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2154,
"cached_tokens": 0,
"completion_tokens": 139,
"tool_ctx_tokens": 1046,
"latency_s": 2.685513973236084
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2564,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1287,
"latency_s": 2.922238826751709
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2863,
"cached_tokens": 0,
"completion_tokens": 136,
"tool_ctx_tokens": 1389,
"latency_s": 2.6890718936920166
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 3123,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1490,
"latency_s": 3.073489189147949
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3408,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1579,
"latency_s": 3.0861823558807373
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3668,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1648,
"latency_s": 2.504854917526245
}
]
},
{
"key": "kv",
"name": "KV 仅缓存(稳定前缀/不压缩)",
"total_cost": 0.0027075,
"component_costs": {
"uncached_input_cost": 0.0010227,
"cached_input_cost": 0.0010176,
"output_cost": 0.0006672,
"uncached_input_tokens": 6818,
"cached_input_tokens": 13568,
"output_tokens": 1112,
"tool_ctx_tokens": 9544
},
"cost_distribution": {
"n": 8,
"mean": 0.0003384375,
"p50": 0.00033465000000000003,
"p95": 0.00040815,
"p99": 0.00040815,
"max": 0.00040815
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1056,
"cached_tokens": 1024,
"completion_tokens": 121,
"tool_ctx_tokens": 276,
"latency_s": 2.1304330825805664
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1763,
"cached_tokens": 0,
"completion_tokens": 117,
"tool_ctx_tokens": 829,
"latency_s": 2.102344036102295
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2130,
"cached_tokens": 1024,
"completion_tokens": 136,
"tool_ctx_tokens": 1046,
"latency_s": 3.537381172180176
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2540,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 1287,
"latency_s": 3.393212080001831
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2834,
"cached_tokens": 2176,
"completion_tokens": 114,
"tool_ctx_tokens": 1389,
"latency_s": 1.906968116760254
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 3081,
"cached_tokens": 2560,
"completion_tokens": 160,
"tool_ctx_tokens": 1490,
"latency_s": 2.7878849506378174
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 3361,
"cached_tokens": 2560,
"completion_tokens": 160,
"tool_ctx_tokens": 1579,
"latency_s": 2.825930118560791
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 3621,
"cached_tokens": 3200,
"completion_tokens": 144,
"tool_ctx_tokens": 1648,
"latency_s": 2.7222893238067627
}
]
},
{
"key": "compress",
"name": "仅压缩(前缀不稳定/摘要)",
"total_cost": 0.00311475,
"component_costs": {
"uncached_input_cost": 0.00242655,
"cached_input_cost": 0.0,
"output_cost": 0.0006882,
"uncached_input_tokens": 16177,
"cached_input_tokens": 0,
"output_tokens": 1147,
"tool_ctx_tokens": 5248
},
"cost_distribution": {
"n": 8,
"mean": 0.00038934375,
"p50": 0.0004059,
"p95": 0.0004497,
"p99": 0.0004497,
"max": 0.0004497
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1112,
"cached_tokens": 0,
"completion_tokens": 127,
"tool_ctx_tokens": 276,
"latency_s": 2.223604917526245
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1829,
"cached_tokens": 0,
"completion_tokens": 91,
"tool_ctx_tokens": 829,
"latency_s": 2.916440963745117
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2164,
"cached_tokens": 0,
"completion_tokens": 129,
"tool_ctx_tokens": 1046,
"latency_s": 2.543590784072876
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2310,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1052,
"latency_s": 2.999290943145752
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2066,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 635,
"latency_s": 5.8334801197052
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2146,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 551,
"latency_s": 2.787147045135498
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2192,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 430,
"latency_s": 3.470608949661255
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2358,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 429,
"latency_s": 2.5834107398986816
}
]
},
{
"key": "both",
"name": "B 优化(KV缓存+压缩)",
"total_cost": 0.00264285,
"component_costs": {
"uncached_input_cost": 0.00148365,
"cached_input_cost": 0.0004608,
"output_cost": 0.0006984000000000001,
"uncached_input_tokens": 9891,
"cached_input_tokens": 6144,
"output_tokens": 1164,
"tool_ctx_tokens": 5248
},
"cost_distribution": {
"n": 8,
"mean": 0.00033035625,
"p50": 0.00033525000000000004,
"p95": 0.00044249999999999997,
"p99": 0.00044249999999999997,
"max": 0.00044249999999999997
},
"spans": [
{
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 1056,
"cached_tokens": 1024,
"completion_tokens": 139,
"tool_ctx_tokens": 276,
"latency_s": 2.405411958694458
},
{
"step": "turn-2",
"tool": "query_logistics",
"kind": "llm",
"prompt_tokens": 1781,
"cached_tokens": 0,
"completion_tokens": 112,
"tool_ctx_tokens": 829,
"latency_s": 2.1826939582824707
},
{
"step": "turn-3",
"tool": "check_refund_policy",
"kind": "llm",
"prompt_tokens": 2143,
"cached_tokens": 1024,
"completion_tokens": 151,
"tool_ctx_tokens": 1046,
"latency_s": 2.5647878646850586
},
{
"step": "turn-4",
"tool": "query_knowledge_base",
"kind": "llm",
"prompt_tokens": 2310,
"cached_tokens": 0,
"completion_tokens": 160,
"tool_ctx_tokens": 1052,
"latency_s": 3.0318493843078613
},
{
"step": "turn-5",
"tool": "query_user_history",
"kind": "llm",
"prompt_tokens": 2060,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 635,
"latency_s": 2.475983142852783
},
{
"step": "turn-6",
"tool": "issue_refund",
"kind": "llm",
"prompt_tokens": 2143,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 551,
"latency_s": 2.7105050086975098
},
{
"step": "turn-7",
"tool": "send_notification",
"kind": "llm",
"prompt_tokens": 2188,
"cached_tokens": 1024,
"completion_tokens": 160,
"tool_ctx_tokens": 430,
"latency_s": 2.7884562015533447
},
{
"step": "turn-8",
"tool": "close_ticket",
"kind": "llm",
"prompt_tokens": 2354,
"cached_tokens": 1024,
"completion_tokens": 122,
"tool_ctx_tokens": 429,
"latency_s": 2.3781700134277344
}
]
}
]
}
@@ -0,0 +1,9 @@
"""Test import bootstrap for the agent-cost-analysis experiment."""
from pathlib import Path
import sys
EXPERIMENT_ROOT = Path(__file__).resolve().parents[1]
if str(EXPERIMENT_ROOT) not in sys.path:
sys.path.insert(0, str(EXPERIMENT_ROOT))
@@ -0,0 +1,44 @@
"""Tracer.chat must tolerate response.usage == None (OpenAI-compatible providers)."""
from types import SimpleNamespace
import config
from tracer import Tracer
class _FakeClient:
def __init__(self, usage):
self.chat = SimpleNamespace(
completions=SimpleNamespace(create=self._create)
)
self._usage = usage
def _create(self, **_kwargs):
return SimpleNamespace(usage=self._usage)
def test_chat_tolerates_null_usage():
tr = Tracer(_FakeClient(None), pricing=config.default_pricing())
resp = tr.chat(step="turn-1", tool="query_order", model="m", messages=[])
assert resp.usage is None
assert len(tr.spans) == 1
s = tr.spans[0]
assert s.prompt_tokens == 0
assert s.completion_tokens == 0
assert s.cost_usd == 0.0
assert s.latency_s >= 0.0
def test_chat_keeps_real_usage():
usage = SimpleNamespace(
prompt_tokens=100,
completion_tokens=20,
prompt_tokens_details=SimpleNamespace(cached_tokens=10),
)
tr = Tracer(_FakeClient(usage), pricing=config.default_pricing())
tr.chat(step="turn-1", tool="query_order", model="m", messages=[])
s = tr.spans[0]
assert s.prompt_tokens == 100
assert s.completion_tokens == 20
assert s.cached_tokens == 10
assert s.cost_usd > 0
@@ -0,0 +1,96 @@
"""
Regression tests for offline trace parsing (实验 7-9 成本分析).
Covers two crash classes found in --offline mode:
- Tracer.from_records: trace JSON with explicit null token fields -> int(None) TypeError
- demo.collect_offline: scenario dict missing the optional "spans" key -> KeyError
"""
import json
import pytest
import config
import demo
from tracer import Tracer
def _span(**overrides):
span = {
"step": "turn-1",
"tool": "query_order",
"kind": "llm",
"prompt_tokens": 100,
"cached_tokens": 10,
"completion_tokens": 12,
"tool_ctx_tokens": 50,
"latency_s": 1.2,
}
span.update(overrides)
return span
def test_from_records_tolerates_null_fields():
"""Explicit JSON nulls in numeric fields are coerced, not int(None) TypeError."""
records = [_span(prompt_tokens=None, cached_tokens=None,
completion_tokens=None, tool_ctx_tokens=None, latency_s=None)]
tr = Tracer.from_records(records, pricing=config.default_pricing())
s = tr.spans[0]
assert s.prompt_tokens == 0
assert s.cached_tokens == 0
assert s.completion_tokens == 0
assert s.tool_ctx_tokens == -1 # null tool_ctx 视为「未知」
assert s.latency_s == 0.0
def test_from_records_tolerates_missing_fields():
"""Minimal span dicts (only step/tool) still parse."""
tr = Tracer.from_records([{"step": "turn-1", "tool": "query_order"}],
pricing=config.default_pricing())
assert tr.spans[0].prompt_tokens == 0
assert tr.spans[0].tool_ctx_tokens == -1
def test_from_records_keeps_real_values():
"""Normal values pass through unchanged (no coercion side effects)."""
tr = Tracer.from_records([_span()], pricing=config.default_pricing())
s = tr.spans[0]
assert (s.prompt_tokens, s.cached_tokens, s.completion_tokens) == (100, 10, 12)
assert s.tool_ctx_tokens == 50
assert s.latency_s == 1.2
def test_from_records_keeps_zero_tool_ctx_tokens():
"""Explicit 0 means known-zero tool context, not unknown (-1)."""
tr = Tracer.from_records([_span(tool_ctx_tokens=0)],
pricing=config.default_pricing())
assert tr.spans[0].tool_ctx_tokens == 0
assert tr.total_tool_ctx_tokens() == 0
def _write_trace(tmp_path, scenarios):
path = tmp_path / "trace.json"
path.write_text(json.dumps({"model": "gpt-5.6-luna", "scenarios": scenarios}),
encoding="utf-8")
return str(path)
def test_collect_offline_skips_scenario_without_spans(tmp_path, capsys):
"""A scenario missing 'spans' is skipped with a warning, not a KeyError crash."""
trace = _write_trace(tmp_path, [
{"key": "naive", "name": "A naive"}, # no spans -> skip
{"key": "both", "name": "B both", "spans": [_span()]}, # valid
])
tracers = demo.collect_offline(["naive", "both"], config.default_pricing(), trace)
assert [k for k, _ in tracers] == ["both"]
assert "缺少 spans" in capsys.readouterr().err
def test_collect_offline_exits_when_no_usable_scenario(tmp_path):
"""When every selected scenario lacks spans, exit cleanly like the empty-trace path."""
trace = _write_trace(tmp_path, [{"key": "naive", "name": "A naive"}])
with pytest.raises(SystemExit):
demo.collect_offline(["naive"], config.default_pricing(), trace)
if __name__ == "__main__":
pytest.main([__file__, "-v"])
+284
View File
@@ -0,0 +1,284 @@
"""
自建的轻量级 tracing / 可观测系统。
设计沿用分布式追踪的 span 树模型(见书 6.x「Agent 的可观测性」):
- 一次 agent 任务 = 一条 Trace
- 每次 LLM 调用 / 工具调用 = 一个 Span
- Span 记录:所属步骤、类型、token 用量(prompt/completion/cached)、时延、成本
用法:
tracer = Tracer(client)
resp = tracer.chat(step="turn-1", tool="query_order",
model=..., messages=..., temperature=0)
tracer.print_breakdown() # 打印按步骤/工具聚合的成本拆解
离线复用(不打模型、只算成本):
tracer = Tracer.from_records(records, pricing=..., name=...)
# records 里是此前真实运行录下的每一步 token 用量(canned token counts
"""
import math
import time
from dataclasses import dataclass, asdict, field
from typing import List, Optional
from config import Pricing, default_pricing
def _percentile(values: List[float], q: float) -> float:
"""最近秩(nearest-rank)百分位,避免引入 numpy 依赖。q 取 0~100。"""
if not values:
return 0.0
xs = sorted(values)
if len(xs) == 1:
return xs[0]
# 最近秩 = ceil(q/100 * N)。不能用 int(round(x + 0.5)):当 q/100*N 恰为
# 整数 k 时,round(k + 0.5) 的银行家舍入会得到 k+1(如 n=100、q=99 时
# 秩变成 100,把 p99 报成最大值)。
rank = max(1, min(len(xs), math.ceil(q / 100.0 * len(xs))))
return xs[rank - 1]
@dataclass
class Span:
"""一次被追踪的调用(这里主要是 LLM 调用)。"""
step: str # 逻辑步骤名,如 "turn-2"
tool: str # 该步骤关联的工具/动作名,用于归因“哪一步最贵”
kind: str = "llm" # span 类型:llm / tool
prompt_tokens: int = 0
cached_tokens: int = 0
completion_tokens: int = 0
reasoning_tokens: int = 0
# 该轮输入里「工具返回结果」占用的累计 token(同一份工具返回会在后续每轮被反复计费)。
# 由上层用 tokenizer 估算并填入;离线复用时从 records 读回。-1 表示未知。
tool_ctx_tokens: int = -1
latency_s: float = 0.0
cost_usd: float = 0.0
response_id: str = ""
response_model: str = ""
response_created: int = 0
@property
def total_tokens(self) -> int:
return self.prompt_tokens + self.completion_tokens
@property
def uncached_prompt_tokens(self) -> int:
return max(self.prompt_tokens - self.cached_tokens, 0)
class Tracer:
"""包裹 OpenAI client,自动记录每次 LLM 调用的 usage / 时延 / 成本。"""
def __init__(self, client=None, name: str = "trace",
pricing: Optional[Pricing] = None):
self.client = client
self.name = name
self.pricing = pricing or default_pricing()
self.spans: List[Span] = []
# ---------- 采集 ----------
def chat(self, step: str, tool: str, tool_ctx_tokens: int = -1, **kwargs):
"""发起一次被追踪的 chat.completions 调用。
kwargs 原样透传给 openai clientmodel / messages / temperature 等)。
tool_ctx_tokens:本轮输入里工具返回结果占用的累计 token(可选,用于成本归因)。
返回原始的 OpenAI response 对象,方便上层取 content。
"""
t0 = time.time()
resp = self.client.chat.completions.create(**kwargs)
latency = time.time() - t0
usage = resp.usage
# Some OpenAI-compatible providers omit usage (null); match from_records coercion.
if usage is None:
span = Span(
step=step,
tool=tool,
kind="llm",
tool_ctx_tokens=tool_ctx_tokens,
latency_s=latency,
cost_usd=0.0,
)
self.spans.append(span)
return resp
# cached_tokens 藏在 prompt_tokens_details 里,注意做防御式读取
cached = 0
details = getattr(usage, "prompt_tokens_details", None)
if details is not None:
cached = getattr(details, "cached_tokens", 0) or 0
prompt_tokens = int(getattr(usage, "prompt_tokens", 0) or 0)
completion_tokens = int(getattr(usage, "completion_tokens", 0) or 0)
completion_details = getattr(usage, "completion_tokens_details", None)
reasoning_tokens = int(getattr(completion_details, "reasoning_tokens", 0) or 0)
span = Span(
step=step,
tool=tool,
kind="llm",
prompt_tokens=prompt_tokens,
cached_tokens=cached,
completion_tokens=completion_tokens,
reasoning_tokens=reasoning_tokens,
tool_ctx_tokens=tool_ctx_tokens,
latency_s=latency,
cost_usd=self.pricing.cost_usd(
prompt_tokens, cached, completion_tokens),
response_id=str(getattr(resp, "id", "") or ""),
response_model=str(getattr(resp, "model", "") or ""),
response_created=int(getattr(resp, "created", 0) or 0),
)
self.spans.append(span)
return resp
# ---------- 离线复用(canned token counts → 重新计成本)----------
@classmethod
def from_records(cls, records: List[dict], name: str = "trace",
pricing: Optional[Pricing] = None) -> "Tracer":
"""用此前录下的 token 用量重建一条 trace,并按给定单价重算成本(不打模型)。"""
tr = cls(client=None, name=name, pricing=pricing)
for r in records:
span = Span(
step=r.get("step", ""),
tool=r.get("tool", ""),
kind=r.get("kind", "llm"),
# Missing/null tool_ctx → -1 (unknown); keep explicit 0 (known-zero).
prompt_tokens=int(r.get("prompt_tokens") or 0),
cached_tokens=int(r.get("cached_tokens") or 0),
completion_tokens=int(r.get("completion_tokens") or 0),
reasoning_tokens=int(r.get("reasoning_tokens") or 0),
tool_ctx_tokens=(-1 if r.get("tool_ctx_tokens") is None
else int(r.get("tool_ctx_tokens"))),
latency_s=float(r.get("latency_s") or 0.0),
response_id=str(r.get("response_id") or ""),
response_model=str(r.get("response_model") or ""),
response_created=int(r.get("response_created") or 0),
)
span.cost_usd = tr.pricing.cost_usd(
span.prompt_tokens, span.cached_tokens, span.completion_tokens)
tr.spans.append(span)
return tr
def to_records(self) -> List[dict]:
"""导出每一步的原始 token 用量(用于落盘成 canned trace,供离线复用)。"""
out = []
for s in self.spans:
d = asdict(s)
d.pop("cost_usd", None) # 成本由单价重算,不落盘固定值
out.append(d)
return out
# ---------- 聚合 ----------
def total_cost(self) -> float:
return sum(s.cost_usd for s in self.spans)
def total_prompt_tokens(self) -> int:
return sum(s.prompt_tokens for s in self.spans)
def total_cached_tokens(self) -> int:
return sum(s.cached_tokens for s in self.spans)
def total_completion_tokens(self) -> int:
return sum(s.completion_tokens for s in self.spans)
def total_uncached_prompt_tokens(self) -> int:
return sum(s.uncached_prompt_tokens for s in self.spans)
def total_tool_ctx_tokens(self) -> int:
return sum(s.tool_ctx_tokens for s in self.spans if s.tool_ctx_tokens >= 0)
def total_latency(self) -> float:
return sum(s.latency_s for s in self.spans)
def cache_rate(self) -> float:
pin = self.total_prompt_tokens()
return self.total_cached_tokens() / pin if pin else 0.0
def component_costs(self) -> dict:
"""把总成本拆成三个成本构成要素(对应书「成本的构成要素」):
- 未缓存输入 / 缓存输入 / 输出
以及输入侧里「工具返回注入」token 占比(若已知)。"""
p = self.pricing
uncached_in = self.total_uncached_prompt_tokens()
cached_in = self.total_cached_tokens()
out = self.total_completion_tokens()
return {
"uncached_input_cost": uncached_in / 1_000_000 * p.input_per_m,
"cached_input_cost": cached_in / 1_000_000 * p.cached_per_m,
"output_cost": out / 1_000_000 * p.output_per_m,
"uncached_input_tokens": uncached_in,
"cached_input_tokens": cached_in,
"output_tokens": out,
"tool_ctx_tokens": self.total_tool_ctx_tokens(),
}
def cost_distribution(self) -> dict:
"""按步骤的单步成本分布(p50/p95/p99)。对应书「成本分布 p50/p95/p99」。"""
costs = [s.cost_usd for s in self.spans]
n = len(costs)
return {
"n": n,
"mean": (sum(costs) / n) if n else 0.0,
"p50": _percentile(costs, 50),
"p95": _percentile(costs, 95),
"p99": _percentile(costs, 99),
"max": max(costs) if costs else 0.0,
}
# ---------- 打印 ----------
def print_breakdown(self, title: Optional[str] = None):
"""打印一次 agent 任务的按步骤成本拆解表,并指出最贵的一步、成本构成与分布。"""
print()
print(f"===== 成本拆解: {title or self.name} =====")
header = (
f"{'步骤':<8} {'工具/动作':<20} {'输入tok':>8} {'缓存tok':>8} "
f"{'工具tok':>8} {'输出tok':>8} {'时延(s)':>8} {'成本($)':>12}"
)
print(header)
print("-" * len(header))
for s in self.spans:
tctx = s.tool_ctx_tokens if s.tool_ctx_tokens >= 0 else "-"
print(
f"{s.step:<8} {s.tool:<20} {s.prompt_tokens:>8} {s.cached_tokens:>8} "
f"{str(tctx):>8} {s.completion_tokens:>8} {s.latency_s:>8.2f} "
f"{s.cost_usd:>12.6f}"
)
print("-" * len(header))
tctx_total = self.total_tool_ctx_tokens() if any(
s.tool_ctx_tokens >= 0 for s in self.spans) else "-"
print(
f"{'合计':<8} {'':<20} {self.total_prompt_tokens():>8} "
f"{self.total_cached_tokens():>8} {str(tctx_total):>8} "
f"{self.total_completion_tokens():>8} "
f"{self.total_latency():>8.2f} {self.total_cost():>12.6f}"
)
# 归因:哪一步最贵
if self.spans:
worst = max(self.spans, key=lambda s: s.cost_usd)
total = self.total_cost()
share = worst.cost_usd / total * 100 if total else 0
print(
f"\n最贵的一步 → {worst.step} / {worst.tool}: "
f"${worst.cost_usd:.6f}(占总成本 {share:.1f}%"
)
# 成本构成拆解(未缓存输入 / 缓存输入 / 输出)
comp = self.component_costs()
total = self.total_cost() or 1e-12
print("成本构成:")
print(f" 未缓存输入 {comp['uncached_input_tokens']:>8} tok "
f"${comp['uncached_input_cost']:.6f} ({comp['uncached_input_cost']/total*100:.1f}%)")
print(f" 缓存输入 {comp['cached_input_tokens']:>8} tok "
f"${comp['cached_input_cost']:.6f} ({comp['cached_input_cost']/total*100:.1f}%)")
print(f" 输出 {comp['output_tokens']:>8} tok "
f"${comp['output_cost']:.6f} ({comp['output_cost']/total*100:.1f}%)")
if comp["tool_ctx_tokens"] > 0:
print(f" 其中「工具返回注入」累计输入 {comp['tool_ctx_tokens']} tok "
f"(同一份工具返回在后续每轮被反复计费)")
# 单步成本分布
dist = self.cost_distribution()
print(f"单步成本分布(n={dist['n']}): 均值 ${dist['mean']:.6f} "
f"p50 ${dist['p50']:.6f} p95 ${dist['p95']:.6f} p99 ${dist['p99']:.6f}")
+426
View File
@@ -0,0 +1,426 @@
# AndroidWorld T3A Evaluation Notes / AndroidWorld T3A 评估分析笔记
> Companion material for *AI Agents in Depth*, Chapter 7 — **Experiment 7-12: Evaluate and improve on AndroidWorld**.
> 配套《深入理解 AI Agent》第 7 章 **实验 7-12 ★★★:AndroidWorld 的评估和改进**。
← [Chapter 7 index / 返回第 7 章目录](../README.md) · 📖 [Read the chapter / 读本章正文](../../book/chapter7.md)[EN](../../book-en/chapter7.md)
---
## English
### What this directory is
This folder is **not** a copy of the [AndroidWorld](https://github.com/google-research/android_world) benchmark codebase. It contains **evaluation artifacts and analysis notes** for a **T3A** (Text-only / accessibility-tree style mobile agent) run plus a companion runner that executes the book's full **diagnose → hypothesize → experiment → decide → iterate** loop against a separate, unmodified upstream checkout.
| Path | Role |
| --- | --- |
| [`t3a_summary.md`](t3a_summary.md) | High-level report: per-task outcomes + capability-tag × difficulty matrix, strengths/weaknesses |
| [`t3a_failed_analysis.md`](t3a_failed_analysis.md) | Failure taxonomy with root-cause write-ups (transcription, complex UI, math/counting, etc.) |
| [`t3a.md`](t3a.md) | Full step traces for runs (including successes): per-step `Action` / `Reason` / `Summary` records |
| [`t3a_failed.md`](t3a_failed.md) | Step traces focused on failed tasks (useful for root-cause replay) |
| [`experiment_core.py`](experiment_core.py) | Evidence aggregation, success/cost decisions, strict completion gates, and five-stage report rendering |
| [`run_controlled_experiment.py`](run_controlled_experiment.py) | Real AndroidWorld control/treatment and candidate-rerun runner; no mock fallback |
| [`merge_candidate_shards.py`](merge_candidate_shards.py) | Strict merger for independent trial shards; rejects overlap, provenance drift, missing reference setup, and duplicate episodes |
| [`test_experiment.py`](test_experiment.py) | Offline checks for redaction, cost decisions, and non-overclaiming gates |
| [`requirements.txt`](requirements.txt) | Installs the adjacent upstream checkout plus the OpenAI-compatible API client |
| `validation/` | Machine-readable real-run evidence and the reports generated from it |
To execute the controlled loop, first clone and configure upstream AndroidWorld (see [Reproduce the benchmark](#reproduce-the-benchmark-optional) below). The large `t3a*.md` files remain reading/analysis inputs; the runner and `validation/` artifacts are the executable evidence layer.
### Background: AndroidWorld + T3A
- **AndroidWorld** evaluates agents that complete real tasks on Android apps (navigation, UI interaction, multi-app flows). Tasks are often **parameterized templates** (anti-contamination, diverse instances) and are scored by **final UI / environment state**, not by matching a fixed action sequence.
- The notes here analyze a **T3A** agent run (logged as `t3a_claude4_sonnet` in the summary tables): the agent plans from UI state (accessibility tree / similar structured observations) and issues discrete actions (`open_app`, `click`, `status`, …).
### Snapshot results (from the included report)
Numbers below come from [`t3a_summary.md`](t3a_summary.md) (116 tasks, one trial each; agent `t3a_claude4_sonnet`, run on 2025-07-02):
| Metric | Value (approx.) |
| --- | --- |
| Overall success rate | **~88%** |
| Fail rate | **~12%** |
| Mean episode length (successful) | **~13.5** steps |
**Where it succeeds:** structured, linear flows—camera/clock/contacts, file ops, Markor notes, many system toggles, multi-app and short-term memorization on easier tags.
**Where it fails (clustered):** SMS reply edge cases, Wi-Fi / combined connectivity, Tasks app queries, VLC playlists, and tasks needing **transcription**, **math/counting**, **complex UI understanding**, **information retrieval**, or **requires_setup**.
### Capability portrait
From the tag × difficulty matrix in the summary:
| Strengths | Critical weaknesses |
| --- | --- |
| `multi_app`, `memorization` (easy ~1.0) | `transcription` (~0.0) |
| Decent `search` on medium | `math_counting` (easy ~0.0) |
| Reliable on standard UI flows | `complex_ui_understanding`, `information_retrieval` (very low) |
| | `requires_setup` (easy ~0.0) |
**One-line portrait:** a strong “operator” on standard linear tasks; weak as a “thinker” when deep vision, counting, non-standard UI, or fragile multi-step state is required.
### Failure categories (see detailed analysis)
Condensed from [`t3a_failed_analysis.md`](t3a_failed_analysis.md):
1. **Transcription** — Navigates gallery/VLC correctly but cannot OCR image/video text; may invent plausible data and “fake success.”
2. **Complex UI** — Sees widgets but lacks a mental model of control logic (e.g. timer digit entry loops after detecting invalid `63s`).
3. **App first-run overhead** — Tutorials / permission wizards burn step budget before the real goal.
4. **Math / counting** — Can scroll and “see” list items but fails to filter + count or sum durations under step limits.
5. **Retrieval + planning** — Dense UIs (calendar grid), multi-delete with state tracking; inefficient recovery (day-by-day instead of reselecting).
Many failures surface as **max steps** (`Agent did not indicate task is done. Reached max number of steps.`)—symptom of loops, inefficient recovery, or missing perception, not merely “too few steps.”
### How to use this material (Experiment 7-12)
Follow the books five-step loop:
1. **Diagnose** — Cross the per-task table with the capability matrix; map surface failures to capability gaps.
2. **Hypothesize** — Layered ideas (surface → mid → deep), e.g. settings navigation hints, fix multimodal input pipe, add UI tree + screenshot, stronger vision model, conditional thinking for count tasks.
3. **Experiment** — Cheap ablations first; measure success **and** latency/cost side effects.
4. **Decide** — Deploy high ROI fixes; reject global “always think” if only a small tag set benefits.
5. **Iterate** — Re-run the suite; new residual failures become the next report.
### Executed controlled loop (2026-07-29 to 2026-08-04)
The companion runner now makes the book's loop executable while leaving the adjacent upstream checkout unmodified. It records the real AndroidWorld evaluator reward, explicit agent termination, actions, steps, wall time, LLM calls, token use, estimated token cost, exact model/runtime provenance, and the installed version of every required app after every episode. A bounded final analysis is requested from the same real configured LLM; the JSON evidence, not that prose, remains authoritative.
The first low-cost phase tested **H1**, a Wi-Fi navigation/state-verification guideline, against the untouched upstream T3A prompt. Its four matched task pairs completed with no runtime errors:
| Phase 1 result | Control | H1 treatment |
| --- | ---: | ---: |
| Successful episodes | 1 / 4 | 1 / 4 |
| Mean evaluator reward | 0.50 | 0.50 |
| Mean latency | 233.47 s | 156.98 s |
| Input + output tokens | 442,619 | 210,039 |
H1 reduced observed latency and token use but produced **no paired success gain**, so it was not promoted. See [phase-1 evidence](validation/paired_wifi_api35_20260729/evidence.json) and its [report](validation/paired_wifi_api35_20260729/report.md).
The residual traces exposed an API-35 observation issue: AndroidWorld's gRPC accessibility feed often returned only status-bar elements after opening the Internet panel, while an independent UIAutomator dump showed the full real Settings hierarchy. **H5** therefore tests a middle-layer input-pipeline change: upstream's `A11yMethod.UIAUTOMATOR` versus the gRPC forwarder, with the same base T3A prompt in both arms. This is an AndroidWorld-supported observation path selected from the companion runner, not an edit to upstream source.
H5 recovered the four-task slice from `1/4` control successes to `4/4` UIAutomator successes with no paired regression and a `0.788×` latency ratio. It was still restricted because its `2.498×` mean-token ratio exceeded the `1.5×` guardrail. The resulting cost-refinement hypothesis **H5C** keeps real UIAutomator observations/actions/evaluators but filters non-semantic container elements before T3A formats the prompt.
The completed H5C paired run preserved `4/4` successes in both arms. Compact UIAutomator used `70,557.5` mean tokens versus `139,439.5` for raw UIAutomator (`0.506×`) and `99.18s` versus `101.20s` mean latency (`0.980×`). It therefore passed the stricter H5C subset gate and became eligible only for a full-suite candidate rerun. At that stage it was **not** deployment approval and did not complete Experiment 7-12's 116-task × five-seed requirement. See the [H5C evidence](validation/paired_h5c_compact_api35_20260729/evidence.json) and [report](validation/paired_h5c_compact_api35_20260729/report.md).
The final reference-environment campaign subsequently completed all five gates: 580/580 unique episodes, 116 tasks × trials 15, zero runtime errors, official setup completed, and the same 24/24 required package versions on every Pixel 6/API-33 shard. The canonical [merged evidence](validation/candidate_h5c_api33_local_qwen_20260804/evidence.json) and [generated report](validation/candidate_h5c_api33_local_qwen_20260804/report.md) record:
| Full candidate result | Value |
| --- | ---: |
| Strict T3A successes | 26 / 580 (`4.4828%`) |
| Evaluator rewards | 77 full (`1.0`) + 1 partial (`0.5`) |
| Mean evaluator reward | `0.133621` |
| Mean steps / LLM calls | `9.672414` / `18.998276` |
| Mean latency | `109.860845s` |
| Mean tokens | `169,069.563793` |
| Total input / output tokens | `97,384,410` / `675,937` |
| Estimated API cost | `$0.00` (local inference) |
Strict success follows the upstream minimal-runner rule: the final evaluator state must equal `1.0` **and** the agent must explicitly declare completion. This is why the 26 strict successes are fewer than the 77 full-reward final states; one additional episode received partial reward `0.5`. Evaluator failures were retained as experimental outcomes; they were not rerun. The merged evidence sets `scope.direct_episode_gate_completed`, `scope.full_suite_completed`, `scope.manuscript_five_seed_gate_completed`, and `experiment_complete` to `true`, but `decision.deployment_approved` remains `false` because the observed result is poor and there is no valid full-suite control comparison.
The candidate used local `qwen2.5-7b-instruct-local` revision `a09a35458c702b33eeacc393d103063234e8bc28`, served by vLLM 0.19.0 on an NVIDIA RTX PRO 6000 Blackwell 96 GB. The H5C paired source used `doubao-seed-1-6-250615`. Consequently, this campaign completes the direct execution/evidence requirement for the promoted observation treatment, but it is **not** a same-model extension and establishes neither comparative uplift nor noninferiority.
The following compatibility treatments are explicitly part of the result boundary:
1. `ContactsNewContactDraft`: UIAutomator does not populate `state.forest`, so `state.ui_elements` is passed to the unchanged official contact predicate.
2. Clipper foreground race: the unchanged clipboard get/set operation is retried once after one second only for the exact documented foreground-access error.
3. `SimpleSmsReplyMostRecent`: the inbox is polled for five additional seconds; if emulator-console injection still leaves it empty, the exact last injected address/body is inserted into the same SMS SQLite database that upstream already clears, then the unchanged evaluator query runs.
4. `RetroPlayingQueue`: only the exact missing `playing_queue` table error from the pinned APK maps to an empty observed queue; the unchanged exact-queue predicate then records evaluator failure.
5. Native 32,768-token context overflow: only after a real provider context error, a deterministic retry retains at most 12,000 characters from the ends of the action-selection UI description or 6,000 from each before/after summary UI description. The goal, history, action, reason, guidance, output format, original retained UI indices, and per-episode truncation/removal counters remain intact. There were 63 such truncations, removing 7,390,498 UI-description characters in total.
6. Runtime-error retries reuse the exact parameters saved in the failed checkpoint, preventing upstream generator drift from changing the task. Completed checkpoints remain canonical when later parameter regeneration drifts, and resume on the same live emulator preserves the completed setup state rather than rerunning setup.
Phase 1 command (shown for reproducibility):
```bash
PYTHONDONTWRITEBYTECODE=1 GRPC_VERBOSITY=ERROR \
python run_controlled_experiment.py \
--mode paired --hypothesis H1 \
--tasks SystemWifiTurnOff,SystemWifiTurnOffVerify,SystemWifiTurnOn,SystemWifiTurnOnVerify \
--trials 1 --seed 42 --model-seed 42 --max-steps 10 \
--transition-pause 0.5 --skip-device-time \
--output-dir validation/paired_wifi_api35_20260729
```
Phase 2 command:
```bash
PYTHONDONTWRITEBYTECODE=1 GRPC_VERBOSITY=ERROR \
python run_controlled_experiment.py \
--mode paired --hypothesis H5 \
--source-phase1-evidence validation/paired_wifi_api35_20260729/evidence.json \
--tasks SystemWifiTurnOff,SystemWifiTurnOffVerify,SystemWifiTurnOn,SystemWifiTurnOnVerify \
--trials 1 --seed 42 --model-seed 42 --max-steps 10 \
--transition-pause 0.5 --skip-device-time \
--output-dir validation/paired_h5_a11y_api35_20260729
```
Cost-refinement command:
```bash
PYTHONDONTWRITEBYTECODE=1 GRPC_VERBOSITY=ERROR \
python run_controlled_experiment.py \
--mode paired --hypothesis H5C \
--source-phase1-evidence validation/paired_wifi_api35_20260729/evidence.json \
--source-phase2-evidence validation/paired_h5_a11y_api35_20260729/evidence.json \
--tasks SystemWifiTurnOff,SystemWifiTurnOffVerify,SystemWifiTurnOn,SystemWifiTurnOnVerify \
--trials 1 --seed 42 --model-seed 42 --max-steps 10 \
--transition-pause 0.5 --skip-device-time \
--output-dir validation/paired_h5c_compact_api35_REPRODUCE
```
The H1/H5 decision gate requires at least four complete pairs, a positive net success delta, no paired regression, and no more than `1.5×` mean latency or token use. H5C instead requires all four compact-treatment pairs to succeed, no regression, at most `1.5×` latency, and at most `0.75×` raw-UIAutomator tokens. Passing either gate permits only a **candidate rerun**, never deployment. Candidate reruns must supply the actual promoted paired-evidence file; a run ID string alone is insufficient. Full experiment completion additionally requires 580 direct candidate records: all 116 tasks × five distinct trial seeds, with no episode error.
For parallel reference-environment execution, keep `--trials 5` and assign one or more 1-based trials with `--trial-indices`; use a distinct `--execution-shard` label for every emulator. Each shard remains incomplete by itself. `merge_candidate_shards.py` accepts completion only when the shard union contains trials 15 exactly once for every task, with the same model, source decision, API-33 environment, completed upstream app setup, and identical required-app versions. Evaluator failures are direct results and are retained; only runtime-error records are eligible for `--resume --retry-errors`.
The historical paired slices used a Pixel 9 Pro API-35 AVD without the complete third-party app bundle; `--skip-device-time` was restricted to those time-independent Wi-Fi evaluators. The full candidate campaign instead uses five isolated Pixel 6/API-33 emulators, the completed upstream setup procedure, and the same 24 required app packages and versions on every shard. The evidence keeps those two environments separate rather than treating the reference-environment candidate run as a same-environment extension of the API-35 slice.
Concrete example trajectories for root-cause practice:
| Example task | File | Lesson |
| --- | --- | --- |
| `ExpenseAddMultipleFromGallery` | failed analysis + `t3a_failed.md` | OCR / multimodal gap; fabricated expenses |
| `ClockTimerEntry` | same | No durable UI model; repeats bad digit sequence |
| `MarkorTranscribeVideo` | same | Video navigation OK, content blind |
| `SportsTracker*Count*` / duration | same | Perception without arithmetic |
| Successful short flows (`CameraTakeVideo`, stopwatch) | `t3a.md` | What “good” step traces look like |
### Directory layout
```text
chapter7/android-world/
├── README.md # This file
├── experiment_core.py # Evidence, decisions, completion gates, report renderer
├── run_controlled_experiment.py # Real AndroidWorld paired/candidate runner
├── merge_candidate_shards.py # Validates and merges parallel trial shards
├── test_experiment.py # Focused offline integrity tests
├── requirements.txt # Adjacent upstream + API client dependency
├── t3a_summary.md # Aggregated metrics + capability matrix
├── t3a_failed_analysis.md # Failure taxonomy & root causes
├── t3a.md # Full (large) run logs
├── t3a_failed.md # Failed-task run logs
└── validation/ # Real evidence.json + generated report.md artifacts
```
### Reproduce the benchmark (optional)
The controlled runner expects a separate adjacent AndroidWorld checkout (the current workspace uses `chapter7/android_world`) plus its configured emulator and model credential. For a clean reproduction:
1. Clone [google-research/android_world](https://github.com/google-research/android_world) (or the fork your course materials specify).
2. Provide an Android emulator / device environment as required by that project.
3. Install the companion requirements in that environment, set the selected provider credential (the default is `ARK_API_KEY`), and run one of the commands above. For the retained local run, the OpenAI-compatible endpoint was `http://127.0.0.1:18111/v1` on the host (`http://host.docker.internal:18111/v1` from the emulator containers). Set `LOCAL_API_KEY` only in the launching process environment; the runner records only its variable name and never persists its value.
4. Provision the exact upstream Pixel 6 / API-33 apps before attempting `--full-suite`; do not use the API-35 Wi-Fi-only deviations for the full benchmark.
Run one trial per isolated emulator, changing `<N>`, ports, and output directory for shards 15:
```bash
python run_controlled_experiment.py \
--android-world-checkout /workspace/android_world \
--mode candidate-rerun --hypothesis H5C \
--source-paired-evidence validation/paired_h5c_compact_api35_20260729/evidence.json \
--full-suite --trials 5 --trial-indices <N> --execution-shard shard-<N>-of-5 \
--seed 42 --model-seed 42 --max-steps 10 --transition-pause 0.5 \
--provider local-vllm --model qwen2.5-7b-instruct-local \
--base-url http://host.docker.internal:18111/v1 --api-key-env LOCAL_API_KEY \
--max-model-tokens 1024 --model-timeout-s 90 --model-retries 2 \
--model-source local_gpu \
--model-revision a09a35458c702b33eeacc393d103063234e8bc28 \
--model-runtime vllm-0.19.0 \
--accelerator NVIDIA_RTX_PRO_6000_Blackwell_96GB \
--perform-emulator-setup \
--output-dir validation/candidate_h5c_api33_local_shard<N>
```
Merge only after all five shards have completed. The merger rejects overlap, missing trials, provenance/setup/app-version drift, runtime errors, and duplicate task/trial keys:
```bash
python merge_candidate_shards.py \
validation/candidate_h5c_api33_local_shard{1,2,3,4,5}/evidence.json \
--source-paired-evidence validation/paired_h5c_compact_api35_20260729/evidence.json \
--output-dir validation/candidate_h5c_api33_local_qwen_20260804
```
Reading order if you only study the notes: **`t3a_summary.md``t3a_failed_analysis.md` → sample episodes in `t3a_failed.md` / `t3a.md`**.
### Related chapter projects
| Project | Relation |
| --- | --- |
| Upstream `android_world` (external) | Runnable benchmark environment |
| [model-benchmark](../model-benchmark/) | API latency / reliability dimensions of “evaluation” |
| [elo-leaderboard](../elo-leaderboard/) | Pairwise ranking instead of absolute task success |
| [public-health-reporting-eval](../public-health-reporting-eval/) | Another structured eval harness in-repo |
---
## 中文
### 本目录是什么
本目录**不是** [AndroidWorld](https://github.com/google-research/android_world) 基准的源码拷贝。它既包含 **T3A** 类移动 Agent 的**评估产物与分析笔记**,也包含一个连接独立、未修改上游 checkout 的配套 runner,用于真实执行书中的完整闭环:**诊断 → 假设 → 实验 → 决策 → 迭代**(对应**实验 7-12**)。
| 路径 | 作用 |
| --- | --- |
| [`t3a_summary.md`](t3a_summary.md) | 总览:逐任务结果 + 能力标签 × 难度矩阵、优势与短板 |
| [`t3a_failed_analysis.md`](t3a_failed_analysis.md) | 失败分类与根因(转录、复杂 UI、数学/计数等) |
| [`t3a.md`](t3a.md) | 完整逐步轨迹(含成功案例):每步记录 `Action` / `Reason` / `Summary` |
| [`t3a_failed.md`](t3a_failed.md) | 失败任务轨迹(适合回放根因) |
| [`experiment_core.py`](experiment_core.py) | 证据聚合、成功/成本决策、严格完成门槛与五阶段报告渲染 |
| [`run_controlled_experiment.py`](run_controlled_experiment.py) | 真实 AndroidWorld 对照/处理与候选重跑 runner;没有 mock fallback |
| [`merge_candidate_shards.py`](merge_candidate_shards.py) | 严格合并独立 trial 分片,并拒绝重叠、来源漂移、缺少参考环境 setup 或重复 episode |
| [`test_experiment.py`](test_experiment.py) | 脱机检查脱敏、成本决策与防止夸大结论的门槛 |
| [`requirements.txt`](requirements.txt) | 安装相邻上游 checkout 与 OpenAI 兼容 API 客户端 |
| `validation/` | 真实运行的机器可读证据及据此生成的报告 |
若要**自己跑**基准,请按上游仓库克隆与配置(见下文[复现基准](#复现基准可选))。本目录以**阅读与分析**为主。
### 背景:AndroidWorld 与 T3A
- **AndroidWorld**:在真实 Android 应用上评测 Agent 的导航、UI 交互与多应用任务。任务多为**参数化模板**(降低泄漏、增加多样性),按**最终 UI / 环境状态**判分,而不是比对固定操作序列。
- 笔记分析的是一次 **T3A** 运行(摘要表中记为 `t3a_claude4_sonnet`):主要依据 UI 状态(无障碍树等结构化观察)规划,并输出离散动作(`open_app``click``status` 等)。
### 结果快照(来自随附报告)
数据摘自 [`t3a_summary.md`](t3a_summary.md)(116 个任务,每任务 1 次 trialAgent 为 `t3a_claude4_sonnet`,运行于 2025-07-02):
| 指标 | 约值 |
| --- | --- |
| 总体成功率 | **~88%** |
| 失败率 | **~12%** |
| 成功任务平均步数 | **~13.5** |
**擅长:** 结构化、线性流程——相机/时钟/联系人、文件操作、Markor 笔记、多数系统开关;在较简单标签上,跨应用与短时记忆表现好。
**短板(失败扎堆):** 短信回复边缘、Wi-Fi/组合连接、Tasks 查询、VLC 播放列表,以及需要**转录**、**数学/计数**、**复杂 UI 理解**、**信息检索**、**requires_setup** 的任务。
### 能力画像
| 优势 | 关键短板 |
| --- | --- |
| `multi_app``memorization`easy ~1.0 | `transcription`~0.0 |
| `search` 在 medium 上较好 | `math_counting`easy ~0.0 |
| 标准 UI 流程稳定 | `complex_ui_understanding``information_retrieval` 很低 |
| | `requires_setup`easy ~0.0 |
**一句话:** 在标准线性任务上是高效的「操作手」;在深度视觉、计数、非标 UI、脆弱多步状态维护上,「思考者」能力明显不足。
### 失败类别(详见分析文)
浓缩自 [`t3a_failed_analysis.md`](t3a_failed_analysis.md)
1. **转录失败** — 图库/VLC 导航正确,但无法 OCR 图/视频文字;可能捏造合理数据「假装成功」。
2. **复杂 UI** — 看得见控件,却没有控件逻辑的心智模型(如计时器输入,发现 `63s` 非法后仍重复错误序列)。
3. **应用首次启动开销** — 教程/权限向导吃掉步数预算。
4. **数学/计数** — 能滚动「看见」列表,却完不成筛选+计数或时长求和。
5. **检索与规划** — 密集日历格、去重删除的状态维护;恢复策略低效(逐天点而不是回月视图重选)。
大量失败以**步数耗尽**呈现(`Reached max number of steps`)——根因往往是循环、低效恢复或感知缺失,而不仅是「上限太小」。
### 如何使用(实验 7-12
按书中五步闭环:
1. **诊断** — 交叉逐任务表与能力矩阵,把表面失败映射到能力缺陷。
2. **假设** — 表层 → 中层 → 深层(如设置导航提示、修复多模态输入管道、截图+UI 树、更强视觉模型、仅对计数任务开思考)。
3. **实验** — 先做低成本对照;同时量成功率与时延/成本副作用。
4. **决策** — 优先部署高 ROI;拒绝为少数标签让全局任务承担数倍延迟/成本。
5. **迭代** — 重跑全集,新失败模式成为下一轮起点。
### 已执行的对照闭环(2026-07-29 至 2026-08-04
配套 runner 会逐 episode 记录真实 AndroidWorld evaluator reward、Agent 是否显式结束、动作、步数、耗时、LLM 调用、token、估算 token 成本、模型/运行时来源,以及每个必需 App 的安装版本。运行结束后,同一个真实配置模型会对聚合证据做受约束分析;JSON 证据始终是权威来源,LLM 文本不能覆盖它。
第一阶段测试低成本表层假设 **H1**:对照组使用原始 T3A prompt,处理组只增加 Wi-Fi 导航和最终状态确认指南。四个配对任务全部正常结束;两组均只成功 `1/4`,平均 evaluator reward 都是 `0.50`。处理组平均延迟由 `233.47s` 降至 `156.98s`,输入+输出 token 由 `442,619` 降至 `210,039`,但**没有配对成功增益**,因此不晋级。证据见 [phase-1 evidence](validation/paired_wifi_api35_20260729/evidence.json) 与 [report](validation/paired_wifi_api35_20260729/report.md)。
残余轨迹暴露了 API 35 观察兼容问题:打开 Internet 面板后,gRPC 无障碍树经常只剩状态栏元素,而独立 UIAutomator dump 能看到完整的真实 Settings 层级。因此第二阶段中层假设 **H5** 对比 gRPC forwarder 与 AndroidWorld 上游已有的 `A11yMethod.UIAUTOMATOR`,两组保持相同原始 T3A prompt、参数、seed、模型和 evaluator。H5 将该四任务切片从对照组 `1/4` 成功提升到 UIAutomator 的 `4/4`,但平均 token 比达到 `2.498×`,超过 `1.5×` 门槛,因此没有晋级。
随后执行的成本优化假设 **H5C** 对比原始 UIAutomator 与过滤非语义容器节点的紧凑 UIAutomator。两组都保持 `4/4` 成功;紧凑组平均 token 从 `139,439.5` 降至 `70,557.5``0.506×`),平均延迟从 `101.20s` 降至 `99.18s``0.980×`)。该结果通过了 H5C 的四任务候选门槛;在当时它仅表示可以在完整参考环境中进行候选重跑,尚不是部署批准,也尚未完成 116 任务 × 5 轮要求。证据见 [H5C JSON](validation/paired_h5c_compact_api35_20260729/evidence.json) 与 [报告](validation/paired_h5c_compact_api35_20260729/report.md);英文部分列出了精确复现命令。
最终参考环境 campaign 已完成全部五项执行门槛:580/580 条唯一 episode116 任务 × 1–5 轮)、零运行时错误、官方 setup 完成,且五个 Pixel 6/API-33 分片均安装相同版本的 24/24 个必需应用。权威结果见[合并 evidence](validation/candidate_h5c_api33_local_qwen_20260804/evidence.json)与[生成报告](validation/candidate_h5c_api33_local_qwen_20260804/report.md)
| 完整候选结果 | 数值 |
| --- | ---: |
| 严格 T3A 成功 | 26 / 580`4.4828%` |
| Evaluator reward | 77 条满分(`1.0`+ 1 条部分分(`0.5` |
| 平均 evaluator reward | `0.133621` |
| 平均步数 / LLM 调用 | `9.672414` / `18.998276` |
| 平均延迟 | `109.860845s` |
| 平均 token | `169,069.563793` |
| 总输入 / 输出 token | `97,384,410` / `675,937` |
| 估算 API 成本 | `$0.00`(本地推理) |
严格成功遵循上游 minimal runner 规则:最终 evaluator 必须为 `1.0`,并且 Agent 必须显式宣告完成。因此 26 条严格成功少于 77 条 evaluator 满分的最终状态;另有 1 条部分 reward `0.5`。Evaluator 失败作为实验结果被完整保留,没有重跑。合并证据中 `scope.direct_episode_gate_completed``scope.full_suite_completed``scope.manuscript_five_seed_gate_completed``experiment_complete` 均为 `true`;但由于观测成绩很低,且没有有效的全集对照,`decision.deployment_approved` 仍为 `false`
候选运行使用本地 `qwen2.5-7b-instruct-local`revision `a09a35458c702b33eeacc393d103063234e8bc28`),由 NVIDIA RTX PRO 6000 Blackwell 96 GB 上的 vLLM 0.19.0 提供服务;H5C 配对源则使用 `doubao-seed-1-6-250615`。因此本 campaign 完成了已晋级观察方案的直接执行/证据要求,但**不是**同模型延续,不能证明比较提升或非劣性。
本结果包含以下明示的兼容处理边界:
1. `ContactsNewContactDraft`UIAutomator 不填充 `state.forest`,因此把 `state.ui_elements` 传给未改动的官方联系人 predicate。
2. Clipper 前台竞态:仅对文档中精确的前台访问错误,在一秒后重试一次未改动的剪贴板读/写操作。
3. `SimpleSmsReplyMostRecent`:多轮询收件箱五秒;若 emulator console 注入后仍为空,将最后一次注入的精确地址/正文写入上游本就直接清理的同一 SMS SQLite 数据库,然后运行未改动的 evaluator query。
4. `RetroPlayingQueue`:仅将固定 APK 缺失 `playing_queue` 表的精确错误映射为空观察队列,再由未改动的精确队列 predicate 记录 evaluator 失败。
5. 原生 32,768-token 上下文溢出:仅在真实 provider context error 之后,确定性重试最多保留 action-selection UI 描述两端共 12,000 个字符,或 before/after summary UI 描述各 6,000 个字符;goal、history、action、reason、guidance、输出格式、保留 UI 的原索引与逐 episode 计数都保留。共发生 63 次截断,移除 7,390,498 个 UI 描述字符。
6. 运行时错误重试复用失败 checkpoint 中保存的精确参数,避免上游生成器漂移改变任务;若后续重生成参数发生漂移,已完成 checkpoint 仍为权威记录。同一存活 emulator 上 resume 时保留已完成的 setup 状态,不重复 setup。
H1/H5 决策门槛要求至少四个完整 pair、净成功增益为正、零配对退化,且平均延迟与 token 都不超过对照的 `1.5×`。H5C 则要求四个紧凑处理组全部成功、零退化、延迟不超过 `1.5×`、token 不超过原始 UIAutomator 的 `0.75×`。通过门槛只允许进入**候选重跑**,绝不等于部署。候选重跑必须提供真实晋级 pair 的 evidence 文件;仅提供 run ID 不够。实验完成还必须有 580 条直接候选记录,即 116 个任务 × 五个不同 trial seed,且没有 episode error。
若要在多个参考环境上并行执行,请保留 `--trials 5`,用 `--trial-indices` 分配一个或多个从 1 开始的 trial,并为每个 emulator 设置不同的 `--execution-shard`。单个分片本身永远不算完成。`merge_candidate_shards.py` 只有在 1–5 号 trial 对每个任务恰好出现一次,并且模型、晋级来源、API-33 环境、上游 App setup 完成状态与必需 App 版本一致时才接受合并。Evaluator 失败属于直接实验结果,必须保留;只有运行时 error 才可通过 `--resume --retry-errors` 重试。
历史配对切片运行在 Pixel 9 Pro API-35 AVD 上,缺少完整第三方 App bundle;`--skip-device-time` 仅用于这些与时间无关的 Wi-Fi evaluator。完整候选 campaign 则改用五个相互隔离的 Pixel 6/API-33 emulator,执行完整上游 setup,并在所有分片上保持相同的 24 个必需 App 包及版本。证据将两种环境明确分开,不把参考环境候选运行描述成 API-35 切片的同环境延伸。
适合精读的轨迹示例:
| 任务 | 材料 | 启示 |
| --- | --- | --- |
| `ExpenseAddMultipleFromGallery` | 失败分析 + `t3a_failed.md` | OCR/多模态缺口;伪造开销条目 |
| `ClockTimerEntry` | 同上 | 无稳定 UI 模型;重复错误输入 |
| `MarkorTranscribeVideo` | 同上 | 会播视频但「看不见」内容 |
| `SportsTracker*` 计数/时长 | 同上 | 有感知无算术 |
| 成功短流程(摄像、秒表等) | `t3a.md` | 对照「正常」轨迹长什么样 |
### 目录结构
```text
chapter7/android-world/
├── README.md # 本文件
├── experiment_core.py # 证据、决策、完成门槛、报告渲染
├── run_controlled_experiment.py # 真实 AndroidWorld 配对/候选 runner
├── merge_candidate_shards.py # 校验并合并并行 trial 分片
├── test_experiment.py # 聚焦的脱机完整性测试
├── requirements.txt # 相邻上游 + API 客户端依赖
├── t3a_summary.md # 汇总指标与能力矩阵
├── t3a_failed_analysis.md # 失败分类与根因
├── t3a.md # 完整运行日志(体积大)
├── t3a_failed.md # 失败任务日志
└── validation/ # 真实 evidence.json + 生成的 report.md
```
### 复现基准(可选)
配套 runner 需要一个独立的相邻 AndroidWorld checkout(当前工作区使用 `chapter7/android_world`)、已配置模拟器以及真实模型凭证。自行重跑请:
1. 克隆 [google-research/android_world](https://github.com/google-research/android_world)(或课程指定 fork)。
2. 按上游文档准备模拟器/真机环境。
3. 在对应环境中安装配套依赖,设置所选 provider 凭证(默认 `ARK_API_KEY`),运行英文部分给出的命令。本地保留运行在 host 使用 `http://127.0.0.1:18111/v1`container 内使用 `http://host.docker.internal:18111/v1``LOCAL_API_KEY` 只设在启动进程环境中,runner 只记录变量名,不保存其值。
4. 只有在完整配置上游 Pixel 6 / API-33 App 后才能尝试 `--full-suite`;不要把 API-35 的 Wi-Fi 专用偏差用于完整 benchmark。
五分片的本地 GPU 命令与严格合并流程见英文复现节;每个隔离 emulator 运行一个 `--trial-indices`,全部完成后再调用 `merge_candidate_shards.py`
仅做笔记研读的推荐顺序:**`t3a_summary.md``t3a_failed_analysis.md` → 抽读 `t3a_failed.md` / `t3a.md` 中的若干 episode**。
### 相关项目
| 项目 | 关系 |
| --- | --- |
| 上游 `android_world`(外部) | 可运行的评测环境 |
| [model-benchmark](../model-benchmark/) | API 时延/可用性维度的评测 |
| [elo-leaderboard](../elo-leaderboard/) | 成对比较式排行,而非绝对任务成功率 |
| [public-health-reporting-eval](../public-health-reporting-eval/) | 仓库内另一套结构化评测脚手架 |
---
## Notes / 说明
- Log files can be **very large** (`t3a.md` ~1MB+). Prefer summary + failed analysis first.
- 日志文件体积很大,建议先读摘要与失败分析。
- Project type: historical **reading / analysis notes** plus a runnable companion that requires a separately provisioned upstream AndroidWorld environment.
- 项目类型:历史**阅读/分析材料** + 可运行配套工具;后者依赖另行配置的上游 AndroidWorld 环境。
+588
View File
@@ -0,0 +1,588 @@
"""Pure reporting helpers for the Experiment 7-12 AndroidWorld loop.
The runtime runner deliberately keeps AndroidWorld imports out of this module so
the evidence checks and report generation can be tested without an emulator.
"""
from __future__ import annotations
from collections import defaultdict
import json
import re
from typing import Any, Iterable, Mapping
BASELINE_TASK_COUNT = 116
WIFI_TASKS = (
"SystemWifiTurnOff",
"SystemWifiTurnOffVerify",
"SystemWifiTurnOn",
"SystemWifiTurnOnVerify",
)
_SECRET_PATTERNS = (
re.compile(r"(?i)(authorization\s*[:=]\s*bearer\s+)[^\s,;]+"),
re.compile(r"(?i)((?:api[_-]?key|access[_-]?token|secret)\s*[:=]\s*)[^\s,;]+"),
re.compile(r"\b(?:sk|ak)-[A-Za-z0-9_-]{12,}\b"),
)
def redact_text(value: object, secrets: Iterable[str] = ()) -> str:
"""Returns a printable error/message with likely credentials removed."""
text = str(value)
for secret in secrets:
if secret:
text = text.replace(secret, "[REDACTED]")
text = _SECRET_PATTERNS[0].sub(r"\1[REDACTED]", text)
text = _SECRET_PATTERNS[1].sub(r"\1[REDACTED]", text)
text = _SECRET_PATTERNS[2].sub("[REDACTED]", text)
return text
def _mean(values: Iterable[float]) -> float | None:
items = list(values)
if not items:
return None
return round(sum(items) / len(items), 6)
def aggregate_episodes(episodes: Iterable[Mapping[str, Any]]) -> dict[str, dict[str, Any]]:
"""Aggregates real episode records by arm."""
groups: dict[str, list[Mapping[str, Any]]] = defaultdict(list)
for episode in episodes:
groups[str(episode["arm"])].append(episode)
output: dict[str, dict[str, Any]] = {}
for arm, rows in sorted(groups.items()):
completed = [row for row in rows if row.get("status") == "completed"]
output[arm] = {
"episodes": len(rows),
"completed_episodes": len(completed),
"error_episodes": len(rows) - len(completed),
"successes": sum(bool(row.get("success")) for row in completed),
"success_rate": _mean(float(bool(row.get("success"))) for row in completed),
"mean_evaluator_reward": _mean(
float(row.get("evaluator_reward", 0.0)) for row in completed
),
"mean_steps": _mean(float(row.get("steps", 0)) for row in completed),
"mean_latency_s": _mean(
float(row.get("elapsed_s", 0.0)) for row in completed
),
"mean_llm_calls": _mean(
float(row.get("llm", {}).get("calls", 0)) for row in completed
),
"mean_llm_latency_s": _mean(
float(row.get("llm", {}).get("latency_s", 0.0)) for row in completed
),
"mean_total_tokens": _mean(
float(row.get("llm", {}).get("input_tokens", 0))
+ float(row.get("llm", {}).get("output_tokens", 0))
for row in completed
),
"total_input_tokens": sum(
int(row.get("llm", {}).get("input_tokens", 0)) for row in completed
),
"total_output_tokens": sum(
int(row.get("llm", {}).get("output_tokens", 0)) for row in completed
),
"total_tokens": sum(
int(row.get("llm", {}).get("input_tokens", 0))
+ int(row.get("llm", {}).get("output_tokens", 0))
for row in completed
),
"estimated_cost_usd": round(
sum(
float(row.get("llm", {}).get("estimated_cost_usd", 0.0))
for row in completed
),
9,
),
}
return output
def paired_rows(episodes: Iterable[Mapping[str, Any]]) -> list[dict[str, Any]]:
"""Builds paired control/treatment comparisons without inventing missing arms."""
groups: dict[str, dict[str, Mapping[str, Any]]] = defaultdict(dict)
for episode in episodes:
if episode.get("arm") in ("control", "treatment"):
groups[str(episode["pair_id"])][str(episode["arm"])] = episode
rows = []
for pair_id, arms in sorted(groups.items()):
if set(arms) != {"control", "treatment"}:
continue
control = arms["control"]
treatment = arms["treatment"]
if control.get("status") != "completed" or treatment.get("status") != "completed":
continue
rows.append({
"pair_id": pair_id,
"task": control["task"],
"trial": control["trial"],
"control_success": bool(control.get("success")),
"treatment_success": bool(treatment.get("success")),
"success_delta": int(bool(treatment.get("success"))) - int(bool(control.get("success"))),
"control_reward": float(control.get("evaluator_reward", 0.0)),
"treatment_reward": float(treatment.get("evaluator_reward", 0.0)),
"reward_delta": round(
float(treatment.get("evaluator_reward", 0.0))
- float(control.get("evaluator_reward", 0.0)),
6,
),
"control_steps": int(control.get("steps", 0)),
"treatment_steps": int(treatment.get("steps", 0)),
"control_latency_s": float(control.get("elapsed_s", 0.0)),
"treatment_latency_s": float(treatment.get("elapsed_s", 0.0)),
})
return rows
def choose_decision(
arm_summary: Mapping[str, Mapping[str, Any]],
pairs: Iterable[Mapping[str, Any]],
*,
minimum_pairs: int = 4,
maximum_latency_ratio: float = 1.5,
maximum_token_ratio: float = 1.5,
) -> dict[str, Any]:
"""Makes a conservative success/cost candidate decision from paired evidence."""
pair_list = list(pairs)
control = arm_summary.get("control")
treatment = arm_summary.get("treatment")
if not control or not treatment or len(pair_list) < minimum_pairs:
return {
"outcome": "insufficient_evidence",
"promote_to_full_suite_candidate": False,
"deployment_approved": False,
"reason": f"Need at least {minimum_pairs} completed pairs; observed {len(pair_list)}.",
}
improvement_count = sum(int(row["success_delta"]) for row in pair_list)
regressions = sum(row["success_delta"] < 0 for row in pair_list)
control_latency = control.get("mean_latency_s")
treatment_latency = treatment.get("mean_latency_s")
latency_ratio = None
if control_latency and treatment_latency is not None:
latency_ratio = round(float(treatment_latency) / float(control_latency), 6)
control_tokens = control.get("mean_total_tokens")
treatment_tokens = treatment.get("mean_total_tokens")
token_ratio = None
if control_tokens and treatment_tokens is not None:
token_ratio = round(float(treatment_tokens) / float(control_tokens), 6)
control_calls = control.get("mean_llm_calls")
treatment_calls = treatment.get("mean_llm_calls")
call_ratio = None
if control_calls and treatment_calls is not None:
call_ratio = round(float(treatment_calls) / float(control_calls), 6)
acceptable_cost = (
latency_ratio is not None
and token_ratio is not None
and latency_ratio <= maximum_latency_ratio
and token_ratio <= maximum_token_ratio
)
if improvement_count > 0 and regressions == 0 and acceptable_cost:
outcome = "promote_candidate_to_full_suite_rerun"
promote = True
reason = (
f"Treatment improved {improvement_count} net paired task(s) with no paired regression. "
"This is a candidate decision, not deployment approval or a full-suite result."
)
elif improvement_count > 0 and regressions == 0:
outcome = "restrict_candidate_due_to_cost"
promote = False
reason = (
"Treatment improved paired success without regressions, but exceeded the "
f"latency/token guardrails ({maximum_latency_ratio:.2f}x / "
f"{maximum_token_ratio:.2f}x). Restrict it to targeted follow-up; do not "
"promote it to the full suite yet."
)
elif improvement_count < 0 or regressions:
outcome = "reject_candidate"
promote = False
reason = (
f"Treatment has {regressions} paired regression(s) and net success delta "
f"{improvement_count}; do not promote."
)
else:
outcome = "inconclusive_no_success_gain"
promote = False
reason = "Treatment produced no paired success gain; keep the upstream control prompt."
return {
"outcome": outcome,
"promote_to_full_suite_candidate": promote,
"deployment_approved": False,
"reason": reason,
"completed_pairs": len(pair_list),
"net_success_delta": improvement_count,
"paired_regressions": regressions,
"mean_latency_ratio_treatment_over_control": latency_ratio,
"mean_token_ratio_treatment_over_control": token_ratio,
"mean_llm_call_ratio_treatment_over_control": call_ratio,
"guardrails": {
"maximum_latency_ratio": maximum_latency_ratio,
"maximum_token_ratio": maximum_token_ratio,
"passed": acceptable_cost,
},
"scope_recommendation": (
"full_suite_candidate_only" if promote else "do_not_deploy"
),
}
def choose_efficiency_decision(
arm_summary: Mapping[str, Mapping[str, Any]],
pairs: Iterable[Mapping[str, Any]],
*,
minimum_pairs: int = 4,
maximum_latency_ratio: float = 1.5,
maximum_token_ratio: float = 0.75,
) -> dict[str, Any]:
"""Promotes a cost refinement only when H5 success is preserved and cost falls."""
pair_list = list(pairs)
control = arm_summary.get("control")
treatment = arm_summary.get("treatment")
if not control or not treatment or len(pair_list) < minimum_pairs:
return {
"outcome": "insufficient_evidence",
"promote_to_full_suite_candidate": False,
"deployment_approved": False,
"reason": f"Need at least {minimum_pairs} completed pairs; observed {len(pair_list)}.",
}
net_success_delta = sum(int(row["success_delta"]) for row in pair_list)
regressions = sum(row["success_delta"] < 0 for row in pair_list)
treatment_successes = sum(bool(row["treatment_success"]) for row in pair_list)
required_treatment_successes = len(pair_list)
success_preserved = treatment_successes == required_treatment_successes
control_latency = control.get("mean_latency_s")
treatment_latency = treatment.get("mean_latency_s")
latency_ratio = (
round(float(treatment_latency) / float(control_latency), 6)
if control_latency and treatment_latency is not None else None
)
control_tokens = control.get("mean_total_tokens")
treatment_tokens = treatment.get("mean_total_tokens")
token_ratio = (
round(float(treatment_tokens) / float(control_tokens), 6)
if control_tokens and treatment_tokens is not None else None
)
control_calls = control.get("mean_llm_calls")
treatment_calls = treatment.get("mean_llm_calls")
call_ratio = (
round(float(treatment_calls) / float(control_calls), 6)
if control_calls and treatment_calls is not None else None
)
passed = (
regressions == 0
and net_success_delta >= 0
and success_preserved
and latency_ratio is not None
and token_ratio is not None
and latency_ratio <= maximum_latency_ratio
and token_ratio <= maximum_token_ratio
)
if passed:
outcome = "promote_efficient_candidate_to_full_suite_rerun"
reason = (
"Treatment preserved paired success with no regression and passed the "
"latency/token efficiency guardrails. This is a candidate decision only."
)
elif not success_preserved or regressions or net_success_delta < 0:
outcome = "reject_efficiency_candidate_due_to_regression"
reason = (
f"Treatment succeeded on {treatment_successes}/{required_treatment_successes} "
f"completed pairs, with {regressions} paired regression(s) and net success "
f"delta {net_success_delta}; it did not preserve the H5 success baseline."
)
else:
outcome = "reject_efficiency_candidate_due_to_cost"
reason = (
"Treatment preserved success but did not reduce tokens to the required "
f"{maximum_token_ratio:.2f}x ratio within the latency guardrail."
)
return {
"outcome": outcome,
"promote_to_full_suite_candidate": passed,
"deployment_approved": False,
"reason": reason,
"completed_pairs": len(pair_list),
"net_success_delta": net_success_delta,
"paired_regressions": regressions,
"treatment_successes": treatment_successes,
"required_treatment_successes": required_treatment_successes,
"success_preservation_passed": success_preserved,
"mean_latency_ratio_treatment_over_control": latency_ratio,
"mean_token_ratio_treatment_over_control": token_ratio,
"mean_llm_call_ratio_treatment_over_control": call_ratio,
"guardrails": {
"objective": "success_noninferiority_and_token_reduction",
"require_all_treatment_pairs_successful": True,
"maximum_latency_ratio": maximum_latency_ratio,
"maximum_token_ratio": maximum_token_ratio,
"passed": passed,
},
"scope_recommendation": (
"full_suite_candidate_only" if passed else "do_not_deploy"
),
}
def enforce_scope_claims(evidence: dict[str, Any]) -> None:
"""Sets completion gates from direct episode evidence, never from counters."""
scope = evidence.setdefault("scope", {})
distinct_tasks = len(set(scope.get("tasks", [])))
configured_trials = int(scope.get("trials_per_task", 0))
mode = scope.get("mode")
tasks = list(dict.fromkeys(str(task) for task in scope.get("tasks", [])))
expected_episode_keys = {
(task, trial)
for task in tasks
for trial in range(1, configured_trials + 1)
}
episodes = evidence.get("episodes", [])
actual_episode_keys = {
(str(row.get("task")), int(row.get("trial", 0))) for row in episodes
}
direct_episode_gate = (
len(episodes) == len(expected_episode_keys)
and actual_episode_keys == expected_episode_keys
and all(
row.get("arm") == "candidate"
and row.get("status") == "completed"
and row.get("evaluator_reward") is not None
and isinstance(row.get("pair_seed"), int)
for row in episodes
)
and all(
len({
row["pair_seed"] for row in episodes if row.get("task") == task
}) == configured_trials
for task in tasks
)
)
full_suite = (
mode == "candidate_rerun"
and distinct_tasks == BASELINE_TASK_COUNT
and configured_trials >= 5
and direct_episode_gate
and evidence.get("environment", {}).get("api_level") == 33
and evidence.get("environment", {}).get("emulator_setup_completed") is True
and evidence.get("environment", {})
.get("app_provisioning", {})
.get("complete")
)
scope["direct_episode_gate_completed"] = direct_episode_gate
scope["full_suite_completed"] = full_suite
scope["manuscript_five_seed_gate_completed"] = full_suite
evidence["experiment_complete"] = bool(
full_suite and evidence.get("decision", {}).get("source_paired_run_id")
)
def _fmt(value: Any) -> str:
if value is None:
return "n/a"
if isinstance(value, float):
return f"{value:.3f}"
return str(value)
def render_report(evidence: Mapping[str, Any]) -> str:
"""Renders the five-stage report from machine-readable evidence."""
scope = evidence["scope"]
environment = evidence["environment"]
arm_summary = evidence.get("arm_summary", {})
decision = evidence.get("decision", {})
pairs = evidence.get("paired_comparison", [])
blockers = evidence.get("environment_boundaries", [])
hypotheses = evidence.get("diagnosis", {}).get("layered_hypotheses", [])
phase = evidence.get("phase", {})
llm_analysis = evidence.get("llm_analysis", {})
controls = (
"same checkout, model, task parameters, generated seed policy, step budget, "
"Pixel 6/API-33 device class, upstream setup, and app versions across isolated shards."
if environment.get("shard_devices")
else "same checkout, model, task parameters, generated seed, step budget, and "
"emulator; arm order alternates by pair."
)
lines = [
"# Experiment 7-12 AndroidWorld iteration report",
"",
f"- Run ID: `{evidence['run_id']}`",
f"- Generated (UTC): `{evidence['generated_at_utc']}`",
f"- Upstream commit: `{environment.get('android_world_commit', 'not reached')}`",
f"- Device: `{environment.get('device_model', 'not reached')}`, API "
f"`{environment.get('api_level', 'not reached')}` (upstream tested reference: API "
f"`{environment.get('upstream_tested_api_level', 33)}`)",
f"- Observation method: `{environment.get('a11y_method', 'a11y_forwarder_app')}`",
f"- Provider/model: `{evidence['model']['provider']}` / `{evidence['model']['model']}`",
f"- Model source/runtime: `{evidence['model'].get('source', 'not recorded')}` / "
f"`{evidence['model'].get('runtime', 'not recorded')}`",
f"- Accelerator: `{evidence['model'].get('accelerator', 'not recorded')}`",
f"- Required apps: "
f"`{environment.get('app_provisioning', {}).get('installed_required_package_count', 'not reached')}/"
f"{environment.get('app_provisioning', {}).get('required_package_count', 'not reached')}`",
f"- Scope: {len(scope['tasks'])} task(s), {scope['trials_per_task']} trial(s), "
f"mode `{scope['mode']}`",
f"- Full 116-task × 5-seed suite completed: **{str(scope['full_suite_completed']).lower()}**",
"",
"The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% "
"numbers are explicitly hypothetical and are not used as rerun results here.",
"",
"## 1. Diagnose",
"",
]
for item in evidence["diagnosis"]["findings"]:
lines.append(f"- {item}")
lines.extend([
"",
"## 2. Hypothesis",
"",
"The diagnosis produced explicit surface, middle, and deep hypotheses. Only one "
"variable is changed in this run; the other hypotheses remain untested.",
"",
"| Layer / ID | Proposed change | Target | Verification | Status |",
"| --- | --- | --- | --- | --- |",
])
for row in hypotheses:
lines.append(
f"| {row.get('layer', 'n/a')} / `{row.get('id', 'n/a')}` | "
f"{row.get('idea', 'n/a')} | {row.get('target', 'n/a')} | "
f"{row.get('verification', 'n/a')} | {row.get('status', 'not tested')} |"
)
lines.extend([
"",
f"Selected hypothesis: `{evidence['hypothesis']['id']}`",
f"- Change: {evidence['hypothesis']['change']}",
f"- Expected measurable result: {evidence['hypothesis']['expected_result']}",
f"- Guardrails: {evidence['hypothesis']['guardrails']}",
"",
"## 3. Controlled experiment",
"",
f"- Phase: `{phase.get('id', 'phase_1_surface')}` — "
f"{phase.get('description', 'low-cost surface prompt ablation')}",
f"- Independent variable: {phase.get('independent_variable', 'task-specific T3A guidelines')}",
f"- Controls: {controls}",
"",
"| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |",
"| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |",
])
for arm, row in sorted(arm_summary.items()):
lines.append(
f"| {arm} | {row['completed_episodes']}/{row['episodes']} | "
f"{_fmt(row['success_rate'])} | {_fmt(row['mean_evaluator_reward'])} | "
f"{_fmt(row['mean_steps'])} | {_fmt(row['mean_latency_s'])} | "
f"{_fmt(row['mean_llm_calls'])} | {_fmt(row.get('mean_total_tokens'))} | "
f"{row['total_input_tokens']} / {row['total_output_tokens']} | "
f"{row.get('estimated_cost_usd', 0.0):.6f} |"
)
if pairs:
lines.extend([
"",
"| Task / trial | Control | Treatment | Δ success | Control→treatment steps |",
"| --- | ---: | ---: | ---: | ---: |",
])
for row in pairs:
lines.append(
f"| {row['task']} / {row['trial']} | {int(row['control_success'])} | "
f"{int(row['treatment_success'])} | {row['success_delta']:+d} | "
f"{row['control_steps']}{row['treatment_steps']} |"
)
lines.extend([
"",
"## 4. Data-driven decision",
"",
f"- Outcome: **`{decision.get('outcome', 'not_applicable')}`**",
f"- Reason: {decision.get('reason', 'This artifact is a candidate rerun, not a paired decision run.')}",
f"- Treatment/control mean latency ratio: "
f"{_fmt(decision.get('mean_latency_ratio_treatment_over_control'))}",
f"- Treatment/control mean token ratio: "
f"{_fmt(decision.get('mean_token_ratio_treatment_over_control'))}",
f"- Treatment/control mean LLM-call ratio: "
f"{_fmt(decision.get('mean_llm_call_ratio_treatment_over_control'))}",
f"- Cost guardrails passed: **{str(decision.get('guardrails', {}).get('passed', False)).lower()}**",
f"- Deployment approved: **{str(decision.get('deployment_approved', False)).lower()}**",
"",
"## 5. Rerun and next report",
"",
])
if scope["full_suite_completed"]:
lines.append(
"The complete 116-task, five-trial candidate rerun gate is satisfied by direct episode evidence."
)
else:
lines.append(
"This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. "
"The next gate is a conditionally enabled candidate rerun over all 116 tasks with five "
"seeds after provisioning the upstream API-33 app environment."
)
failed = [
episode for episode in evidence.get("episodes", [])
if episode.get("status") != "completed" or not episode.get("success")
]
if failed:
lines.append("")
lines.append("Observed residual failures:")
for episode in failed:
if episode.get("error"):
detail = episode["error"]
elif episode.get("evaluator_reward") == 1.0 and not episode.get("agent_declared_done"):
detail = "final evaluator state passed, but the agent never declared completion"
elif episode.get("agent_declared_done") and episode.get("evaluator_reward") != 1.0:
detail = "agent declared completion, but the real evaluator state failed"
else:
detail = "evaluator reward / completion gate was not satisfied"
lines.append(f"- `{episode['arm']} / {episode['task']} / trial {episode['trial']}`: {detail}")
lines.extend(["", "### LLM analysis of this run", ""])
if llm_analysis.get("status") == "completed":
lines.append(
"The following bounded interpretation was produced by the configured real LLM from "
"the aggregate evidence (the JSON remains authoritative):"
)
lines.append("")
lines.append(f"- Summary: {llm_analysis.get('summary', 'n/a')}")
lines.append(
f"- Cost/benefit interpretation: {llm_analysis.get('cost_benefit_interpretation', 'n/a')}"
)
for item in llm_analysis.get("observed_failure_pattern", []):
lines.append(f"- Residual pattern: {item}")
next_hypothesis = llm_analysis.get("next_hypothesis", {})
if next_hypothesis:
lines.append(
f"- Next hypothesis `{next_hypothesis.get('id', 'n/a')}` "
f"({next_hypothesis.get('layer', 'n/a')}): {next_hypothesis.get('idea', 'n/a')} "
f"Target: {next_hypothesis.get('target', 'n/a')} Verification: "
f"{next_hypothesis.get('verification', 'n/a')}"
)
else:
lines.append(
f"No LLM analysis was accepted: {llm_analysis.get('error', 'analysis was not requested for this artifact')}"
)
lines.extend(["", "## Environment boundaries", ""])
if blockers:
lines.extend(f"- {item}" for item in blockers)
else:
lines.append("- None recorded.")
lines.extend([
"",
"The JSON beside this report is the authoritative evidence. It contains episode-level "
"evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; "
"credentials and raw prompts are not stored.",
"",
])
return "\n".join(lines)
def dumps_json(evidence: Mapping[str, Any]) -> str:
return json.dumps(evidence, ensure_ascii=False, indent=2, sort_keys=True) + "\n"
@@ -0,0 +1,271 @@
#!/usr/bin/env python3
"""Merge independently executed Experiment 7-12 trial shards without hiding failures."""
from __future__ import annotations
import argparse
import copy
import hashlib
import json
from pathlib import Path
from typing import Any
from experiment_core import (
BASELINE_TASK_COUNT,
aggregate_episodes,
dumps_json,
enforce_scope_claims,
paired_rows,
render_report,
)
from run_controlled_experiment import (
EXPERIMENT_ID,
OpenAICompatibleLlm,
_generate_llm_analysis,
_utc_now,
)
def _parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("shards", nargs="+", type=Path)
parser.add_argument("--output-dir", type=Path, required=True)
parser.add_argument("--source-paired-evidence", type=Path, required=True)
parser.add_argument("--analysis-base-url")
parser.add_argument("--analysis-api-key-env", default="LOCAL_API_KEY")
parser.add_argument("--analysis-model")
return parser.parse_args()
def _load(path: Path) -> dict[str, Any]:
data = path.read_bytes()
evidence = json.loads(data)
if evidence.get("experiment") != EXPERIMENT_ID:
raise RuntimeError(f"Not direct Experiment {EXPERIMENT_ID} evidence: {path}")
return evidence
def _app_versions(evidence: dict[str, Any]) -> list[tuple[str, Any, Any]]:
apps = evidence.get("environment", {}).get("app_provisioning", {}).get("apps", [])
return sorted(
(row.get("package"), row.get("version_code"), row.get("version_name"))
for row in apps
)
def main() -> int:
args = _parse_args()
if args.output_dir.exists():
raise RuntimeError(f"Output directory already exists: {args.output_dir}")
paths = [path.resolve() for path in args.shards]
shards = [_load(path) for path in paths]
if len(shards) < 2:
raise RuntimeError("At least two independent shards are required")
reference = shards[0]
reference_tasks = reference.get("scope", {}).get("tasks", [])
reference_trials = int(reference.get("scope", {}).get("trials_per_task", 0))
if len(reference_tasks) != BASELINE_TASK_COUNT or reference_trials < 5:
raise RuntimeError("Shards must declare the complete 116-task, five-trial scope")
if reference.get("scope", {}).get("mode") != "candidate_rerun":
raise RuntimeError("Only candidate-rerun shards can be merged")
selected_trials: set[int] = set()
episodes: list[dict[str, Any]] = []
seen_keys: set[tuple[str, int, str]] = set()
shard_rows = []
reference_model = reference.get("model")
reference_source = reference.get("decision", {}).get("source_paired_run_id")
reference_versions = _app_versions(reference)
merged_boundaries: list[str] = []
merged_retry_history: list[dict[str, Any]] = []
merged_parameter_drift: list[dict[str, Any]] = []
merged_retry_parameter_drift: list[dict[str, Any]] = []
for path, shard in zip(paths, shards):
scope = shard.get("scope", {})
if scope.get("tasks") != reference_tasks:
raise RuntimeError(f"Task ordering differs in shard: {path}")
if int(scope.get("trials_per_task", 0)) != reference_trials:
raise RuntimeError(f"Trial scope differs in shard: {path}")
if shard.get("model") != reference_model:
raise RuntimeError(f"Model configuration differs in shard: {path}")
if shard.get("decision", {}).get("source_paired_run_id") != reference_source:
raise RuntimeError(f"Promoted paired source differs in shard: {path}")
environment = shard.get("environment", {})
if environment.get("api_level") != 33:
raise RuntimeError(f"Shard is not on reference API 33: {path}")
if environment.get("emulator_setup_completed") is not True:
raise RuntimeError(f"Shard did not complete official emulator/app setup: {path}")
if not environment.get("app_provisioning", {}).get("complete"):
raise RuntimeError(f"Shard has an incomplete official app bundle: {path}")
if _app_versions(shard) != reference_versions:
raise RuntimeError(f"Official app versions differ in shard: {path}")
if shard.get("credentials_persisted") is not False:
raise RuntimeError(f"Shard does not attest credential-free evidence: {path}")
shard_trials = {int(value) for value in scope.get("selected_trials", [])}
if not shard_trials or selected_trials.intersection(shard_trials):
raise RuntimeError(f"Missing or overlapping selected trials in shard: {path}")
selected_trials.update(shard_trials)
for episode in shard.get("episodes", []):
trial = int(episode.get("trial", 0))
key = (str(episode.get("task")), trial, str(episode.get("arm")))
if trial not in shard_trials:
raise RuntimeError(f"Episode lies outside its declared trial shard: {key}")
if key in seen_keys:
raise RuntimeError(f"Duplicate direct episode across shards: {key}")
seen_keys.add(key)
episodes.append(copy.deepcopy(episode))
shard_rows.append({
"path": str(path),
"sha256": hashlib.sha256(path.read_bytes()).hexdigest(),
"run_id": shard.get("run_id"),
"execution_shard": environment.get("execution_shard"),
"selected_trials": sorted(shard_trials),
"episodes": len(shard.get("episodes", [])),
"completed_episodes": sum(
row.get("status") == "completed" for row in shard.get("episodes", [])
),
"error_episodes": sum(
row.get("status") != "completed" for row in shard.get("episodes", [])
),
})
for boundary in shard.get("environment_boundaries", []):
if boundary not in merged_boundaries:
merged_boundaries.append(boundary)
for retry in shard.get("retry_history", []):
merged_retry_history.append({
"execution_shard": environment.get("execution_shard"),
**copy.deepcopy(retry),
})
for drift in shard.get("resume_parameter_drift", []):
row = {
"execution_shard": environment.get("execution_shard"),
**copy.deepcopy(drift),
}
if row not in merged_parameter_drift:
merged_parameter_drift.append(row)
for drift in shard.get("retry_parameter_drift", []):
row = {
"execution_shard": environment.get("execution_shard"),
**copy.deepcopy(drift),
}
if row not in merged_retry_parameter_drift:
merged_retry_parameter_drift.append(row)
expected_trials = set(range(1, reference_trials + 1))
if selected_trials != expected_trials:
raise RuntimeError(
f"Shard trial union is {sorted(selected_trials)}; expected {sorted(expected_trials)}"
)
task_order = {task: index for index, task in enumerate(reference_tasks)}
episodes.sort(key=lambda row: (task_order[str(row["task"])], int(row["trial"])))
merged = copy.deepcopy(reference)
merged["run_id"] = "exp7-12-merged-" + _utc_now().replace(":", "").replace("-", "")
merged["generated_at_utc"] = _utc_now()
merged["command"] = ["merge_candidate_shards.py", *map(str, paths)]
merged["scope"]["selected_trials"] = sorted(selected_trials)
merged["episodes"] = episodes
merged["shards"] = shard_rows
merged["environment_boundaries"] = merged_boundaries
merged["retry_history"] = merged_retry_history
if merged_parameter_drift:
merged["resume_parameter_drift"] = merged_parameter_drift
else:
merged.pop("resume_parameter_drift", None)
if merged_retry_parameter_drift:
merged["retry_parameter_drift"] = merged_retry_parameter_drift
else:
merged.pop("retry_parameter_drift", None)
merged["environment"]["execution_shard"] = "merged"
merged["environment"]["shard_devices"] = [
{
"run_id": shard.get("run_id"),
"execution_shard": shard.get("environment", {}).get("execution_shard"),
"device_serial": shard.get("environment", {}).get("device_serial"),
"avd_name": shard.get("environment", {}).get("avd_name"),
"api_level": shard.get("environment", {}).get("api_level"),
}
for shard in shards
]
merged["scope"]["completed_episodes"] = sum(
row.get("status") == "completed" for row in episodes
)
merged["scope"]["error_episodes"] = sum(
row.get("status") != "completed" for row in episodes
)
merged["arm_summary"] = aggregate_episodes(episodes)
merged["paired_comparison"] = paired_rows(episodes)
enforce_scope_claims(merged)
paired_source = json.loads(args.source_paired_evidence.read_text(encoding="utf-8"))
if paired_source.get("run_id") != reference_source:
raise RuntimeError("Supplied paired evidence does not match the shard source run ID")
merged["decision"]["source_paired_evidence"] = str(
args.source_paired_evidence.resolve()
)
merged["decision"]["source_paired_model"] = paired_source.get("model")
paired_model = paired_source.get("model", {}).get("model")
candidate_model = merged.get("model", {}).get("model")
if paired_model and paired_model != candidate_model:
message = (
f"The full-suite candidate uses model {candidate_model}, while the promoted paired "
f"H5C source used {paired_model}. This user-requested local-GPU campaign evaluates "
"the promoted observation treatment but is not a same-model extension of the paired result."
)
if message not in merged.setdefault("environment_boundaries", []):
merged["environment_boundaries"].append(message)
if merged["scope"]["full_suite_completed"]:
merged["decision"].update({
"outcome": "full_candidate_rerun_completed",
"deployment_approved": False,
"reason": (
"The direct 116-task x five-trial candidate rerun completed on five "
"independent reference-environment shards. Negative evaluator results are retained."
),
})
else:
merged["decision"].update({
"outcome": "candidate_rerun_has_errors",
"deployment_approved": False,
"reason": (
"The merged candidate evidence contains missing or error episodes and does not "
"satisfy the strict completion gate."
),
})
if args.analysis_base_url and args.analysis_model:
import os
api_key = os.environ.get(args.analysis_api_key_env)
if not api_key:
raise RuntimeError(
f"Analysis credential variable is unset: {args.analysis_api_key_env}"
)
llm = OpenAICompatibleLlm(
api_key=api_key,
base_url=args.analysis_base_url,
model=args.analysis_model,
seed=int(reference_model.get("seed", 42)),
max_tokens=int(reference_model.get("max_tokens", 1024)),
timeout_s=120,
retries=1,
input_cost_per_million_usd=0.0,
output_cost_per_million_usd=0.0,
)
merged["llm_analysis"] = _generate_llm_analysis(merged, llm, [api_key])
else:
merged["llm_analysis"] = {
"status": "not_run",
"error": "Merged analysis endpoint was not configured.",
}
args.output_dir.mkdir(parents=True)
(args.output_dir / "evidence.json").write_text(dumps_json(merged), encoding="utf-8")
(args.output_dir / "report.md").write_text(render_report(merged), encoding="utf-8")
print(f"Evidence: {args.output_dir / 'evidence.json'}")
print(f"Report: {args.output_dir / 'report.md'}")
return 0 if merged["scope"]["full_suite_completed"] else 2
if __name__ == "__main__":
raise SystemExit(main())
+2
View File
@@ -0,0 +1,2 @@
-e ../android_world
openai>=1.30.0,<3
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,147 @@
### 类别一:转录失败 (Transcription Failure)
这类失败的核心是 Agent 无法从非文本格式(如图片、视频)中提取和理解文字信息。
**1. 任务: `ExpenseAddMultipleFromGallery` (task 13)**
* **目标**: 从图库中的 `expenses.jpg` 图片里读取开销项目,并将其添加到“Pro Expense”应用中。
* **失败过程分析**:
1. **操作正确**: Agent 成功地打开了“Simple Gallery Pro”,找到了 `expenses.jpg` 并将其全屏显示。这一系列的UI导航操作是完全正确的。
2. **认知断层**: 在第8步,Agent 的思考(`Reason`)是:“我需要先打开开销追踪应用,为添加开销数据做准备”。这表明它**知道**下一步该做什么,但完全**没有提及**它从图片中看到了什么。它无法“读取”图片上的文字。
3. **行为“伪装”**: 从第13步开始,Agent 在“Pro Expense”应用中输入了“Coffee”、“4.50”、“Lunch”、“12.75”等条目。这些数据并非来自图片,而是 Agent 根据任务描述“添加开销”而**凭空生成或从其训练数据中联想出的通用示例**。它在模拟一个成功的流程,尽管它没有获取到完成任务所必需的真实信息。
4. **最终失败**: 由于 Agent 添加的数据与 `expenses.jpg` 中的真实数据完全不符,最终的任务验证(通过检查应用数据库)失败。
* **根本原因**: Agent 的视觉模型缺乏 OCR 能力。它能识别UI元素(按钮、列表),但无法将图片中的像素信息转译为结构化的文本数据。
**2. 任务: `MarkorTranscribeVideo` (task 36)**
* **目标**: 观看 `ZwUN_moment_70_.mp4` 视频,并将视频帧上显示的文字序列,以逗号分隔的形式记录到 `Markor` 笔记中。
* **失败过程分析**:
1. **导航能力出色**: Agent 展现了复杂的导航能力。它成功打开VLC,处理了权限请求,浏览到 `Download` 文件夹,并正确地播放了目标视频。它甚至还处理了VLC首次启动时的教程浮层。
2. **内容理解为零**: 在第12步到第20步之间,Agent 反复尝试与播放器交互(点击播放、暂停),但其思考日志中从未提及视频中出现的任何文字内容。它就像一个盲人,可以熟练地操作播放器,却看不见屏幕上播放的内容。
3. **陷入无效循环**: 它似乎知道任务尚未完成,所以不断重复“点击播放器”这个动作,希望能触发下一步,但由于无法获取新信息(视频文字),它无法推进任务。
4. **最终失败**: 任务因步骤超限而失败。
* **根本原因**: 与图片转录类似,Agent 完全不具备从视频帧中提取视觉信息并将其转换为文本的能力。这是其感知模块的根本性缺陷。
---
### 类别二:复杂UI理解失败 (Complex UI Understanding Failure)
这类失败的核心是 Agent 无法理解非标准、复杂或有特殊状态逻辑的UI。它能看到按钮,但不知道按下按钮后会发生什么,或者如何通过一系列操作达到特定UI状态。
**1. 任务: `ClockTimerEntry` (task 9)**
* **目标**: 设置一个“0小时16分35秒”的计时器。
* **失败过程分析**:
1. **初步操作正确**: Agent 打开时钟,切换到计时器标签页,准备输入。
2. **发现错误**: 它尝试输入`1635`。在输入`1``6``3`后,屏幕显示为 `00h 01m 63s`。在第5步,它的思考日志写道:“我看到当前显示的是‘00h 01m 63s’,这是一个无效的时间格式”。这表明它**具备基础的UI阅读能力**,能识别出错误,但是并不理解输入显示的逻辑应当是什么样的,缺失基于预期的长久规划,所以错误的终止了继续输入的行为,而是转向了错误处理。
3. **能够纠正,但不能学习**: 它接着正确地使用了三次退格键,将输入清零。
4. **陷入逻辑循环**: 从第9步开始,它完全重复了第一次的错误输入序列,再次输入`1``6`...。它没有从上一次的失败中学习到这个UI的输入机制
* **根本原因**: Agent 对这个UI的理解是表层的。它知道“输入数字”和“删除”,但它没有形成一个关于“计时器数字如何进位”的心智模型。因此,当它的简单策略失败时,它无法调整策略,只能重复无效的尝试。
### 类别三:应用初始化导致的步数限制
**1. 任务: `MarkorAddNoteHeader` (task 23)**
* **目标**: 在一个笔记文件的现有内容**之前**,添加一行新文字和一个空行。
* **失败过程分析**:
1. **导航与初级编辑正确**: Agent 成功打开 `Markor`,找到并打开了目标笔记。
2. Step 1-7: 处理首次启动。Agent 打开 Markor 应用,然后像一个真实用户一样,耐心地点击了5次“下一步”来完成教程,最后关闭了弹出的更新日志。这7步是正确的,但对于完成核心任务来说是“开销”。
3. **高级操作的困惑**: 在第11和12步,它通过长按和点击,成功地“全选”了所有文本。这是一个相对复杂的操作,它执行得很好。
* **根本原因**: 应用初次启动时需要进行初始化设置,这部分步数消耗了模型的有效步数。
---
### 类别四:数学/计数失败 (Math/Counting Failure)
这类失败的核心是 Agent 缺乏在UI操作环境中执行基本数学或逻辑运算的能力。
**1. 任务: `SportsTrackerActivitiesCountForWeek` (task 87)**
* **目标**: 计算“本周”(从周一开始)总共有多少次“跑步”活动。
* **失败过程分析**:
1. **数据收集成功**: Agent 成功打开 `OpenTracks` 应用,并在日志中通过上下滚动,完整地浏览了本周的所有活动列表。它把完成任务所需的所有原始信息都“看”了一遍。
2. **认知处理失败**: 在第10步,它给出了答案 `2`。令人疑惑的地方在于:Agent 声称自己完成了任务,而错误的原因却是:Agent did not indicate task is done. Reached max number of steps.
* **可能原因**: Agent 能够完成感知(看到列表)和操作(滚动列表)的部分,但没能完成认知处理(筛选“跑步”活动,并对结果进行**计数**)。
**2. 任务: `SportsTrackerTotalDurationForCategoryThisWeek` (task 92)**
* **目标**: 计算本周所有“跑步”活动的总时长。
* **失败过程分析**:
1. **数据收集成功**: 和上一个任务一样,Agent 成功打开应用并浏览了所有本周的活动。它看到了每个活动的名称和时长(例如 `30:00`, `1:15:00`)。
2. **认知处理失败**: 它没能在限定步数内完成这个任务,最终超时。因为它需要执行比计数更复杂的认知操作:
* **筛选**: 找到所有“跑步”活动。
* **提取**: 记下每个跑步活动的时长。
* **求和**: 将所有时长加在一起。
* **可能原因**: Agent 不具备执行数学**求和**的能力。它可以看到数字,但无法在操作的上下文中对它们进行计算。
---
### 类别五:信息检索与复杂逻辑失败
这类失败的核心是 Agent 无法在复杂场景下定位信息,或无法执行需要多个逻辑步骤和状态维护的复杂任务。
**1. 任务: `SimpleCalendarAnyEventsOnDate` (task 68)**
* **目标**: 检查“10月28日”是否有任何事件。
* **失败过程分析**:
1. **界面理解失败**: Agent 打开了日历的月视图。这是一个信息密集的界面。在第2步,它试图点击10月28日,但错误地点到了10月13日。
2. **低效的恢复策略**: 发现错误后,它没有返回月视图重新选择,而是陷入了一种非常低效的策略:在日视图中,一次又一次地点击“向后一天”的箭头,试图从13日逐天“走到”28日。
3. **最终失败**: 这种策略耗时过长,导致任务在到达目标日期前就因步骤超限而失败。
* **根本原因**: 这是**信息检索**和**复杂UI理解**的复合失败。首先,它无法在密集网格中准确定位目标(检索失败)。其次,它没有更好的错误恢复逻辑,只会采用最原始、最低效的线性导航策略。
**2. 任务: `RecipeDeleteDuplicateRecipes3` (task 52)**
* **目标**: 删除所有重复的食谱,但要确保每种独特的食谱至少保留一个实例。
* **失败过程分析**:
1. **子任务成功**: 日志显示,Agent 在初期是成功的。它能识别出“Avocado Toast with Egg”是重复的,并成功删除了一个副本。它也能识别并删除一个“BBQ Chicken Quesadillas”的副本。
2. **状态维护失败**: **失败点在于**,在成功执行了几次删除后,它似乎“忘记”了自己已经处理过哪些食谱,或者被列表更新后的新布局搞糊涂了。它开始陷入循环,反复尝试删除“Butternut Squash Soup”,或者在已经清理过的项目上再次尝试,最终在长达34步的无效操作后超时。
* **根本原因**: 这是一个典型的**复杂任务规划和状态维护**的失败。Agent 可以执行“删除一个重复项”这个简单的子程序,但它无法管理一个更高级的、需要记忆和迭代的循环逻辑(“*当还有未处理的重复项时,找到下一组并处理*”)。它的任务计划非常脆弱,一旦UI状态因自身操作而改变(列表项变少),它就无法适应新的状态,导致逻辑混乱。
*
### 超时类型
好的,我们来对 `t3a_failed.md` 文件中所有任务进行一次全面的统计,专门找出那些因为“达到最大步数限制”而失败的案例。
失败日志中对应的标志性语句是:
`Agent did not indicate task is done. Reached max number of steps.`
经过对整个 `t3a_failed.md` 文件内容的统计,在所有失败的任务中,**总共有 28 个任务**是因为步数耗尽而失败的。
单纯说“步数不足”是表象,更深层的原因是 Agent **为何会**耗尽步数。我们可以将这 28 个失败案例归为以下几种典型的“低效模式”:
#### 1. 陷入无效的逻辑循环或低效策略 (9个任务)
Agent 知道目标,但它采用的策略是错误的或极其低效的,导致它在重复的无用功中耗尽了所有步数。
* **`ClockTimerEntry`**: 发现输入错误后,能够清零,但随即又**重复了一遍完全相同的错误输入**,陷入循环。
* **`RecipeDeleteDuplicateRecipes` (3个相关任务)**: 能够成功删除一两个重复项,但随后就**逻辑混乱**,无法系统性地清理所有重复项,开始无效点击,直到超时。
* **`SimpleCalendar...` (4个相关任务)**: 在日历中定位日期时,一旦首次点击错误,就陷入了“从A日逐天点击到B日”的**最低效导航策略**,而不是返回月视图重新选择,快速消耗了步数。
* **`MarkorTranscribeVideo`**: 由于无法从视频中获取信息,它只能反复“点击播放器”,期待状态改变,陷入了**无信息输入的无效等待循环**。
#### 2. 被复杂或非标UI交互所困 (7个任务)
Agent 在面对不熟悉的、需要精细操作的UI控件时,会反复尝试错误的交互方式,直至失败。
* **`MarkorAddNoteHeader`, `MarkorChangeNoteContent`, `MarkorEditNote`**: 这三个任务都失败在**文本编辑器**上。Agent 不理解如何精确地定位光标、插入文本、替换部分文本,而是错误地选择了“全选”等不适用的策略。
* **`MarkorMoveNote`**: 在文件移动的对话框中,虽然最终找到了正确的文件夹,但之前的导航步骤过于繁琐,耗尽了步数。
* **`OsmAndMarker`**: 在地图应用中,它完全无法理解如何添加“标记点”,在长按、点击等操作中浪费了所有步数。
* **`SimpleDrawProCreateDrawing`**: 在保存文件的对话框中,它无法正确地**清空并重命名**文件,反复进行无效的删除和输入尝试。
#### 3. 信息处理能力缺失导致的操作停滞 (5个任务)
Agent 能够看到数据,但无法在认知层面进行处理(计数、求和、比较),导致它只能反复地、漫无目的地滚动浏览,无法给出答案。
* **`SportsTrackerActivitiesCountForWeek`**
* **`SportsTrackerActivityDuration`**
* **`SportsTrackerLongestDistanceActivity`**
* **`SportsTrackerTotalDistanceForCategoryOverInterval`**
* **`SportsTrackerTotalDurationForCategoryThisWeek`**
在这5个任务中,Agent 的行为模式高度一致:成功打开应用,然后花费大量步数(通常是7-8步)反复上下滚动列表,最后因为无法进行**计数、求和、比较大小**等数学运算而超时失败。
#### 4. 简单的导航或执行失败 (7个任务)
在这些看似直接的任务中,Agent 依然会因为某些原因“迷路”或卡住,最终超时。
* **`MarkorCreateFolder`, `MarkorCreateNoteAndSms`, `MarkorCreateNoteFromClipboard`**: 在创建文件/文件夹的流程中,Agent 在处理文件名、扩展名或后续分享步骤时被卡住。
* **`MarkorDeleteNewestNote`, `MarkorDeleteNote`**: 在删除文件的流程中,它在排序或选择文件后,未能及时执行删除操作。
* **`RetroPlaylistDuration`**: 在添加歌曲到播放列表时,它能成功添加几首,但似乎没有一个明确的停止条件和检查机制(检查总时长),导致无限添加下去直至超时。
* **`SimpleCalendarDeleteEvents`**: 在删除多个事件时,它成功删除了一个,但之后没能有效地继续删除下一个,浪费了步骤。
### 结论
**“步数耗尽”更像是一个“症状”而非“病因”。**
这 28 个案例清晰地表明,Agent 的失败并非因为步数限制本身过于苛刻,而是因为它缺乏高效解决问题的核心能力。当面对其能力短板(如复杂UI理解、数学运算、高级任务规划)时,它无法制定出简洁有效的执行路径,只能通过大量低效甚至无效的尝试来“凑步数”,最终不可避免地走向失败。
+254
View File
@@ -0,0 +1,254 @@
### 核心内容概括
1. **逐任务性能列表 (Per-Task Performance List):**
- 详细列出了 Agent 在116个不同任务上的表现。
- 关键指标包括:成功率、平均任务长度(步数)、耗时等。
2. **按能力标签与难度划分的性能分析 (Performance Analysis by Capability Tag and Difficulty):**
- 将所有任务按照所需的核心能力和难度进行分类。
- 统计并展示 Agent 在每种能力/难度组合下的平均成功率。
### 第一部分:逐任务性能分析
---
### **1. 表格列释义**
我们首先明确每个数据列的含义:
- `task`: 任务的唯一名称,清晰地描述了任务目标。例如 `AudioRecorderRecordAudio` (录音机录制音频) 或 `ContactsAddContact` (联系人应用中添加联系人)。
- `task_num`: 任务的数字编号,从0到115。
- `num_complete_trials`: 已完成的尝试次数。在这里,每个任务都只尝试了1次。
- `mean_success_rate`: 平均成功率。由于每个任务只尝试1次,`1.00` 代表成功,`0.00` 代表失败。
- `mean_episode_length`: 平均回合长度。一个“回合”指 Agent 为完成一次任务所执行的**总步数**(如点击、输入、滑动等)。这个数值越高,通常意味着任务越复杂或 Agent 走了弯路。对于失败的任务,此值为 `NaN` (Not a Number),表示没有成功的记录。
- `total_runtime_s`: 完成该任务总共消耗的时间,单位是秒 (s)。
- `num_fail_trials`: 失败的尝试次数。`0.00` 表示该次尝试成功,`1.00` 表示失败。
### **2. 洞察与发现**
1. 压倒性的成功率:
从 task_num 0 到 81,以及从 93 到 101Agent 的 mean_success_rate 均为 1.00。这表明 Agent 在处理绝大多数结构化、明确的任务时表现得非常出色且稳定。这些任务涵盖了:
- **基础应用操作**:相机、时钟、录音机、联系人。
- **文件管理**:删除文件、移动文件。
- **内容创作与编辑**:在 `Markor`(一个笔记应用)中创建/编辑笔记。
- **系统设置**:开关蓝牙、调节亮度。
2. 明确的失败群:
失败的案例非常集中,主要出现在 task_num 82 以及 102 至 115 的任务中。
- **`SimpleSmsReplyMostRecent` (task 82):** 回复最近一条短信,失败。
- **`SystemWifiTurn...` (tasks 102, 103, 104, 105):** 与开关Wi-Fi及验证状态相关的任务,有多次失败。
- **`Tasks...` (tasks 106 - 111):** 与 `Tasks`(待办事项)应用相关的一系列查询任务,全部失败。
- **`Turn...` (tasks 112, 113):** 涉及Wi-Fi和蓝牙组合操作的任务,失败。
- **`Vlc...` (tasks 114, 115):** 在VLC播放器中创建播放列表的任务,失败。
3. 任务复杂度的体现:
观察 mean_episode_length(平均步数),我们可以评估任务的复杂度。
- **简单任务**`ClockStopWatchPausedVerify` (task 7) 仅需 3 步。
- **复杂任务**`RecipeAddMultipleRecipesFromMarkor2` (task 48) 需要 44 步,`OsmAndTrack` (task 44) 需要 41 步。这说明 Agent 能够维持一个较长的操作序列来完成复杂的目标。
**d. 总体性能平均值 (Average)**
- `mean_success_rate`: **`0.88`**,即 88% 的总体成功率。这是一个非常高的指标,说明 Agent 具有很强的泛化能力和执行能力。
- `num_fail_trials`: **`0.12`**,即 12% 的失败率,与成功率相对应。
- `mean_episode_length`: **`13.45`** 步,所有成功任务的平均操作步数。
### 第二部分:按能力标签与难度划分的性能分析
### **1. 表格解读**
这张表格是一个诊断矩阵。它不再关注单个任务的成败,而是将任务解构成一系列所需的核心能力(`tags`),并在不同的难度等级(`easy`, `medium`, `hard`)下评估 Agent 的表现。表格中的数值是对应类别下所有任务的**平均成功率**。
- **`tags`**: 描述任务所需要的一种或多种核心能力。例如:
- `complex_ui_understanding`: 理解复杂或非标准的界面布局。
- `math_counting`: 需要进行数学计算或计数。
- `transcription`: 需要从一种形式(如图像、视频)转录信息到文本。
- `information_retrieval`: 从屏幕上寻找并提取特定信息。
- **`difficulty`**: 任务的难度级别。
- **数值**: 该 `tag``difficulty` 组合下所有任务的平均成功率。`1.0` 表示100%成功, `0.0` 表示0%成功, 或 `NaN` 表示数据集中没有该组合的任务。
### **2. Agent 的能力画像**
### **核心优势 (Core Strengths)**
Agent 在以下几个方面表现出了近乎完美或非常强的能力:
- **跨应用操作 (`multi_app`)**: 在 `easy` 级别下成功率为 `1.00`。这表明 Agent 能够可靠地在不同应用程序之间进行切换以完成任务。
- **记忆 (`memorization`)**: 在 `easy` 级别下成功率为 `1.00`。Agent 具备有效的短期记忆能力,能够记住先前步骤的信息(例如,一个文件名)并在后续步骤中使用它。
- **搜索 (`search`)**: 在 `medium` 级别下成功率为 `1.00`,在 `easy` 级别下也有 `0.60` 的表现。这说明 Agent 擅长使用应用内或系统级的搜索功能来定位信息。
### **关键短板 (Critical Weaknesses)**
- **转录 (`transcription`)**: **成功率为 `0.00`**。这是最严重的失败,表明 Agent **完全不具备**从图像或视频等非结构化源头准确提取并转录信息到文本字段的能力。这可能源于其视觉模型(Vision Model)在光学字符识别(OCR)上的缺陷。
- **数学/计数 (`math_counting`)**: 在 `easy` 级别下成功率为 **`0.00`**,在中等和困难级别下也仅有 `0.33`。这是一个重大的认知缺陷。Agent 似乎无法在手机操作的情境中执行简单的数学运算或对界面元素进行计数。
- **需要预设 (`requires_setup`)**: 在 `easy` 级别下成功率为 **`0.00`**。Agent 无法处理那些需要先进行特定环境设置(例如,确保某个文件存在或某个设置开启)才能开始的任务。它可能缺乏检查前置条件并根据情况采取修正动作的能力。
- **复杂UI理解 (`complex_ui_understanding`)**: 成功率普遍很低 (`easy` 0.17, `hard` 0.14)。这是另一个核心弱点。Agent 的操作严重依赖于**标准、规范的UI设计**。一旦遇到布局复杂、控件非主流或信息密度高的界面,它就很容易“迷路”,无法准确定位到正确的交互元素。
- **信息检索 (`information_retrieval`)**: 在 `easy` 级别下成功率仅为 `0.17`。这与 `complex_ui_understanding` 弱点高度相关。即使在简单的场景下,如果信息没有以一种简单明了的方式呈现,Agent 也很难从中找到并提取出需要的内容。
---
### 第三部分:总体结论与推断
现在,我们可以将两部分的分析联系起来,形成一个完整的结论。
**`t3a_claude4_sonnet` Agent 的总体画像是:一个在执行标准、线性流程任务方面非常高效的“操作手”,但在需要深度视觉理解、逻辑推理和适应非标准环境等高级认知能力的“思考者”角色上存在明显不足。**
**为什么那些任务会失败?**
- **`SystemWifiTurn...` (tasks 102-105) 和 `Tasks...` (tasks 106-111) 等的失败**:可以高度归因于 **`complex_ui_understanding`** 和 **`information_retrieval`** 的双重失败。系统设置界面、某些设计不佳的应用(可能是 `Tasks` 应用)的UI可能不符合Agent的“预期”,导致它无法找到正确的开关或读取到正确的状态信息。
- **`SimpleSmsReplyMostRecent` (task 82) 的失败**:可能涉及 `information_retrieval`(需要准确识别出“哪一条是最近的”)和 `complex_ui_understanding`
- 所有涉及**数学、转录、需要预设**的任务失败,其根本原因已在第二部分中清晰揭示。
### **下一步的分析与优化建议**
基于以上分析,后续的优化路径非常清晰:
1. **根因分析 (Root Cause Analysis)**: 我们目前是基于统计摘要进行的宏观推断。下一步最关键的操作,就是深入到**具体失败任务的详细操作日志**中。通过逐帧查看 Agent 的“观察 (`Observation`) -> 思考 (`Thought`) -> 行动 (`Action`)”链条,我们可以精确地看到它是在哪一步、因为什么样的错误感知或推理而导致了任务失败。
2. **模型与算法优化 (Model & Algorithm Optimization)**:
- 针对 **`complex_ui`** 问题,需要用更多样化、更复杂的UI布局数据来训练 Agent,或者开发更鲁棒的UI解析模块(例如,将UI元素解析为图结构而非仅仅是位置和文本)。
- 针对 **`transcription`** 和 **`math_counting`** 问题,可能需要在 Agent 的工具集中集成更强大的专用工具,例如一个高精度的OCR服务或一个计算器工具,并教会 Agent 何时以及如何调用这些工具。
```shell
task_num num_complete_trials mean_success_rate mean_episode_length total_runtime_s num_fail_trials
task
AudioRecorderRecordAudio 0 1.00 0.0 10.00 101.80 0.00
AudioRecorderRecordAudioWithFileName 1 1.00 0.0 20.00 326.70 0.00
BrowserDraw 2 1.00 0.0 20.00 391.40 0.00
BrowserMaze 3 1.00 1.0 16.00 195.50 0.00
BrowserMultiply 4 1.00 1.0 15.00 177.30 0.00
CameraTakePhoto 5 1.00 1.0 4.00 41.20 0.00
CameraTakeVideo 6 1.00 1.0 7.00 83.80 0.00
ClockStopWatchPausedVerify 7 1.00 1.0 3.00 30.10 0.00
ClockStopWatchRunning 8 1.00 1.0 4.00 40.80 0.00
ClockTimerEntry 9 1.00 0.0 10.00 127.10 0.00
ContactsAddContact 10 1.00 1.0 8.00 106.60 0.00
ContactsNewContactDraft 11 1.00 1.0 8.00 108.40 0.00
ExpenseAddMultiple 12 1.00 1.0 25.00 356.10 0.00
ExpenseAddMultipleFromGallery 13 1.00 0.0 32.00 455.00 0.00
ExpenseAddMultipleFromMarkor 14 1.00 0.0 24.00 325.70 0.00
ExpenseAddSingle 15 1.00 1.0 10.00 140.40 0.00
ExpenseDeleteDuplicates 16 1.00 1.0 10.00 151.50 0.00
ExpenseDeleteDuplicates2 17 1.00 1.0 15.00 381.10 0.00
ExpenseDeleteMultiple 18 1.00 1.0 11.00 152.90 0.00
ExpenseDeleteMultiple2 19 1.00 1.0 14.00 198.30 0.00
ExpenseDeleteSingle 20 1.00 1.0 5.00 65.30 0.00
FilesDeleteFile 21 1.00 1.0 10.00 129.60 0.00
FilesMoveFile 22 1.00 1.0 13.00 164.00 0.00
MarkorAddNoteHeader 23 1.00 0.0 12.00 175.70 0.00
MarkorChangeNoteContent 24 1.00 0.0 12.00 174.10 0.00
MarkorCreateFolder 25 1.00 0.0 10.00 137.10 0.00
MarkorCreateNote 26 1.00 1.0 13.00 190.50 0.00
MarkorCreateNoteAndSms 27 1.00 0.0 18.00 264.50 0.00
MarkorCreateNoteFromClipboard 28 1.00 0.0 14.00 194.60 0.00
MarkorDeleteAllNotes 29 1.00 1.0 13.00 154.80 0.00
MarkorDeleteNewestNote 30 1.00 0.0 10.00 124.70 0.00
MarkorDeleteNote 31 1.00 0.0 10.00 120.80 0.00
MarkorEditNote 32 1.00 0.0 12.00 169.60 0.00
MarkorMergeNotes 33 1.00 0.0 31.00 452.30 0.00
MarkorMoveNote 34 1.00 0.0 14.00 185.70 0.00
MarkorTranscribeReceipt 35 1.00 0.0 18.00 232.30 0.00
MarkorTranscribeVideo 36 1.00 0.0 20.00 266.70 0.00
NotesIsTodo 37 1.00 1.0 3.00 43.30 0.00
NotesMeetingAttendeeCount 38 1.00 1.0 7.00 87.80 0.00
NotesRecipeIngredientCount 39 1.00 1.0 6.00 82.70 0.00
NotesTodoItemCount 40 1.00 1.0 7.00 102.50 0.00
OpenAppTaskEval 41 1.00 1.0 3.00 35.20 0.00
OsmAndFavorite 42 1.00 1.0 10.00 148.60 0.00
OsmAndMarker 43 1.00 0.0 20.00 275.00 0.00
OsmAndTrack 44 1.00 0.0 41.00 607.30 0.00
RecipeAddMultipleRecipes 45 1.00 1.0 36.00 540.30 0.00
RecipeAddMultipleRecipesFromImage 46 1.00 0.0 24.00 316.70 0.00
RecipeAddMultipleRecipesFromMarkor 47 1.00 0.0 25.00 363.80 0.00
RecipeAddMultipleRecipesFromMarkor2 48 1.00 0.0 44.00 719.30 0.00
RecipeAddSingleRecipe 49 1.00 1.0 13.00 190.90 0.00
RecipeDeleteDuplicateRecipes 50 1.00 0.0 10.00 153.60 0.00
RecipeDeleteDuplicateRecipes2 51 1.00 0.0 24.00 348.60 0.00
RecipeDeleteDuplicateRecipes3 52 1.00 0.0 34.00 433.90 0.00
RecipeDeleteMultipleRecipes 53 1.00 0.0 24.00 328.80 0.00
RecipeDeleteMultipleRecipesWithConstraint 54 1.00 1.0 11.00 164.00 0.00
RecipeDeleteMultipleRecipesWithNoise 55 1.00 1.0 20.00 252.80 0.00
RecipeDeleteSingleRecipe 56 1.00 1.0 6.00 68.50 0.00
RecipeDeleteSingleWithRecipeWithNoise 57 1.00 1.0 7.00 84.60 0.00
RetroCreatePlaylist 58 1.00 1.0 21.00 270.50 0.00
RetroPlayingQueue 59 1.00 1.0 17.00 235.80 0.00
RetroPlaylistDuration 60 1.00 0.0 30.00 386.10 0.00
RetroSavePlaylist 61 1.00 1.0 27.00 321.30 0.00
SaveCopyOfReceiptTaskEval 62 1.00 1.0 11.00 132.40 0.00
SimpleCalendarAddOneEvent 63 1.00 0.0 14.00 202.30 0.00
SimpleCalendarAddOneEventInTwoWeeks 64 1.00 0.0 18.00 233.70 0.00
SimpleCalendarAddOneEventRelativeDay 65 1.00 0.0 17.00 231.50 0.00
SimpleCalendarAddOneEventTomorrow 66 1.00 0.0 18.00 235.70 0.00
SimpleCalendarAddRepeatingEvent 67 1.00 0.0 16.00 224.10 0.00
SimpleCalendarAnyEventsOnDate 68 1.00 0.0 10.00 164.80 0.00
SimpleCalendarDeleteEvents 69 1.00 0.0 14.00 210.00 0.00
SimpleCalendarDeleteEventsOnRelativeDay 70 1.00 0.0 3.00 45.60 0.00
SimpleCalendarDeleteOneEvent 71 1.00 0.0 12.00 186.20 0.00
SimpleCalendarEventOnDateAtTime 72 1.00 0.0 6.00 104.20 0.00
SimpleCalendarEventsInNextWeek 73 1.00 0.0 6.00 85.80 0.00
SimpleCalendarEventsInTimeRange 74 1.00 0.0 10.00 174.50 0.00
SimpleCalendarEventsOnDate 75 1.00 1.0 6.00 198.00 0.00
SimpleCalendarFirstEventAfterStartTime 76 1.00 0.0 10.00 167.50 0.00
SimpleCalendarLocationOfEvent 77 1.00 1.0 6.00 110.80 0.00
SimpleCalendarNextEvent 78 1.00 0.0 7.00 105.50 0.00
SimpleCalendarNextMeetingWithPerson 79 1.00 0.0 4.00 60.60 0.00
SimpleDrawProCreateDrawing 80 1.00 0.0 18.00 307.90 0.00
SimpleSmsReply 81 1.00 0.0 9.00 211.30 0.00
SimpleSmsReplyMostRecent 82 0.00 NaN NaN 17.30 1.00
SimpleSmsResend 83 1.00 0.0 8.00 142.40 0.00
SimpleSmsSend 84 1.00 0.0 8.00 132.70 0.00
SimpleSmsSendClipboardContent 85 1.00 0.0 8.00 132.50 0.00
SimpleSmsSendReceivedAddress 86 1.00 0.0 18.00 276.50 0.00
SportsTrackerActivitiesCountForWeek 87 1.00 0.0 10.00 180.20 0.00
SportsTrackerActivitiesOnDate 88 1.00 0.0 5.00 83.40 0.00
SportsTrackerActivityDuration 89 1.00 0.0 10.00 156.40 0.00
SportsTrackerLongestDistanceActivity 90 1.00 0.0 10.00 191.60 0.00
SportsTrackerTotalDistanceForCategoryOverInterval 91 1.00 0.0 20.00 351.00 0.00
SportsTrackerTotalDurationForCategoryThisWeek 92 1.00 0.0 10.00 187.50 0.00
SystemBluetoothTurnOff 93 1.00 1.0 8.00 90.20 0.00
SystemBluetoothTurnOffVerify 94 1.00 1.0 7.00 82.20 0.00
SystemBluetoothTurnOn 95 1.00 1.0 5.00 68.80 0.00
SystemBluetoothTurnOnVerify 96 1.00 1.0 4.00 57.10 0.00
SystemBrightnessMax 97 1.00 0.0 10.00 120.50 0.00
SystemBrightnessMaxVerify 98 1.00 0.0 10.00 131.20 0.00
SystemBrightnessMin 99 1.00 0.0 10.00 125.60 0.00
SystemBrightnessMinVerify 100 1.00 0.0 10.00 125.80 0.00
SystemCopyToClipboard 101 1.00 0.0 5.00 79.30 0.00
SystemWifiTurnOff 102 1.00 0.0 10.00 146.10 0.00
SystemWifiTurnOffVerify 103 0.00 NaN NaN 29.10 1.00
SystemWifiTurnOn 104 0.00 NaN NaN 29.90 1.00
SystemWifiTurnOnVerify 105 0.00 NaN NaN 28.80 1.00
TasksCompletedTasksForDate 106 0.00 NaN NaN 40.30 1.00
TasksDueNextWeek 107 0.00 NaN NaN 30.60 1.00
TasksDueOnDate 108 0.00 NaN NaN 30.30 1.00
TasksHighPriorityTasks 109 0.00 NaN NaN 33.80 1.00
TasksHighPriorityTasksDueOnDate 110 0.00 NaN NaN 29.20 1.00
TasksIncompleteTasksOnDate 111 0.00 NaN NaN 30.40 1.00
TurnOffWifiAndTurnOnBluetooth 112 0.00 NaN NaN 29.30 1.00
TurnOnWifiAndOpenApp 113 0.00 NaN NaN 39.20 1.00
VlcCreatePlaylist 114 0.00 NaN NaN 9.10 1.00
VlcCreateTwoPlaylists 115 0.00 NaN NaN 9.30 1.00
========= Average ========= 0 0.88 0.4 13.45 174.96 0.12
mean_success_rate
difficulty easy medium hard
tags
complex_ui_understanding 0.17 0.5 0.14
data_edit 0.36 0.5 0.0
data_entry 0.27 0.44 0.0
game_playing 0.50 - -
information_retrieval 0.17 0.4 0.0
math_counting 0.00 0.33 0.33
memorization 1.00 0.0 0.25
multi_app 1.00 0.0 0.0
parameterized 0.37 0.44 0.22
repetition 0.50 0.5 0.4
requires_setup 0.00 0.5 0.0
screen_reading 0.40 0.5 0.33
search 0.60 1.0 0.17
transcription 0.00 0.0 0.0
untagged 0.60 1.0 -
verification 0.60 - -
```
+474
View File
@@ -0,0 +1,474 @@
"""Focused, offline checks for Experiment 7-12 evidence/reporting."""
from __future__ import annotations
from argparse import Namespace
import sqlite3
import pytest
from experiment_core import (
BASELINE_TASK_COUNT,
aggregate_episodes,
choose_decision,
choose_efficiency_decision,
enforce_scope_claims,
paired_rows,
redact_text,
render_report,
)
from run_controlled_experiment import (
_context_safe_output_cap,
_missing_retro_queue_as_empty,
_read_nonempty_with_retry,
_retry_clipper_foreground,
_truncate_current_ui_section,
_validate_resume_evidence,
)
def _episode(arm: str, task: str, success: bool, latency: float) -> dict:
return {
"pair_id": task + ":trial-1",
"task": task,
"trial": 1,
"arm": arm,
"status": "completed",
"success": success,
"evaluator_reward": float(success),
"steps": 3 if success else 10,
"elapsed_s": latency,
"llm": {"calls": 4, "input_tokens": 100, "output_tokens": 20},
}
def test_redaction_covers_explicit_and_pattern_credentials() -> None:
secret = "definitely-not-for-output"
text = redact_text(
f"api_key={secret} Authorization: Bearer abcdefghijk sk-example123456789",
[secret],
)
assert secret not in text
assert "abcdefghijk" not in text
assert "sk-example123456789" not in text
assert text.count("[REDACTED]") >= 3
def test_retro_missing_queue_schema_becomes_empty_observation() -> None:
def missing_queue(_env: object) -> list[str]:
raise sqlite3.OperationalError("no such table: playing_queue")
assert _missing_retro_queue_as_empty(missing_queue)(object()) == []
def test_retro_compatibility_does_not_hide_other_sqlite_errors() -> None:
def corrupt_database(_env: object) -> list[str]:
raise sqlite3.OperationalError("database disk image is malformed")
with pytest.raises(sqlite3.OperationalError, match="malformed"):
_missing_retro_queue_as_empty(corrupt_database)(object())
def test_context_cap_keeps_headroom_for_provider_lower_bound() -> None:
error = (
"This model's maximum context length is 32768 tokens. However, you "
"requested 1024 output tokens and your prompt contains at least 31745 "
"input tokens."
)
assert _context_safe_output_cap(error, 1024) == 991
assert _context_safe_output_cap("unrelated provider error", 1024) is None
def test_context_truncation_is_limited_to_middle_of_current_ui() -> None:
prefix = "prefix and goal"
ui = "A" * 9000 + "M" * 16384 + "Z" * 9000
suffix = "guidance and output format"
prompt = (
prefix
+ "\n\nHere is a list of descriptions for some UI elements on the current screen:\n"
+ ui
+ "\nHere are some useful guidelines you need to follow:\n"
+ suffix
)
result = _truncate_current_ui_section(prompt)
assert result is not None
truncated, removed = result
assert prefix in truncated and suffix in truncated
assert "A" * 1000 in truncated and "Z" * 1000 in truncated
assert removed > 0
assert len(truncated) < len(prompt)
def test_context_truncation_handles_before_and_after_summary_ui() -> None:
before = "B" * 12000
after = "A" * 12000
prompt = (
"goal and summary rules\n"
"Here is the description for the before screenshot:\n"
+ before
+ "\nHere is the description for the after screenshot:\n"
+ after
+ "\nThis is the action you picked: click\nBased on the reason: test"
)
result = _truncate_current_ui_section(prompt)
assert result is not None
truncated, removed = result
assert "goal and summary rules" in truncated
assert "This is the action you picked: click" in truncated
assert "B" * 500 in truncated and "A" * 500 in truncated
assert removed > 0
def test_sms_inbox_poll_preserves_empty_then_observed_result() -> None:
reads = iter([[], [], ["Row: 0, address=123, body=hello"]])
assert _read_nonempty_with_retry(
lambda: next(reads), attempts=3, delay_s=0
) == ["Row: 0, address=123, body=hello"]
def test_clipper_retry_is_limited_to_exact_foreground_error() -> None:
attempts = iter([
RuntimeError(
"Clipper app must be in the foreground to access clipboard. "
"Additionally, app privileges must be granted manually."
),
"clipboard value",
])
def flaky_call() -> str:
result = next(attempts)
if isinstance(result, Exception):
raise result
return result
assert _retry_clipper_foreground(flaky_call, delay_s=0) == "clipboard value"
with pytest.raises(RuntimeError, match="unrelated"):
_retry_clipper_foreground(
lambda: (_ for _ in ()).throw(RuntimeError("unrelated")), delay_s=0
)
def test_paired_comparison_and_conservative_candidate_decision() -> None:
episodes = []
for index in range(4):
task = f"wifi-{index}"
episodes.extend([
_episode("control", task, index > 0, 10.0),
_episode("treatment", task, True, 11.0),
])
summary = aggregate_episodes(episodes)
pairs = paired_rows(episodes)
decision = choose_decision(summary, pairs)
assert len(pairs) == 4
assert decision["net_success_delta"] == 1
assert decision["paired_regressions"] == 0
assert decision["promote_to_full_suite_candidate"] is True
assert decision["outcome"] == "promote_candidate_to_full_suite_rerun"
def test_subset_can_never_claim_full_suite_completion() -> None:
evidence = {
"scope": {
"mode": "candidate_rerun",
"tasks": ["a", "b", "c", "d"],
"trials_per_task": 5,
"completed_episodes": 20,
"error_episodes": 0,
},
"decision": {"source_paired_run_id": "paired-real"},
}
enforce_scope_claims(evidence)
assert evidence["scope"]["full_suite_completed"] is False
assert evidence["experiment_complete"] is False
def test_full_suite_gate_requires_direct_116_by_5_evidence() -> None:
tasks = [f"task-{index}" for index in range(BASELINE_TASK_COUNT)]
episodes = [
{
"task": task,
"trial": trial,
"pair_seed": task_index * 1009 + trial,
"arm": "candidate",
"status": "completed",
"evaluator_reward": 1.0,
}
for task_index, task in enumerate(tasks)
for trial in range(1, 6)
]
evidence = {
"scope": {
"mode": "candidate_rerun",
"tasks": tasks,
"trials_per_task": 5,
"completed_episodes": BASELINE_TASK_COUNT * 5,
"error_episodes": 0,
},
"episodes": episodes,
"decision": {"source_paired_run_id": "paired-real"},
"environment": {
"api_level": 33,
"emulator_setup_completed": True,
"app_provisioning": {"complete": True},
},
}
enforce_scope_claims(evidence)
assert evidence["scope"]["full_suite_completed"] is True
assert evidence["experiment_complete"] is True
def test_full_suite_gate_requires_reference_api_and_apps() -> None:
tasks = [f"task-{index}" for index in range(BASELINE_TASK_COUNT)]
episodes = [
{
"task": task,
"trial": trial,
"pair_seed": task_index * 1009 + trial,
"arm": "candidate",
"status": "completed",
"evaluator_reward": 0.0,
}
for task_index, task in enumerate(tasks)
for trial in range(1, 6)
]
evidence = {
"scope": {
"mode": "candidate_rerun",
"tasks": tasks,
"trials_per_task": 5,
},
"episodes": episodes,
"decision": {"source_paired_run_id": "paired-real"},
"environment": {
"api_level": 35,
"emulator_setup_completed": False,
"app_provisioning": {"complete": False},
},
}
enforce_scope_claims(evidence)
assert evidence["scope"]["direct_episode_gate_completed"] is True
assert evidence["scope"]["full_suite_completed"] is False
assert evidence["experiment_complete"] is False
def test_full_suite_gate_rejects_counters_without_direct_episodes() -> None:
evidence = {
"scope": {
"mode": "candidate_rerun",
"tasks": [f"task-{index}" for index in range(BASELINE_TASK_COUNT)],
"trials_per_task": 5,
"completed_episodes": BASELINE_TASK_COUNT * 5,
"error_episodes": 0,
},
"episodes": [],
"decision": {"source_paired_run_id": "paired-real"},
}
enforce_scope_claims(evidence)
assert evidence["scope"]["direct_episode_gate_completed"] is False
assert evidence["scope"]["full_suite_completed"] is False
assert evidence["experiment_complete"] is False
def test_success_gain_over_cost_guardrail_is_not_promoted() -> None:
episodes = []
for index in range(4):
task = f"wifi-{index}"
control = _episode("control", task, index > 0, 10.0)
treatment = _episode("treatment", task, True, 20.0)
treatment["llm"]["input_tokens"] = 1000
treatment["llm"]["output_tokens"] = 200
episodes.extend([control, treatment])
summary = aggregate_episodes(episodes)
decision = choose_decision(summary, paired_rows(episodes))
assert decision["outcome"] == "restrict_candidate_due_to_cost"
assert decision["guardrails"]["passed"] is False
assert decision["promote_to_full_suite_candidate"] is False
assert decision["deployment_approved"] is False
def test_efficiency_refinement_can_promote_without_inventing_success_gain() -> None:
episodes = []
for index in range(4):
task = f"wifi-{index}"
control = _episode("control", task, True, 10.0)
treatment = _episode("treatment", task, True, 9.0)
treatment["llm"]["input_tokens"] = 40
treatment["llm"]["output_tokens"] = 10
episodes.extend([control, treatment])
decision = choose_efficiency_decision(
aggregate_episodes(episodes), paired_rows(episodes)
)
assert decision["net_success_delta"] == 0
assert decision["paired_regressions"] == 0
assert decision["guardrails"]["passed"] is True
assert decision["promote_to_full_suite_candidate"] is True
assert decision["deployment_approved"] is False
def test_efficiency_refinement_rejects_cheap_but_unsuccessful_treatment() -> None:
episodes = []
for index in range(4):
task = f"wifi-{index}"
control = _episode("control", task, False, 10.0)
treatment = _episode("treatment", task, False, 9.0)
treatment["llm"]["input_tokens"] = 40
treatment["llm"]["output_tokens"] = 10
episodes.extend([control, treatment])
decision = choose_efficiency_decision(
aggregate_episodes(episodes), paired_rows(episodes)
)
assert decision["mean_token_ratio_treatment_over_control"] < 0.75
assert decision["success_preservation_passed"] is False
assert decision["guardrails"]["passed"] is False
assert decision["outcome"] == "reject_efficiency_candidate_due_to_regression"
assert decision["promote_to_full_suite_candidate"] is False
def test_report_labels_historical_and_hypothetical_numbers() -> None:
evidence = {
"run_id": "test-run",
"generated_at_utc": "2026-07-29T00:00:00Z",
"environment": {
"android_world_commit": "abc123",
"device_model": "emulator",
"api_level": 35,
"upstream_tested_api_level": 33,
},
"model": {"provider": "real-provider", "model": "real-model"},
"scope": {
"tasks": ["SystemWifiTurnOn"],
"trials_per_task": 1,
"mode": "paired",
"full_suite_completed": False,
},
"diagnosis": {"findings": ["Historical finding."]},
"hypothesis": {
"id": "H1",
"change": "Add a task guideline.",
"expected_result": "Improve paired reward.",
"guardrails": "Same task and model.",
},
"arm_summary": {},
"paired_comparison": [],
"decision": {
"outcome": "insufficient_evidence",
"reason": "Need four pairs.",
},
"episodes": [],
"environment_boundaries": ["API mismatch."],
"llm_analysis": {
"status": "completed",
"summary": "Observed subset summary.",
"observed_failure_pattern": ["One bounded residual pattern."],
"cost_benefit_interpretation": "No deployment approval.",
"next_hypothesis": {
"id": "H5",
"layer": "middle",
"idea": "Test the input path.",
"target": "One paired gain.",
"verification": "Matched paired run.",
},
},
}
report = render_report(evidence)
assert "historical input evidence" in report
assert "explicitly hypothetical" in report
assert "not the complete AndroidWorld benchmark" in report
assert "Full 116-task × 5-seed suite completed: **false**" in report
assert "Observed subset summary." in report
assert "No deployment approval." in report
def test_resume_rejects_changed_configuration() -> None:
evidence = {
"experiment": "7-12",
"hypothesis": {"id": "H5"},
"scope": {
"mode": "paired",
"tasks": ["SystemWifiTurnOff"],
"trials_per_task": 1,
"max_steps": 10,
},
"model": {
"model": "real-model",
"seed": 42,
"provider": "real-provider",
"base_url": "https://provider.invalid/v1",
"max_tokens": 1024,
},
"environment": {
"skip_device_time": True,
"device_serial": "emulator-5554",
"grpc_port": 8554,
},
"episodes": [],
}
args = Namespace(
tasks="SystemWifiTurnOff",
hypothesis="H5",
mode="paired",
trials=1,
max_steps=11,
model="real-model",
model_seed=42,
provider="real-provider",
base_url="https://provider.invalid/v1",
max_model_tokens=1024,
transition_pause=None,
skip_device_time=True,
console_port=5554,
grpc_port=8554,
seed=42,
)
with pytest.raises(RuntimeError, match="max_steps"):
_validate_resume_evidence(evidence, args)
def test_resume_rejects_changed_pair_seed() -> None:
evidence = {
"experiment": "7-12",
"hypothesis": {"id": "H5C"},
"scope": {
"mode": "paired",
"tasks": ["SystemWifiTurnOff"],
"trials_per_task": 1,
"max_steps": 10,
},
"model": {
"model": "real-model",
"seed": 42,
"provider": "real-provider",
"base_url": "https://provider.invalid/v1",
"max_tokens": 1024,
},
"environment": {
"skip_device_time": True,
"device_serial": "emulator-5554",
"grpc_port": 8554,
},
"episodes": [{
"task": "SystemWifiTurnOff",
"trial": 1,
"arm": "control",
"pair_seed": 42,
}],
}
args = Namespace(
tasks="SystemWifiTurnOff",
hypothesis="H5C",
mode="paired",
trials=1,
max_steps=10,
model="real-model",
model_seed=42,
provider="real-provider",
base_url="https://provider.invalid/v1",
max_model_tokens=1024,
transition_pause=None,
skip_device_time=True,
console_port=5554,
grpc_port=8554,
seed=43,
)
with pytest.raises(RuntimeError, match="Resume seed mismatch"):
_validate_resume_evidence(evidence, args)
@@ -0,0 +1,652 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-merged-20260804T092058Z`
- Generated (UTC): `2026-08-04T09:20:58Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **true**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed policy, step budget, Pixel 6/API-33 device class, upstream setup, and app versions across isolated shards.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 580/580 | 0.045 | 0.134 | 9.672 | 109.861 | 18.998 | 169069.564 | 97384410 / 675937 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`full_candidate_rerun_completed`**
- Reason: The direct 116-task x five-trial candidate rerun completed on five independent reference-environment shards. Negative evaluator results are retained.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
The complete 116-task, five-trial candidate rerun gate is satisfied by direct episode evidence.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudio / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudio / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / AudioRecorderRecordAudio / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudio / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / AudioRecorderRecordAudioWithFileName / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / BrowserMultiply / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchPausedVerify / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / ClockStopWatchPausedVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchPausedVerify / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchRunning / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchRunning / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchRunning / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchRunning / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchRunning / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / ContactsAddContact / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / ContactsAddContact / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / ContactsAddContact / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / ExpenseAddSingle / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteDuplicates / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteDuplicates2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteMultiple2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteSingle / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteSingle / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteSingle / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteSingle / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / FilesDeleteFile / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateFolder / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateFolder / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateFolder / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateFolder / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorDeleteAllNotes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorDeleteNewestNote / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorDeleteNewestNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorDeleteNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorEditNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / NotesIsTodo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesIsTodo / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / NotesIsTodo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / NotesIsTodo / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / NotesMeetingAttendeeCount / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesMeetingAttendeeCount / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesMeetingAttendeeCount / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / NotesRecipeIngredientCount / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesRecipeIngredientCount / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / OpenAppTaskEval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / OpenAppTaskEval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndMarker / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndMarker / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndTrack / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleRecipe / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / RetroCreatePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroCreatePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroCreatePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroCreatePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroCreatePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarAnyEventsOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarAnyEventsOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarAnyEventsOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarEventsInTimeRange / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarLocationOfEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleDrawProCreateDrawing / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleDrawProCreateDrawing / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleDrawProCreateDrawing / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleDrawProCreateDrawing / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleDrawProCreateDrawing / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsReply / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsReplyMostRecent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsResend / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSend / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSendClipboardContent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSendReceivedAddress / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivitiesCountForWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivitiesOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivityDuration / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerLongestDistanceActivity / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOffVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOffVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOffVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOffVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOn / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOn / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOn / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOn / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOnVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOnVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOnVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOnVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMax / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMax / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMax / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMax / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMax / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMaxVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMaxVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMaxVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMin / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMinVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMinVerify / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / SystemCopyToClipboard / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemCopyToClipboard / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemCopyToClipboard / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SystemCopyToClipboard / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SystemCopyToClipboard / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOff / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOff / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOffVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOffVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOffVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOn / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOn / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOn / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOn / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOnVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOnVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOnVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / TasksCompletedTasksForDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksCompletedTasksForDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksCompletedTasksForDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksCompletedTasksForDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksCompletedTasksForDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / TasksHighPriorityTasksDueOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 5`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate rerun on 116 tasks across five independent shards showed a success rate of 4.48% with 26 successes out of 580 episodes. The mean latency was 109.86 seconds, and the mean number of LLM calls was 19. The mean total tokens were 169,069.56, with a mean evaluator reward of 0.133621.
- Cost/benefit interpretation: The cost of running the candidate on 116 tasks was relatively high, with a mean latency of 109.86 seconds and a mean of 19 LLM calls per episode. The benefit, in terms of success rate, was minimal, with only 26 successes out of 580 episodes. The low success rate and high cost suggest that the current approach may not be efficient or effective.
- Residual pattern: Multiple tasks, such as AudioRecorderRecordAudio, CameraTakePhoto, and ExpenseAddMultiple, consistently failed across multiple episodes.
- Residual pattern: The failure rate was particularly high for tasks involving complex interactions with the UI, such as ContactsAddContact and ExpenseDeleteMultiple2.
- Residual pattern: The evaluator reward was low for most episodes, indicating that the tasks were not being completed as expected.
- Next hypothesis `H5C` (middle): Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements. Target: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. Verification: Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes 8,192 characters only from the middle of the current-screen indexed UI section. Prompt prefix, goal, history, leading and trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 8,192 characters for action selection or 4,096 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 16,384 characters for action selection or 8,192 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- Run stopped by environment/runtime blocker: Generated params do not match checkpoint for MarkorCreateFolder:trial-4:seed-25270
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This user-requested local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,209 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260804T045559Z`
- Generated (UTC): `2026-08-04T09:13:52Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 116/116 | 0.026 | 0.129 | 9.698 | 110.113 | 19.103 | 174058.216 | 20054856 / 135897 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`candidate_subset_rerun_completed`**
- Reason: A real modified candidate subset rerun completed, but it is not the 116-task × five-trial gate and cannot approve deployment.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchRunning / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteMultiple / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteSingle / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / FilesDeleteFile / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateFolder / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / OsmAndFavorite / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndTrack / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / RetroCreatePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarDeleteEvents / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleDrawProCreateDrawing / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsResend / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivitiesCountForWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOnVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMax / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMaxVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMin / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMinVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOff / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOffVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOnVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `candidate / TasksCompletedTasksForDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 1`: agent declared completion, but the real evaluator state failed
- `candidate / SystemCopyToClipboard / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesIsTodo / trial 1`: evaluator reward / completion gate was not satisfied
- `candidate / NotesMeetingAttendeeCount / trial 1`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate subset of 116 tasks was rerun, achieving a success rate of 2.59% with an estimated cost of $0.00. The mean latency was 110.11 seconds, and the mean number of LLM calls was 19.10. The mean total tokens used were 174,058.22.
- Cost/benefit interpretation: The cost of running the 116 tasks was minimal, with an estimated cost of $0.00. However, the low success rate of 2.59% indicates that the current approach is not cost-effective. The high latency and token usage suggest that the model may need optimization to reduce computational overhead.
- Residual pattern: Most tasks failed to achieve success, with only 3 out of 116 tasks succeeding.
- Residual pattern: Tasks involving complex interactions with the UI, such as 'RetroPlayingQueue' and 'SportsTrackerTotalDistanceForCategoryOverInterval', had the highest failure rates.
- Residual pattern: Tasks that required multiple steps or complex conditions, like 'RecipeDeleteMultipleRecipesWithNoise', were particularly challenging.
- Next hypothesis `H5C` (middle): Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements. Target: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. Verification: Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes 8,192 characters only from the middle of the current-screen indexed UI section. Prompt prefix, goal, history, leading and trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 8,192 characters for action selection or 4,096 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 16,384 characters for action selection or 8,192 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,206 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260804T045559Z`
- Generated (UTC): `2026-08-04T09:14:33Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 116/116 | 0.034 | 0.121 | 9.569 | 106.154 | 18.690 | 163049.526 | 18780790 / 132955 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`candidate_subset_rerun_completed`**
- Reason: A real modified candidate subset rerun completed, but it is not the 116-task × five-trial gate and cannot approve deployment.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / CameraTakePhoto / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / ClockStopWatchRunning / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseAddMultiple / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / ExpenseDeleteDuplicates / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteSingle / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / ContactsNewContactDraft / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateFolder / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorDeleteNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorMergeNotes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndTrack / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleRecipe / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / RetroCreatePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarDeleteEvents / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleDrawProCreateDrawing / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSend / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSendClipboardContent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivityDuration / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMax / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMaxVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMin / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemCopyToClipboard / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOffVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOn / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOnVerify / trial 2`: final evaluator state passed, but the agent never declared completion
- `candidate / TasksCompletedTasksForDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 2`: agent declared completion, but the real evaluator state failed
- `candidate / TasksIncompleteTasksOnDate / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesMeetingAttendeeCount / trial 2`: evaluator reward / completion gate was not satisfied
- `candidate / NotesRecipeIngredientCount / trial 2`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate subset of 116 tasks was rerun, and all tasks completed without errors. The success rate was 3.45%, with 4 successes out of 116 episodes. The mean latency was 106.15 seconds, and the mean total tokens were 163,049.53.
- Cost/benefit interpretation: The cost in terms of latency and token usage is high, but the benefit in terms of preserving paired success without regression is maintained. However, the low success rate suggests that the current approach may not be effective.
- Residual pattern: Most tasks failed to achieve success as defined by the evaluator reward.
- Residual pattern: Tasks such as 'SystemBrightnessMax', 'SystemBrightnessMin', and 'SystemBrightnessMinVerify' had high latency and token usage.
- Next hypothesis `H5C` (middle): Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements. Target: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. Verification: Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes 8,192 characters only from the middle of the current-screen indexed UI section. Prompt prefix, goal, history, leading and trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 8,192 characters for action selection or 4,096 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,203 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260804T045559Z`
- Generated (UTC): `2026-08-04T09:14:42Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 116/116 | 0.052 | 0.142 | 9.698 | 108.721 | 19.078 | 166715.009 | 19204766 / 134175 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`candidate_subset_rerun_completed`**
- Reason: A real modified candidate subset rerun completed, but it is not the 116-task × five-trial gate and cannot approve deployment.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / AudioRecorderRecordAudioWithFileName / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchRunning / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteSingle / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / FilesDeleteFile / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateFolder / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorDeleteNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / RetroCreatePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarNextEvent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleDrawProCreateDrawing / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOnVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMax / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMaxVerify / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMin / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SystemWifiTurnOn / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / TasksCompletedTasksForDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / NotesRecipeIngredientCount / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleSmsSendClipboardContent / trial 3`: evaluator reward / completion gate was not satisfied
- `candidate / SystemCopyToClipboard / trial 3`: agent declared completion, but the real evaluator state failed
- `candidate / NotesIsTodo / trial 3`: final evaluator state passed, but the agent never declared completion
- `candidate / NotesTodoItemCount / trial 3`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate subset of 116 tasks was rerun, achieving a success rate of 5.17% with 6 successes out of 116 episodes. The mean latency was 108.72 seconds, and the mean total tokens were 166,715. The cost-benefit analysis indicates that while the candidate approach preserved success with no regression, it did not meet the latency and token efficiency guardrails.
- Cost/benefit interpretation: The candidate approach preserved success with no regression, but it did not meet the latency and token efficiency guardrails. The higher latency and token usage suggest that the candidate approach may be less efficient than the paired comparison, which could impact the overall cost and performance of the system.
- Residual pattern: Most tasks failed to achieve success, with only a few tasks (e.g., OpenAppTaskEval, SimpleCalendarNextMeetingWithPerson) succeeding.
- Residual pattern: The mean latency and mean total tokens were higher than the guardrails specified for the paired comparison.
- Residual pattern: The candidate approach did not meet the latency and token efficiency requirements, as the mean latency was 1.5 times the paired comparison and the mean tokens were 0.75 times the raw-UIAutomator tokens.
- Next hypothesis `H5C_extended` (middle): Further refine the input pipeline by optimizing the element list to reduce latency and token usage while maintaining success. Target: Reduce the mean latency to 1.25 times the paired comparison and the mean tokens to 0.65 times the raw-UIAutomator tokens. Verification: Conduct a full-suite candidate rerun with the optimized element list to validate the hypothesis.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: 8,192 characters for action selection or 4,096 from each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,204 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260804T045559Z`
- Generated (UTC): `2026-08-04T09:19:57Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 116/116 | 0.052 | 0.129 | 9.716 | 109.254 | 19.069 | 169894.517 | 19570600 / 137164 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`candidate_subset_rerun_completed`**
- Reason: A real modified candidate subset rerun completed, but it is not the 116-task × five-trial gate and cannot approve deployment.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / AudioRecorderRecordAudioWithFileName / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchRunning / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsAddContact / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseAddMultiple / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateFolder / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorCreateNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / MarkorDeleteNewestNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / MarkorEditNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / ContactsNewContactDraft / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / OsmAndFavorite / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndTrack / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithConstraint / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroCreatePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarDeleteEvents / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SimpleCalendarEventsOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleDrawProCreateDrawing / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesCountForWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerLongestDistanceActivity / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOnVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMax / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOffVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemWifiTurnOn / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOnVerify / trial 4`: final evaluator state passed, but the agent never declared completion
- `candidate / TasksCompletedTasksForDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / SystemCopyToClipboard / trial 4`: agent declared completion, but the real evaluator state failed
- `candidate / NotesIsTodo / trial 4`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 4`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate subset rerun completed with 116 episodes, achieving a success rate of 5.17% and an estimated cost of $0.00. The mean latency was 109.25 seconds, and the mean number of LLM calls was 19.07.
- Cost/benefit interpretation: The cost of running the candidate subset is minimal, with an estimated cost of $0.00. However, the success rate is low, and the latency and number of LLM calls are high, indicating that the current approach may not be efficient or effective.
- Residual pattern: Most tasks failed to achieve success, with only a few tasks (e.g., NotesMeetingAttendeeCount, SystemBrightnessMaxVerify) succeeding.
- Residual pattern: The majority of tasks took longer than the mean latency of 109.25 seconds, indicating potential inefficiencies.
- Residual pattern: The mean number of LLM calls per task was 19.07, which is relatively high, suggesting that the model may be making multiple calls to achieve a task.
- Next hypothesis `H5C_extended` (middle): Implement a more aggressive filtering of the UIAutomator hierarchy to reduce the number of LLM calls and improve latency while maintaining success rates. Target: Reduce the mean number of LLM calls per task to 10 and decrease the mean latency to 80 seconds. Verification: Run a full-suite candidate rerun with the new filtering strategy and compare the success rate, latency, and number of LLM calls to the current results.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Run stopped by environment/runtime blocker: Generated params do not match checkpoint for MarkorCreateFolder:trial-4:seed-25270
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes 8,192 characters only from the middle of the current-screen indexed UI section. Prompt prefix, goal, history, leading and trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,202 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260804T045559Z`
- Generated (UTC): `2026-08-04T09:13:41Z`
- Upstream commit: `d9c569f764b3a5629321858de03ff653d0f24056`
- Device: `sdk_gphone64_x86_64`, API `33` (upstream tested reference: API `33`)
- Observation method: `uiautomator_compact`
- Provider/model: `local-vllm` / `qwen2.5-7b-instruct-local`
- Model source/runtime: `local_gpu` / `vllm-0.19.0`
- Accelerator: `NVIDIA_RTX_PRO_6000_Blackwell_96GB`
- Required apps: `24/24`
- Scope: 116 task(s), 5 trial(s), mode `candidate_rerun`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens | Est. cost (USD) |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| candidate | 116/116 | 0.060 | 0.147 | 9.681 | 115.062 | 19.052 | 171630.552 | 19773398 / 135746 | 0.000000 |
## 4. Data-driven decision
- Outcome: **`candidate_subset_rerun_completed`**
- Reason: A real modified candidate subset rerun completed, but it is not the 116-task × five-trial gate and cannot approve deployment.
- Treatment/control mean latency ratio: n/a
- Treatment/control mean token ratio: n/a
- Treatment/control mean LLM-call ratio: n/a
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `candidate / AudioRecorderRecordAudio / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / AudioRecorderRecordAudioWithFileName / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserDraw / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMaze / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / BrowserMultiply / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakePhoto / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / CameraTakeVideo / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ClockStopWatchPausedVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ClockStopWatchRunning / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ClockTimerEntry / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / ContactsAddContact / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultiple / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromGallery / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddSingle / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteDuplicates / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteDuplicates2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteMultiple / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ExpenseDeleteMultiple2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseDeleteSingle / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / ContactsNewContactDraft / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / ExpenseAddMultipleFromMarkor / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / FilesDeleteFile / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / FilesMoveFile / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorAddNoteHeader / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorChangeNoteContent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteAndSms / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorCreateNoteFromClipboard / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteAllNotes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNewestNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorDeleteNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorEditNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMergeNotes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorMoveNote / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeReceipt / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / MarkorTranscribeVideo / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / NotesTodoItemCount / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OpenAppTaskEval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndFavorite / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / OsmAndMarker / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / OsmAndTrack / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromImage / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddMultipleRecipesFromMarkor2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeAddSingleRecipe / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes2 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteDuplicateRecipes3 / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipes / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteMultipleRecipesWithNoise / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RecipeDeleteSingleWithRecipeWithNoise / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / RetroCreatePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlaylistDuration / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroSavePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SaveCopyOfReceiptTaskEval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventInTwoWeeks / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventRelativeDay / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddOneEventTomorrow / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAddRepeatingEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarAnyEventsOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEvents / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteEventsOnRelativeDay / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarDeleteOneEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventOnDateAtTime / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInNextWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsInTimeRange / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarEventsOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarFirstEventAfterStartTime / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarLocationOfEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextEvent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleCalendarNextMeetingWithPerson / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SimpleDrawProCreateDrawing / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReply / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsResend / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSend / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendClipboardContent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsSendReceivedAddress / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / SportsTrackerActivitiesCountForWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivitiesOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerActivityDuration / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerLongestDistanceActivity / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDistanceForCategoryOverInterval / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SportsTrackerTotalDurationForCategoryThisWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOff / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOffVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBluetoothTurnOn / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBluetoothTurnOnVerify / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / SystemBrightnessMax / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMin / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemBrightnessMinVerify / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / SystemCopyToClipboard / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOff / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SystemWifiTurnOn / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / TasksCompletedTasksForDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueNextWeek / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksDueOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasks / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksHighPriorityTasksDueOnDate / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TasksIncompleteTasksOnDate / trial 5`: agent declared completion, but the real evaluator state failed
- `candidate / TurnOffWifiAndTurnOnBluetooth / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / TurnOnWifiAndOpenApp / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreatePlaylist / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / VlcCreateTwoPlaylists / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / RetroPlayingQueue / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / SimpleSmsReplyMostRecent / trial 5`: evaluator reward / completion gate was not satisfied
- `candidate / NotesIsTodo / trial 5`: final evaluator state passed, but the agent never declared completion
- `candidate / NotesMeetingAttendeeCount / trial 5`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: The candidate subset of 116 tasks was rerun, achieving a success rate of 6.03% with 7 successes out of 116 episodes. The mean latency was 115.06 seconds, and the mean total tokens were 171,630.55. The cost-benefit analysis indicates that while the candidate subset showed some success, the high latency and token usage suggest further optimization is needed.
- Cost/benefit interpretation: The candidate subset showed some success but with high latency and token usage. Further optimization is needed to reduce these metrics while maintaining or improving success rates.
- Residual pattern: Most tasks failed to achieve success, with only 7 out of 116 tasks completing successfully.
- Residual pattern: The mean latency of 115.06 seconds is significantly high, indicating potential inefficiencies in the current implementation.
- Residual pattern: The mean total tokens of 171,630.55 are also high, suggesting that the model is using more resources than necessary.
- Next hypothesis `H5C` (middle): Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements. Target: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. Verification: Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator.
## Environment boundaries
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- The full-suite candidate uses model qwen2.5-7b-instruct-local, while the promoted paired H5C source used doubao-seed-1-6-250615. This local-GPU campaign evaluates the promoted observation treatment but is not a same-model extension of the paired result.
- ContactsNewContactDraft's official success predicate was fed the upstream UIAutomator state.ui_elements because that observation mode does not populate state.forest; the predicate and requested contact fields were not changed.
- Clipboard get/set retries once after the exact Clipper foreground-access runtime error; the operation, content, task, and evaluator are unchanged.
- SimpleSmsReplyMostRecent polls the unchanged inbox query for up to five additional seconds because emulator-injected SMS delivery can lag past upstream's fixed wait. If the inbox remains empty, the exact last injected address/body is inserted into the same SMS database that upstream clears directly; task data and the evaluator are unchanged.
- The pinned official Retro Music APK omits the playing_queue table, a known upstream runtime error. Only that exact missing-table condition was mapped to an empty observed queue so the unchanged exact queue predicate records an evaluator failure instead of losing the episode.
- Runtime-error retries reuse the exact task parameters retained in the discarded error checkpoints; upstream parameter-generator drift cannot silently change the retried task.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes 8,192 characters only from the middle of the current-screen indexed UI section. Prompt prefix, goal, history, leading and trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
- If compact UIAutomator still exceeds the pinned model's native 32,768-token context, the retry removes a bounded middle span only from indexed UI descriptions: at most 12,000 retained characters for action selection or 6,000 for each before/after summary screen. Prompt prefix, goal, history, action, reason, leading/trailing UI elements and indices, guidance, and output format remain; per-episode removal counters are retained.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
@@ -0,0 +1,136 @@
{
"schema_version": 1,
"experiment": "7-12",
"status": "complete",
"generated_at_utc": "2026-08-04T09:35:34.274750+00:00",
"run_dir": "chapter7/android-world/validation",
"git_commit": "0d0df3e758872c7e73a998e57d2e529b16ece095",
"command": "Five isolated Pixel 6/API-33 trial shards used run_controlled_experiment.py with local-vLLM credentials supplied only by environment-variable name; merge_candidate_shards.py strictly merged trials 1-5.",
"provider_receipt_count": 0,
"status_reasons": [
"580/580 direct episodes completed across 116 tasks x five trials with zero runtime errors; evaluator failures are retained.",
"Every shard completed official Pixel 6/API-33 setup and recorded the same 24/24 required package versions.",
"The result is negative: 26/580 strict successes and deployment is not approved.",
"The candidate used local Qwen2.5-7B while the paired H5C source used Doubao, so no same-model uplift or noninferiority is claimed.",
"All disclosed evaluator, emulator-race, context-limit, resume, and exact-parameter retry compatibility treatments remain in retained evidence."
],
"inputs": [
{
"path": "chapter7/android-world/experiment_core.py",
"bytes": 24201,
"sha256": "752c1f181f0bef3cb0726fcc35eddc8c715624c4804c39b9a2d371518f4b3ac6"
},
{
"path": "chapter7/android-world/run_controlled_experiment.py",
"bytes": 76108,
"sha256": "ebc5417066fadbcd3de15d878d8b95a6ab9bff7f1b8e9eec392d148dbc59817e"
},
{
"path": "chapter7/android-world/merge_candidate_shards.py",
"bytes": 11368,
"sha256": "c7b0519c2246e88750f91859603fd34fea72a67f04c1e8eacb8fc5d2d63cd0f1"
},
{
"path": "chapter7/android-world/t3a_summary.md",
"bytes": 28800,
"sha256": "e5616507e2b0658b64dc5d0e2c08e6207eba12da8345a31d1cb85bf68957f002"
},
{
"path": "chapter7/android-world/t3a_failed_analysis.md",
"bytes": 14448,
"sha256": "7486112a3bf0791a17ca72b81b3745dcd92936416b5c3da937618c8d7a6eb00d"
}
],
"artifacts": [
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json",
"bytes": 3932422,
"sha256": "526e73e061bbafa3802484934dd29512a3b94a6a066d0efa10fc44c3f8138083"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_qwen_20260804/report.md",
"bytes": 70854,
"sha256": "4b4222ba4597eb57f35f40a30de180284d00a8fa66824e5492c75c4c8ce8b96a"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard1/evidence.json",
"bytes": 833946,
"sha256": "8c9bad94ee2141c4de2e666e291f41e1f7d7cf7c53a8ad77a73ab17341bb68a3"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard1/report.md",
"bytes": 23232,
"sha256": "6113f7d935c4c1f085a53a4a2f0be5e9d8053e47c5aa4434d56a6a5d7c56e6a4"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard2/evidence.json",
"bytes": 805331,
"sha256": "2cf319e6eb7f99758c7dd093a46d445afa29ff35a919afc5f08f01c0a2085f44"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard2/report.md",
"bytes": 22455,
"sha256": "573f6fa57a846d6b231db2702069afff277434f39518deec8b64bdc0720eac92"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard3/evidence.json",
"bytes": 815150,
"sha256": "9b4bfed4165696044768c9e05fa00d696469ea99a06e9d3f694ea817a1041911"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard3/report.md",
"bytes": 22157,
"sha256": "8e37b4fdaf97eaa0bc14a6a161bcf14914ae6512782a94b5901c7bbfc361806a"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard4/evidence.json",
"bytes": 820872,
"sha256": "4da385617b9b2f7a3d684884fd79184970e23d862583ab6b622b48f975e15696"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard4/report.md",
"bytes": 21949,
"sha256": "509a9224d6ce823c6d9ccff15a962fafdc0ed0b9724c053c2d83754517764f02"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard5/evidence.json",
"bytes": 808670,
"sha256": "245f99f7a8924d1b8340e510a6b12f345ac35ea779099605755d15d180109ce3"
},
{
"path": "chapter7/android-world/validation/candidate_h5c_api33_local_shard5/report.md",
"bytes": 21696,
"sha256": "4d22f0cdeb7b5f82560241b2f91169b94629c1f9cca2f320de4f54f5f1b99479"
},
{
"path": "chapter7/android-world/validation/paired_h5_a11y_api35_20260729/evidence.json",
"bytes": 47072,
"sha256": "b2276e35e1f58d32775ea70ac04142a7d3075c074b9f31d40b153e1be10f93da"
},
{
"path": "chapter7/android-world/validation/paired_h5_a11y_api35_20260729/report.md",
"bytes": 8239,
"sha256": "5d5b0db9580020c286096ca53617cbd2d0f7fb811e48c4cf51688212ef638272"
},
{
"path": "chapter7/android-world/validation/paired_h5c_compact_api35_20260729/evidence.json",
"bytes": 39643,
"sha256": "a60204b6c0ac187b0ae9a68182b5ada302420fc31b2fc22c308e10eb1242e390"
},
{
"path": "chapter7/android-world/validation/paired_h5c_compact_api35_20260729/report.md",
"bytes": 8930,
"sha256": "55abde0ef9a6d5ed090f26bc99e424346c42a5e98036a278bd50edcf4bf48f2e"
},
{
"path": "chapter7/android-world/validation/paired_wifi_api35_20260729/evidence.json",
"bytes": 45941,
"sha256": "a1665bffafed7b8935db03dc4369435015f784d6f6465fe09c7da86fc0504635"
},
{
"path": "chapter7/android-world/validation/paired_wifi_api35_20260729/report.md",
"bytes": 4096,
"sha256": "067f176e4722082afd45bedbd79a392b26d677d54013a8b14a81cea1fbd70d57"
}
]
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,93 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260729T122904Z`
- Generated (UTC): `2026-07-29T13:17:47Z`
- Upstream commit: `0e95d641e244504c22087cc29b013f3b2428a261`
- Device: `sdk_gphone64_arm64`, API `35` (upstream tested reference: API `33`)
- Observation method: `varies_by_arm:a11y_forwarder_app_vs_uiautomator`
- Provider/model: `ark` / `doubao-seed-1-6-250615`
- Scope: 4 task(s), 1 trial(s), mode `paired`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5`
- Change: Select AndroidWorld's UIAUTOMATOR observation method in the companion runner without changing upstream source.
- Expected measurable result: At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, no regression, and at most 1.5x mean latency and tokens; never treat a subset gain as full-suite success or deployment approval.
## 3. Controlled experiment
- Phase: `phase_2_middle` — middle-layer input-pipeline ablation prompted by phase-1 residual traces
- Independent variable: accessibility observation pipeline (gRPC forwarder versus UIAutomator)
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| control | 4/4 | 0.250 | 0.500 | 8.000 | 172.849 | 14.000 | 68310.500 | 251587 / 21655 |
| treatment | 4/4 | 1.000 | 1.000 | 5.750 | 136.235 | 11.000 | 170673.500 | 671404 / 11290 |
| Task / trial | Control | Treatment | Δ success | Control→treatment steps |
| --- | ---: | ---: | ---: | ---: |
| SystemWifiTurnOff / 1 | 0 | 1 | +1 | 10→6 |
| SystemWifiTurnOffVerify / 1 | 0 | 1 | +1 | 10→6 |
| SystemWifiTurnOn / 1 | 0 | 1 | +1 | 10→7 |
| SystemWifiTurnOnVerify / 1 | 1 | 1 | +0 | 2→4 |
## 4. Data-driven decision
- Outcome: **`restrict_candidate_due_to_cost`**
- Reason: Treatment improved paired success without regressions, but exceeded the latency/token guardrails (1.50x / 1.50x). Restrict it to targeted follow-up; do not promote it to the full suite yet.
- Treatment/control mean latency ratio: 0.788
- Treatment/control mean token ratio: 2.498
- Treatment/control mean LLM-call ratio: 0.786
- Cost guardrails passed: **false**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `control / SystemWifiTurnOff / trial 1`: evaluator reward / completion gate was not satisfied
- `control / SystemWifiTurnOffVerify / trial 1`: final evaluator state passed, but the agent never declared completion
- `control / SystemWifiTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: Experiment 7-12 compared control (a11y-forwarder observation) and treatment (UIAutomator observation) arms in a paired setup with 4 Wi-Fi system Settings tasks. Treatment achieved 100% success (4/4) vs control's 25% (1/4), reduced mean latency (136.2s vs 172.8s) and LLM calls (11.0 vs 14.0), but had a mean token ratio (treatment/control) of 2.498, exceeding the 1.5x guardrail. Environment boundaries include API 35 AVD (vs upstream API 33 reference), restriction to Settings tasks, UIAutomator as a compatibility path (not reference config), and skipped device-time setting due to non-root AVD limitations.
- Cost/benefit interpretation: Treatment provides substantial benefit via improved success rate (net +3) and reduced latency/LLM calls, but incurs significantly higher token cost (2.498x control), violating token guardrails and limiting deployment despite success gains.
- Residual pattern: Treatment mean token ratio (2.498x) exceeds 1.5x guardrail
- Residual pattern: Control arm has low success rate (25%, 1/4 completed episodes)
- Next hypothesis `H6` (middle): Optimize UIAutomator observation pipeline to reduce token usage while maintaining treatment success rate Target: Mean token ratio (treatment/control) ≤1.5x and success rate ≥1.0 in paired Wi-Fi tasks Verification: Conduct paired run with optimized UIAutomator pipeline vs control, using same 4 Wi-Fi tasks, API 35 AVD environment, and guardrails; measure token ratio and success rate
## Environment boundaries
- The available AVD is API 35, while upstream is tested on Pixel 6 / API 33. Results are real but not reference-environment comparable.
- The full third-party AndroidWorld app bundle was not provisioned. This run is restricted to system Settings tasks.
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Per-task device-time setting was skipped because the non-root API-35 AVD rejects `adb shell date`; Wi-Fi evaluators do not depend on time.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
@@ -0,0 +1,940 @@
{
"arm_summary": {
"control": {
"completed_episodes": 4,
"episodes": 4,
"error_episodes": 0,
"mean_evaluator_reward": 1.0,
"mean_latency_s": 101.198336,
"mean_llm_calls": 8.5,
"mean_llm_latency_s": 62.453747,
"mean_steps": 4.75,
"mean_total_tokens": 139439.5,
"success_rate": 1.0,
"successes": 4,
"total_input_tokens": 549928,
"total_output_tokens": 7830,
"total_tokens": 557758
},
"treatment": {
"completed_episodes": 4,
"episodes": 4,
"error_episodes": 0,
"mean_evaluator_reward": 1.0,
"mean_latency_s": 99.184035,
"mean_llm_calls": 8.5,
"mean_llm_latency_s": 59.748216,
"mean_steps": 4.75,
"mean_total_tokens": 70557.5,
"success_rate": 1.0,
"successes": 4,
"total_input_tokens": 274067,
"total_output_tokens": 8163,
"total_tokens": 282230
}
},
"baseline": {
"agent": "t3a_claude4_sonnet",
"provenance": "historical bundled report; not generated by this runner",
"reported_success_rate_approx": 0.88,
"run_date": "2025-07-02",
"source": "t3a_summary.md and t3a_failed_analysis.md",
"tasks": 116,
"trials_per_task": 1
},
"command": [
"run_controlled_experiment.py",
"--mode",
"paired",
"--hypothesis",
"H5C",
"--source-phase1-evidence",
"validation/paired_wifi_api35_20260729/evidence.json",
"--source-phase2-evidence",
"validation/paired_h5_a11y_api35_20260729/evidence.json",
"--tasks",
"SystemWifiTurnOff,SystemWifiTurnOffVerify,SystemWifiTurnOn,SystemWifiTurnOnVerify",
"--trials",
"1",
"--seed",
"42",
"--model-seed",
"42",
"--max-steps",
"10",
"--transition-pause",
"0.5",
"--skip-device-time",
"--output-dir",
"validation/paired_h5c_compact_api35_20260729"
],
"credentials_persisted": false,
"decision": {
"completed_pairs": 4,
"deployment_approved": false,
"guardrails": {
"maximum_latency_ratio": 1.5,
"maximum_token_ratio": 0.75,
"objective": "success_noninferiority_and_token_reduction",
"passed": true,
"require_all_treatment_pairs_successful": true
},
"mean_latency_ratio_treatment_over_control": 0.980096,
"mean_llm_call_ratio_treatment_over_control": 1.0,
"mean_token_ratio_treatment_over_control": 0.506008,
"net_success_delta": 0,
"outcome": "promote_efficient_candidate_to_full_suite_rerun",
"paired_regressions": 0,
"promote_to_full_suite_candidate": true,
"reason": "Treatment preserved paired success with no regression and passed the latency/token efficiency guardrails. This is a candidate decision only.",
"required_treatment_successes": 4,
"scope_recommendation": "full_suite_candidate_only",
"success_preservation_passed": true,
"treatment_successes": 4
},
"diagnosis": {
"findings": [
"The historical run evaluated 116 tasks once each and reports approximately 88% overall success.",
"Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.",
"The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.",
"The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause."
],
"layered_hypotheses": [
{
"id": "H1",
"idea": "Add Wi-Fi Settings navigation and final-state verification guidance.",
"layer": "surface",
"status": "tested in source phase 1",
"target": "At least one net paired success across the four Wi-Fi tasks, with no regression.",
"verification": "Paired upstream-prompt versus task-guideline ablation with matched seeds."
},
{
"id": "H2",
"idea": "Add application-specific recognition rules for the non-standard Tasks UI.",
"layer": "surface",
"status": "not tested",
"target": "Improve at least two of the six historical Tasks failures with no regression.",
"verification": "Paired Tasks-only prompt/tool-description ablation after app provisioning."
},
{
"id": "H3",
"idea": "Repair and validate the multimodal input path for transcription tasks.",
"layer": "middle",
"status": "not tested",
"target": "Raise transcription success above the historical 0% while bounding added tokens and latency.",
"verification": "Paired screenshot-disabled versus screenshot-enabled transcription run."
},
{
"id": "H4",
"idea": "Conditionally enable deeper thinking for counting tasks.",
"layer": "middle",
"status": "not tested",
"target": "Improve math/counting success without applying the cost to unrelated tasks.",
"verification": "Paired tag-routed thinking-mode ablation with latency and token guardrails."
},
{
"id": "H5",
"idea": "Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path.",
"layer": "middle",
"status": "tested in source phase 2",
"target": "At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens.",
"verification": "Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds."
},
{
"id": "H5C",
"idea": "Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost.",
"layer": "middle",
"status": "tested in this run",
"target": "Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.",
"verification": "Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator."
},
{
"id": "H6",
"idea": "Combine screenshots with the structured UI tree and compare stronger vision-capable models.",
"layer": "deep",
"status": "not tested",
"target": "Improve complex-UI success enough to justify multimodal latency and token cost.",
"verification": "Factorial UI-tree/screenshot/model ablation on the full tagged slice."
}
]
},
"environment": {
"a11y_method": "varies_by_arm:uiautomator_vs_uiautomator_compact",
"android_world_checkout": "/Users/boj/book/ai-agent-book/chapter7/android_world",
"android_world_checkout_clean": true,
"android_world_commit": "0e95d641e244504c22087cc29b013f3b2428a261",
"api_level": 35,
"avd_name": "Pixel_9_Pro_API_35",
"device_model": "sdk_gphone64_arm64",
"device_serial": "emulator-5554",
"grpc_port": 8554,
"perform_emulator_setup": false,
"physical_size": "Physical size: 1280x2856",
"protobuf_bootstrap": {
"generated_protobufs_in_temporary_directory": true,
"temporary_directory_persisted": false
},
"skip_device_time": true,
"upstream_tested_api_level": 33,
"upstream_tested_device": "Pixel 6"
},
"environment_boundaries": [
"The available AVD is API 35, while upstream is tested on Pixel 6 / API 33. Results are real but not reference-environment comparable.",
"The full third-party AndroidWorld app bundle was not provisioned. This run is restricted to system Settings tasks.",
"UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.",
"Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.",
"Per-task device-time setting was skipped because the non-root API-35 AVD rejects `adb shell date`; Wi-Fi evaluators do not depend on time."
],
"episodes": [
{
"agent_declared_done": true,
"arm": "control",
"elapsed_s": 105.382271,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi off.",
"initial_evaluator_reward": 0.0,
"llm": {
"calls": 9,
"input_tokens": 132349,
"latency_s": 64.872256,
"output_tokens": 2008,
"reasoning_tokens": 1488,
"system_fingerprints": []
},
"observation_method": "uiautomator",
"order_position": 1,
"pair_id": "SystemWifiTurnOff:trial-1:seed-42",
"pair_seed": 42,
"params": {
"on_or_off": "off"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOff",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 22.742624,
"goal_status": null,
"reason": "To turn off Wi-Fi, the first step is to open the Settings app. On the current home screen, UI element 26 is the \"Settings\" app, which is clickable. Clicking it will open the Settings menu where Wi-Fi settings can be accessed.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 21.091666,
"goal_status": null,
"reason": "The current screen shows an error dialog \"com.google.androidenv.accessibilityforwarder keeps stopping\" which is blocking access to the Settings menu. To proceed with turning off Wi-Fi, I need to dismiss this dialog by closing the problematic app. The \"Close app\" button (index 11) is visible and clickable, so clicking it will remove the dialog and allow me to continue navigating to Wi-Fi settings.",
"step": 2
},
{
"action_type": "input_text",
"done": false,
"elapsed_s": 24.93577,
"goal_status": null,
"reason": "To proceed with turning off Wi-Fi, I need to locate the Wi-Fi settings. Since the current screen is the settings search screen, the most efficient way is to search for \"Wi-Fi\" using the search bar. The search bar (UI element 15) is an editable text field and is focused, making it suitable for input.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 25.397996,
"goal_status": null,
"reason": "The current screen shows the Wi-Fi settings with a checked (on) switch at index 32. To turn off Wi-Fi, I need to click this switch to toggle it to the off position.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 8.938998,
"goal_status": "complete",
"reason": "The Wi-Fi switch (index 32) is now unchecked (is_checked=False), indicating that Wi-Fi has been turned off. The user's goal of turning Wi-Fi off has been successfully achieved.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "treatment",
"elapsed_s": 103.549168,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi off.",
"initial_evaluator_reward": 0.0,
"llm": {
"calls": 9,
"input_tokens": 75681,
"latency_s": 65.635331,
"output_tokens": 2101,
"reasoning_tokens": 1568,
"system_fingerprints": []
},
"observation_method": "uiautomator_compact",
"order_position": 2,
"pair_id": "SystemWifiTurnOff:trial-1:seed-42",
"pair_seed": 42,
"params": {
"on_or_off": "off"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOff",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 20.105315,
"goal_status": null,
"reason": "To turn off Wi-Fi, the first step is to open the Settings app. On the current home screen, UI element 12 is the \"Settings\" app, which is visible and clickable. Clicking it will open the Settings menu where Wi-Fi settings can be accessed.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 23.230104,
"goal_status": null,
"reason": "To turn off Wi-Fi, I need to access the Wi-Fi settings. The \"Network & internet\" option (UI element 6) is visible and clickable, and its summary mentions \"WiFi\", indicating it contains Wi-Fi settings. Clicking this will lead to the next screen where Wi-Fi can be toggled off.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 21.825863,
"goal_status": null,
"reason": "To access the Wi-Fi settings and turn it off, I need to click on the \"Internet\" option, which is part of the clickable linear layout (UI element 3) on the Network & internet settings page.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 27.327625,
"goal_status": null,
"reason": "The current screen shows the Internet settings where the Wi-Fi switch is visible. The Wi-Fi option is contained in UI element 9 (a LinearLayout), which is clickable. This element includes the \"Wi-Fi\" text (UI element 10) and the switch (UI element 11) that is currently on (is_checked=True). Clicking this LinearLayout will toggle the Wi-Fi switch off.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 8.69858,
"goal_status": "complete",
"reason": "The Wi-Fi switch is now off (is_checked=False as per UI element 11), so the task of turning Wi-Fi off has been successfully completed.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "treatment",
"elapsed_s": 102.209211,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi off.",
"initial_evaluator_reward": 1.0,
"llm": {
"calls": 9,
"input_tokens": 64462,
"latency_s": 62.502274,
"output_tokens": 2130,
"reasoning_tokens": 1620,
"system_fingerprints": []
},
"observation_method": "uiautomator_compact",
"order_position": 1,
"pair_id": "SystemWifiTurnOffVerify:trial-1:seed-1051",
"pair_seed": 1051,
"params": {
"on_or_off": "off"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOffVerify",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 20.831033,
"goal_status": null,
"reason": "The current screen shows an error message pop-up that is blocking access to the main interface. To proceed with turning off Wi-Fi, I need to close this pop-up first. The \"Close app\" button (UI element 2) is visible and clickable, so clicking it will dismiss the pop-up.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 23.270511,
"goal_status": null,
"reason": "To turn off Wi-Fi, the next step is to open the Settings app, as Wi-Fi settings are typically located there. The \"Settings\" app is visible on the home screen (UI element 12), which is clickable. Clicking it will allow access to the settings menu where Wi-Fi can be managed.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 23.028096,
"goal_status": null,
"reason": "To turn off Wi-Fi, I need to access the network settings. The \"Network & internet\" option (UI element 6, which is a clickable LinearLayout containing the \"Network & internet\" title and summary) is the appropriate next step to navigate to Wi-Fi controls.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 20.913882,
"goal_status": null,
"reason": "To access the Wi-Fi controls, I need to navigate to the \"Internet\" settings. The \"Internet\" option is contained within the clickable LinearLayout (UI element 3), which is visible and interactive on the current \"Network & internet\" screen. Clicking this element will likely lead to the page where Wi-Fi can be toggled off.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 11.76325,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 11) is already in the off position (is_checked=False), so the goal of turning Wi-Fi off has been achieved.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "control",
"elapsed_s": 112.724126,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi off.",
"initial_evaluator_reward": 1.0,
"llm": {
"calls": 9,
"input_tokens": 150834,
"latency_s": 67.904524,
"output_tokens": 2205,
"reasoning_tokens": 1649,
"system_fingerprints": []
},
"observation_method": "uiautomator",
"order_position": 2,
"pair_id": "SystemWifiTurnOffVerify:trial-1:seed-1051",
"pair_seed": 1051,
"params": {
"on_or_off": "off"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOffVerify",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 18.772365,
"goal_status": null,
"reason": "To turn off Wi-Fi, the first step is to open the Settings app. On the home screen, UI element 26 is the \"Settings\" app, which is clickable. Clicking it will open the Settings menu where Wi-Fi settings can be accessed.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 25.71773,
"goal_status": null,
"reason": "To turn off Wi-Fi, I need to access the Wi-Fi settings. The \"Network & internet\" option (UI element 24) in the Settings menu is the appropriate entry point as it mentions \"WiFi\" in its summary. The parent LinearLayout (UI element 20) is clickable, so I will tap on it to proceed to the Network & internet settings page.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 24.548097,
"goal_status": null,
"reason": "The goal is to turn off Wi-Fi. Currently, on the Settings home screen, the \"Network & internet\" section (UI element 20) is visible, which is the correct path to access Wi-Fi settings. Since the previous attempt to click this element didn't result in a screen change (possibly due to a delay), retrying the click on UI element 20 is necessary to navigate to the Network & internet settings where Wi-Fi can be toggled off.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 26.737918,
"goal_status": null,
"reason": "To access the Wi-Fi settings, I need to navigate into the \"Internet\" section under \"Network & internet\". The \"Internet\" option (UI element 15) is clickable and likely contains Wi-Fi controls, so clicking it will allow me to proceed toward turning off Wi-Fi.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 13.895503,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 31) is already in the off position (is_checked=False), so the task of turning Wi-Fi off has been completed.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "control",
"elapsed_s": 100.007401,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi on.",
"initial_evaluator_reward": 0.0,
"llm": {
"calls": 9,
"input_tokens": 146625,
"latency_s": 60.454149,
"output_tokens": 1997,
"reasoning_tokens": 1484,
"system_fingerprints": []
},
"observation_method": "uiautomator",
"order_position": 1,
"pair_id": "SystemWifiTurnOn:trial-1:seed-2060",
"pair_seed": 2060,
"params": {
"on_or_off": "on"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOn",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 21.700473,
"goal_status": null,
"reason": "To turn on Wi-Fi, the first step is to open the Settings app. On the current home screen, UI element 26 is the \"Settings\" app, which is clickable. Clicking it will allow access to the settings menu where Wi-Fi can be enabled.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 20.450628,
"goal_status": null,
"reason": "To proceed towards enabling Wi-Fi, I need to access the \"Network & internet\" settings, which is the logical next step as it contains Wi-Fi options. The \"Network & internet\" element (index 24) is visible and clickable on the current screen.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 22.812015,
"goal_status": null,
"reason": "To proceed towards enabling Wi-Fi, I need to access the Internet settings which typically include Wi-Fi options. The \"Internet\" option (contained in the clickable LinearLayout at index 15) is visible on the screen and is the logical next step to find the Wi-Fi toggle.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 21.686302,
"goal_status": null,
"reason": "The current screen shows the Internet settings page where the Wi-Fi option is present. The Wi-Fi switch (UI element 31) is currently unchecked (is_checked=False), indicating Wi-Fi is off. To turn Wi-Fi on, I need to click this switch to toggle it to the on position.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 10.361169,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 35) is now checked, indicating Wi-Fi is enabled. The screen also shows \"AndroidWifi\" as connected, confirming the task of turning Wi-Fi on has been successfully completed.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "treatment",
"elapsed_s": 106.717192,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi on.",
"initial_evaluator_reward": 0.0,
"llm": {
"calls": 9,
"input_tokens": 74268,
"latency_s": 60.267918,
"output_tokens": 2150,
"reasoning_tokens": 1681,
"system_fingerprints": []
},
"observation_method": "uiautomator_compact",
"order_position": 2,
"pair_id": "SystemWifiTurnOn:trial-1:seed-2060",
"pair_seed": 2060,
"params": {
"on_or_off": "on"
},
"status": "completed",
"steps": 5,
"success": true,
"task": "SystemWifiTurnOn",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "open_app",
"done": false,
"elapsed_s": 23.770885,
"goal_status": null,
"reason": "To turn on Wi-Fi, the first step is to open the Settings app where network settings are located. The current screen shows a \"Settings\" app icon (UI element 12), and using the `open_app` action is the most direct way to launch the app.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 27.515651,
"goal_status": null,
"reason": "To proceed to the Wi-Fi settings, I need to access the \"Network & internet\" section. The UI element 6 is a clickable LinearLayout that contains the \"Network & internet\" text (elements 7 and 8), making it the appropriate element to interact with.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 22.324583,
"goal_status": null,
"reason": "To access the Wi-Fi settings, I need to navigate into the \"Internet\" section, which is likely to contain Wi-Fi controls. The \"Internet\" option is part of the clickable LinearLayout (UI element 3) on the current screen.",
"step": 3
},
{
"action_type": "click",
"done": false,
"elapsed_s": 19.494808,
"goal_status": null,
"reason": "The Wi-Fi switch (UI element 11) is currently off (is_checked=False). To turn Wi-Fi on, I need to click this switch to toggle it to the on position.",
"step": 4
},
{
"action_type": "status",
"done": true,
"elapsed_s": 11.27913,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 11) is now checked (is_checked=True), indicating Wi-Fi is turned on. Additionally, \"AndroidWifi\" is shown as connected (UI element 12), confirming the task is completed.",
"step": 5
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "treatment",
"elapsed_s": 84.26057,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi on.",
"initial_evaluator_reward": 1.0,
"llm": {
"calls": 7,
"input_tokens": 59656,
"latency_s": 50.587342,
"output_tokens": 1782,
"reasoning_tokens": 1390,
"system_fingerprints": []
},
"observation_method": "uiautomator_compact",
"order_position": 1,
"pair_id": "SystemWifiTurnOnVerify:trial-1:seed-3069",
"pair_seed": 3069,
"params": {
"on_or_off": "on"
},
"status": "completed",
"steps": 4,
"success": true,
"task": "SystemWifiTurnOnVerify",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 20.295946,
"goal_status": null,
"reason": "To turn on Wi-Fi, the first step is to open the Settings app. On the current home screen, UI element 12 is the \"Settings\" app, which is visible and clickable. Clicking on it will allow access to the settings menu where Wi-Fi can be enabled.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 26.66904,
"goal_status": null,
"reason": "To proceed with turning on Wi-Fi, I need to access the \"Network & internet\" settings, as it contains Wi-Fi options. The \"Network & internet\" section is represented by the clickable LinearLayout (UI element 6) on the current screen. Clicking this will navigate to the relevant submenu where Wi-Fi can be enabled.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 26.764651,
"goal_status": null,
"reason": "To proceed with turning on Wi-Fi, I need to access the Wi-Fi settings. The \"Internet\" section (contained in UI element 3, a clickable LinearLayout) is the logical next step, as it likely leads to Wi-Fi configuration options.",
"step": 3
},
{
"action_type": "status",
"done": true,
"elapsed_s": 8.154307,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 11) is already checked (is_checked=True), indicating that Wi-Fi is enabled. The user's goal of turning Wi-Fi on has been achieved.",
"step": 4
}
],
"trial": 1
},
{
"agent_declared_done": true,
"arm": "control",
"elapsed_s": 86.679544,
"error": null,
"evaluator_reward": 1.0,
"goal": "Turn wifi on.",
"initial_evaluator_reward": 1.0,
"llm": {
"calls": 7,
"input_tokens": 120120,
"latency_s": 56.58406,
"output_tokens": 1620,
"reasoning_tokens": 1195,
"system_fingerprints": []
},
"observation_method": "uiautomator",
"order_position": 2,
"pair_id": "SystemWifiTurnOnVerify:trial-1:seed-3069",
"pair_seed": 3069,
"params": {
"on_or_off": "on"
},
"status": "completed",
"steps": 4,
"success": true,
"task": "SystemWifiTurnOnVerify",
"task_guidelines_applied": false,
"trace": [
{
"action_type": "click",
"done": false,
"elapsed_s": 21.802337,
"goal_status": null,
"reason": "To turn on Wi-Fi, the first step is to open the Settings app. On the home screen, UI element 26 is the \"Settings\" app, which is clickable. Clicking it will allow access to the settings menu where Wi-Fi can be enabled.",
"step": 1
},
{
"action_type": "click",
"done": false,
"elapsed_s": 24.619676,
"goal_status": null,
"reason": "To proceed with enabling Wi-Fi, I need to access the \"Network & internet\" settings, as it contains Wi-Fi options. The \"Network & internet\" section is a clickable element (UI element 20) on the current Settings screen. Clicking this will navigate to the relevant submenu where Wi-Fi can be toggled on.",
"step": 2
},
{
"action_type": "click",
"done": false,
"elapsed_s": 26.561039,
"goal_status": null,
"reason": "To proceed with enabling Wi-Fi, I need to access the Wi-Fi settings. The \"Internet\" option (UI element 19) with \"AndroidWifi\" as the summary (UI element 20) is likely the entry point to Wi-Fi settings. The parent linear layout (UI element 15) is clickable, so clicking it will navigate to the Wi-Fi settings screen where the toggle can be accessed.",
"step": 3
},
{
"action_type": "status",
"done": true,
"elapsed_s": 11.348688,
"goal_status": "complete",
"reason": "The Wi-Fi switch (UI element 35) is already in the checked state, indicating that Wi-Fi is turned on. The user's goal of turning Wi-Fi on has been achieved.",
"step": 4
}
],
"trial": 1
}
],
"experiment": "7-12",
"experiment_complete": false,
"generated_at_utc": "2026-07-29T14:12:27Z",
"hypothesis": {
"change": "Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.",
"expected_result": "Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.",
"guardrails": "Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.",
"guideline_sha256": null,
"guidelines": [],
"id": "H5C",
"layer": "middle",
"verification": "Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator."
},
"llm_analysis": {
"cost_benefit_interpretation": "The treatment provides significant token efficiency (mean token ratio 0.506) with preserved success and marginal latency improvement (mean latency ratio 0.980) in the tested 4-task subset. However, interpretation is bounded by the API/app environment: results are from an API 35 AVD (not upstream API 33 reference), restricted to system Settings tasks (no full third-party app bundle), and use UIAutomator as a compatibility path (not upstream reference configuration), limiting generalizability beyond the tested scope.",
"llm": {
"calls": 1,
"input_tokens": 3257,
"latency_s": 34.44844,
"output_tokens": 1101,
"reasoning_tokens": 647,
"system_fingerprints": []
},
"next_hypothesis": {
"id": "H5C-full",
"idea": "Evaluate compact UIAutomator (semantic-filtered elements) across the full AndroidWorld task suite to verify token efficiency and success preservation beyond the 4-task Settings subset.",
"layer": "middle",
"target": "Full task suite (all 116 tasks) with upstream reference environment (Pixel 6 / API 33) and provisioned full third-party AndroidWorld app bundle.",
"verification": "Paired run comparing compact UIAutomator (treatment) vs raw UIAutomator (control) across all tasks, ensuring guardrails (mean token ratio ≤0.75, mean latency ratio ≤1.5, success non-inferiority) hold in the reference environment."
},
"observed_failure_pattern": [],
"source": "real configured LLM over aggregate direct evidence",
"status": "completed",
"summary": "A paired experiment comparing raw UIAutomator (control) and compact UIAutomator (treatment) on 4 system Settings tasks (SystemWifiTurnOff, SystemWifiTurnOffVerify, SystemWifiTurnOn, SystemWifiTurnOnVerify) in an API 35 AVD environment. Both arms completed 4 episodes with 100% success rate, identical mean steps (4.75) and LLM calls (8.5). Treatment showed lower mean total tokens (50.6% of control) and slightly lower mean latency (98.0% of control). The decision was to promote the treatment as a full-suite candidate rerun, as it preserved success, passed latency/token guardrails, but remains a subset with environment limitations."
},
"model": {
"base_url": "https://ark.cn-beijing.volces.com/api/v3",
"max_tokens": 1024,
"model": "doubao-seed-1-6-250615",
"provider": "ark",
"seed": 42,
"temperature": 0
},
"paired_comparison": [
{
"control_latency_s": 105.382271,
"control_reward": 1.0,
"control_steps": 5,
"control_success": true,
"pair_id": "SystemWifiTurnOff:trial-1:seed-42",
"reward_delta": 0.0,
"success_delta": 0,
"task": "SystemWifiTurnOff",
"treatment_latency_s": 103.549168,
"treatment_reward": 1.0,
"treatment_steps": 5,
"treatment_success": true,
"trial": 1
},
{
"control_latency_s": 112.724126,
"control_reward": 1.0,
"control_steps": 5,
"control_success": true,
"pair_id": "SystemWifiTurnOffVerify:trial-1:seed-1051",
"reward_delta": 0.0,
"success_delta": 0,
"task": "SystemWifiTurnOffVerify",
"treatment_latency_s": 102.209211,
"treatment_reward": 1.0,
"treatment_steps": 5,
"treatment_success": true,
"trial": 1
},
{
"control_latency_s": 100.007401,
"control_reward": 1.0,
"control_steps": 5,
"control_success": true,
"pair_id": "SystemWifiTurnOn:trial-1:seed-2060",
"reward_delta": 0.0,
"success_delta": 0,
"task": "SystemWifiTurnOn",
"treatment_latency_s": 106.717192,
"treatment_reward": 1.0,
"treatment_steps": 5,
"treatment_success": true,
"trial": 1
},
{
"control_latency_s": 86.679544,
"control_reward": 1.0,
"control_steps": 4,
"control_success": true,
"pair_id": "SystemWifiTurnOnVerify:trial-1:seed-3069",
"reward_delta": 0.0,
"success_delta": 0,
"task": "SystemWifiTurnOnVerify",
"treatment_latency_s": 84.26057,
"treatment_reward": 1.0,
"treatment_steps": 4,
"treatment_success": true,
"trial": 1
}
],
"phase": {
"description": "middle-layer input-pipeline cost refinement after the H5 success/cost result",
"id": "phase_2_cost_refinement",
"independent_variable": "raw versus semantic-filtered UIAutomator element list",
"source_phase1_decision": {
"completed_pairs": 4,
"mean_latency_ratio_treatment_over_control": 0.672411,
"net_success_delta": 0,
"outcome": "inconclusive_no_success_gain",
"paired_regressions": 0,
"promote_to_full_suite_candidate": false,
"reason": "Treatment produced no paired success gain; keep the upstream control prompt."
},
"source_phase1_evidence": "/Users/boj/book/ai-agent-book/chapter7/android-world/validation/paired_wifi_api35_20260729/evidence.json",
"source_phase1_run_id": "exp7-12-20260729T114648Z",
"source_phase2_decision": {
"completed_pairs": 4,
"deployment_approved": false,
"guardrails": {
"maximum_latency_ratio": 1.5,
"maximum_token_ratio": 1.5,
"passed": false
},
"mean_latency_ratio_treatment_over_control": 0.788177,
"mean_llm_call_ratio_treatment_over_control": 0.785714,
"mean_token_ratio_treatment_over_control": 2.498496,
"net_success_delta": 3,
"outcome": "restrict_candidate_due_to_cost",
"paired_regressions": 0,
"promote_to_full_suite_candidate": false,
"reason": "Treatment improved paired success without regressions, but exceeded the latency/token guardrails (1.50x / 1.50x). Restrict it to targeted follow-up; do not promote it to the full suite yet.",
"scope_recommendation": "do_not_deploy"
},
"source_phase2_evidence": "/Users/boj/book/ai-agent-book/chapter7/android-world/validation/paired_h5_a11y_api35_20260729/evidence.json",
"source_phase2_run_id": "exp7-12-20260729T122904Z"
},
"resume_commands": [
[
"run_controlled_experiment.py",
"--mode",
"paired",
"--hypothesis",
"H5C",
"--source-phase1-evidence",
"validation/paired_wifi_api35_20260729/evidence.json",
"--source-phase2-evidence",
"validation/paired_h5_a11y_api35_20260729/evidence.json",
"--tasks",
"SystemWifiTurnOff,SystemWifiTurnOffVerify,SystemWifiTurnOn,SystemWifiTurnOnVerify",
"--trials",
"1",
"--seed",
"42",
"--model-seed",
"42",
"--max-steps",
"10",
"--transition-pause",
"0.5",
"--skip-device-time",
"--output-dir",
"validation/paired_h5c_compact_api35_20260729",
"--resume"
]
],
"run_id": "exp7-12-20260729T131911Z",
"schema_version": 1,
"scope": {
"base_pair_seed": 42,
"completed_episodes": 8,
"direct_episode_gate_completed": false,
"error_episodes": 0,
"full_suite_completed": false,
"manuscript_five_seed_gate_completed": false,
"max_steps": 10,
"mode": "paired",
"tasks": [
"SystemWifiTurnOff",
"SystemWifiTurnOffVerify",
"SystemWifiTurnOn",
"SystemWifiTurnOnVerify"
],
"transition_pause_s": 0.5,
"trials_per_task": 1
}
}
@@ -0,0 +1,88 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260729T131911Z`
- Generated (UTC): `2026-07-29T14:12:27Z`
- Upstream commit: `0e95d641e244504c22087cc29b013f3b2428a261`
- Device: `sdk_gphone64_arm64`, API `35` (upstream tested reference: API `33`)
- Observation method: `varies_by_arm:uiautomator_vs_uiautomator_compact`
- Provider/model: `ark` / `doubao-seed-1-6-250615`
- Scope: 4 task(s), 1 trial(s), mode `paired`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
The diagnosis produced explicit surface, middle, and deep hypotheses. Only one variable is changed in this run; the other hypotheses remain untested.
| Layer / ID | Proposed change | Target | Verification | Status |
| --- | --- | --- | --- | --- |
| surface / `H1` | Add Wi-Fi Settings navigation and final-state verification guidance. | At least one net paired success across the four Wi-Fi tasks, with no regression. | Paired upstream-prompt versus task-guideline ablation with matched seeds. | tested in source phase 1 |
| surface / `H2` | Add application-specific recognition rules for the non-standard Tasks UI. | Improve at least two of the six historical Tasks failures with no regression. | Paired Tasks-only prompt/tool-description ablation after app provisioning. | not tested |
| middle / `H3` | Repair and validate the multimodal input path for transcription tasks. | Raise transcription success above the historical 0% while bounding added tokens and latency. | Paired screenshot-disabled versus screenshot-enabled transcription run. | not tested |
| middle / `H4` | Conditionally enable deeper thinking for counting tasks. | Improve math/counting success without applying the cost to unrelated tasks. | Paired tag-routed thinking-mode ablation with latency and token guardrails. | not tested |
| middle / `H5` | Replace the API-35-incompatible gRPC accessibility feed with upstream's UIAutomator observation path. | At least one net paired Wi-Fi success with no regression and at most 1.5x latency/tokens. | Paired a11y-forwarder versus UIAutomator run with the same upstream T3A prompt and matched seeds. | tested in source phase 2 |
| middle / `H5C` | Filter non-semantic UIAutomator container nodes after H5 exposed excessive prompt-token cost. | Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency. | Paired raw-UIAutomator versus compact-UIAutomator run with matched tasks, seeds, prompt, and evaluator. | tested in this run |
| deep / `H6` | Combine screenshots with the structured UI tree and compare stronger vision-capable models. | Improve complex-UI success enough to justify multimodal latency and token cost. | Factorial UI-tree/screenshot/model ablation on the full tagged slice. | not tested |
Selected hypothesis: `H5C`
- Change: Use the real upstream UIAutomator hierarchy but retain only visible text, descriptions, and actionable/scrollable elements.
- Expected measurable result: Preserve H5 paired success with no regression while using at most 0.75x raw-UIAutomator tokens and 1.5x latency.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; require at least four completed pairs, every compact-UIAutomator treatment pair successful, no paired regression, at most 1.5x mean latency, and at most 0.75x raw-UIAutomator mean tokens. Passing a paired gate permits only a full-suite candidate rerun; it is not deployment approval, and a subset must never be reported as full-suite success.
## 3. Controlled experiment
- Phase: `phase_2_cost_refinement` — middle-layer input-pipeline cost refinement after the H5 success/cost result
- Independent variable: raw versus semantic-filtered UIAutomator element list
- Controls: same checkout, model, task parameters, generated seed, step budget, and emulator; arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Mean tokens | Input / output tokens |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| control | 4/4 | 1.000 | 1.000 | 4.750 | 101.198 | 8.500 | 139439.500 | 549928 / 7830 |
| treatment | 4/4 | 1.000 | 1.000 | 4.750 | 99.184 | 8.500 | 70557.500 | 274067 / 8163 |
| Task / trial | Control | Treatment | Δ success | Control→treatment steps |
| --- | ---: | ---: | ---: | ---: |
| SystemWifiTurnOff / 1 | 1 | 1 | +0 | 5→5 |
| SystemWifiTurnOffVerify / 1 | 1 | 1 | +0 | 5→5 |
| SystemWifiTurnOn / 1 | 1 | 1 | +0 | 5→5 |
| SystemWifiTurnOnVerify / 1 | 1 | 1 | +0 | 4→4 |
## 4. Data-driven decision
- Outcome: **`promote_efficient_candidate_to_full_suite_rerun`**
- Reason: Treatment preserved paired success with no regression and passed the latency/token efficiency guardrails. This is a candidate decision only.
- Treatment/control mean latency ratio: 0.980
- Treatment/control mean token ratio: 0.506
- Treatment/control mean LLM-call ratio: 1.000
- Cost guardrails passed: **true**
- Deployment approved: **false**
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
### LLM analysis of this run
The following bounded interpretation was produced by the configured real LLM from the aggregate evidence (the JSON remains authoritative):
- Summary: A paired experiment comparing raw UIAutomator (control) and compact UIAutomator (treatment) on 4 system Settings tasks (SystemWifiTurnOff, SystemWifiTurnOffVerify, SystemWifiTurnOn, SystemWifiTurnOnVerify) in an API 35 AVD environment. Both arms completed 4 episodes with 100% success rate, identical mean steps (4.75) and LLM calls (8.5). Treatment showed lower mean total tokens (50.6% of control) and slightly lower mean latency (98.0% of control). The decision was to promote the treatment as a full-suite candidate rerun, as it preserved success, passed latency/token guardrails, but remains a subset with environment limitations.
- Cost/benefit interpretation: The treatment provides significant token efficiency (mean token ratio 0.506) with preserved success and marginal latency improvement (mean latency ratio 0.980) in the tested 4-task subset. However, interpretation is bounded by the API/app environment: results are from an API 35 AVD (not upstream API 33 reference), restricted to system Settings tasks (no full third-party app bundle), and use UIAutomator as a compatibility path (not upstream reference configuration), limiting generalizability beyond the tested scope.
- Next hypothesis `H5C-full` (middle): Evaluate compact UIAutomator (semantic-filtered elements) across the full AndroidWorld task suite to verify token efficiency and success preservation beyond the 4-task Settings subset. Target: Full task suite (all 116 tasks) with upstream reference environment (Pixel 6 / API 33) and provisioned full third-party AndroidWorld app bundle. Verification: Paired run comparing compact UIAutomator (treatment) vs raw UIAutomator (control) across all tasks, ensuring guardrails (mean token ratio ≤0.75, mean latency ratio ≤1.5, success non-inferiority) hold in the reference environment.
## Environment boundaries
- The available AVD is API 35, while upstream is tested on Pixel 6 / API 33. Results are real but not reference-environment comparable.
- The full third-party AndroidWorld app bundle was not provisioned. This run is restricted to system Settings tasks.
- UIAutomator is an upstream AndroidWorld observation option selected by the companion runner; it preserves real UI actions/evaluators but is a compatibility path, not the upstream API-33 reference configuration.
- Compact UIAutomator removes only non-semantic container nodes; observations, coordinates, Android actions, and AndroidWorld evaluators remain real.
- Per-task device-time setting was skipped because the non-root API-35 AVD rejects `adb shell date`; Wi-Fi evaluators do not depend on time.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,67 @@
# Experiment 7-12 AndroidWorld iteration report
- Run ID: `exp7-12-20260729T114648Z`
- Generated (UTC): `2026-07-29T12:13:15Z`
- Upstream commit: `0e95d641e244504c22087cc29b013f3b2428a261`
- Device: `sdk_gphone64_arm64`, API `35` (upstream tested reference: API `33`)
- Provider/model: `ark` / `doubao-seed-1-6-250615`
- Scope: 4 task(s), 1 trial(s), mode `paired`
- Full 116-task × 5-seed suite completed: **false**
The bundled ~88% baseline is historical input evidence. The manuscript's 88%→94% numbers are explicitly hypothetical and are not used as rerun results here.
## 1. Diagnose
- The historical run evaluated 116 tasks once each and reports approximately 88% overall success.
- Wi-Fi is a concentrated failure cluster: three of the four SystemWifiTurn* rows failed in the bundled per-task table.
- The capability matrix links the cluster to weak complex_ui_understanding, information_retrieval, and requires_setup behavior.
- The failed traces show navigation/state-verification loops; increasing the step cap alone would treat a symptom rather than the cause.
## 2. Hypothesis
- ID: `H1`
- Change: Use upstream T3A.set_task_guidelines to add only Wi-Fi Settings navigation and final-state verification guidance.
- Expected measurable result: At least one net paired Wi-Fi success, with no paired regression; record reward, steps, latency, calls, and tokens.
- Guardrails: Same model, seed, task parameters, emulator, checkout, and step budget; do not treat a subset gain as full-suite success.
## 3. Controlled experiment
Control and treatment use the same checkout, model, task parameters, step budget, and emulator. Only the task-specific T3A guidelines differ. Arm order alternates by pair.
| Arm | Episodes | Success | Reward | Steps | Latency (s) | LLM calls | Input / output tokens |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| control | 4/4 | 0.250 | 0.500 | 8.500 | 233.465 | 15.750 | 411525 / 31094 |
| treatment | 4/4 | 0.250 | 0.500 | 8.000 | 156.985 | 12.500 | 190519 / 19520 |
| Task / trial | Control | Treatment | Δ success | Control→treatment steps |
| --- | ---: | ---: | ---: | ---: |
| SystemWifiTurnOff / 1 | 0 | 0 | +0 | 10→10 |
| SystemWifiTurnOffVerify / 1 | 0 | 0 | +0 | 10→10 |
| SystemWifiTurnOn / 1 | 0 | 0 | +0 | 10→10 |
| SystemWifiTurnOnVerify / 1 | 1 | 1 | +0 | 4→2 |
## 4. Data-driven decision
- Outcome: **`inconclusive_no_success_gain`**
- Reason: Treatment produced no paired success gain; keep the upstream control prompt.
- Treatment/control mean latency ratio: 0.672
## 5. Rerun and next report
This run is a real controlled subset/smoke rerun, not the complete AndroidWorld benchmark. The next gate is a conditionally enabled candidate rerun over all 116 tasks with five seeds after provisioning the upstream API-33 app environment.
Observed residual failures:
- `control / SystemWifiTurnOff / trial 1`: evaluator reward / completion gate was not satisfied
- `treatment / SystemWifiTurnOff / trial 1`: evaluator reward / completion gate was not satisfied
- `treatment / SystemWifiTurnOffVerify / trial 1`: evaluator reward / completion gate was not satisfied
- `control / SystemWifiTurnOffVerify / trial 1`: evaluator reward / completion gate was not satisfied
- `control / SystemWifiTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
- `treatment / SystemWifiTurnOn / trial 1`: evaluator reward / completion gate was not satisfied
## Environment boundaries
- The available AVD is API 35, while upstream is tested on Pixel 6 / API 33. Results are real but not reference-environment comparable.
- The full third-party AndroidWorld app bundle was not provisioned. This run is restricted to system Settings tasks.
- Per-task device-time setting was skipped because the non-root API-35 AVD rejects `adb shell date`; Wi-Fi evaluators do not depend on time.
The JSON beside this report is the authoritative evidence. It contains episode-level evaluator rewards, actions, timing, token counts, configuration, and explicit completion gates; credentials and raw prompts are not stored.
+62
View File
@@ -0,0 +1,62 @@
# Data files
*.json
*.jsonl
*.csv
# Python
__pycache__/
*.py[cod]
*$py.class
*.so
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
*.egg-info/
.installed.cfg
*.egg
# Virtual environments
venv/
ENV/
env/
# IDE
.vscode/
.idea/
*.swp
*.swo
*~
# Generated visualizations
*.png
*.html
!index.html
# Canonical Experiment 7-7 evidence is intentionally versioned. The public
# Arena input remains ignored because it is ~2 GB; manifests bind it by URL,
# byte size, record count, and SHA-256 instead.
!validation/
!validation/**/*.json
!validation/**/*.png
!validation/**/*.html
# OS
.DS_Store
Thumbs.db
# Jupyter
.ipynb_checkpoints/
*.ipynb
# Logs
*.log
+786
View File
@@ -0,0 +1,786 @@
# Elo Rating Leaderboard from Pairwise Comparisons
## English
**Experiment 7-7**: Building Model Leaderboard from Pairwise Comparison Data
This project implements an Elo rating system from scratch to analyze model performance using Chatbot Arena's public voting data. The implementation demonstrates how the Bradley-Terry model extracts relative model capabilities from millions of pairwise comparison votes.
## Overview
The Elo rating system is a method for calculating the relative skill levels of players (or in this case, AI models) in zero-sum games. Originally developed for chess, it has been adapted to rank AI language models based on head-to-head comparisons from user votes.
### Key Features
- **High-performance implementation**: NumPy + Numba JIT + parallel processing for optimal speed
- **Real voting data analysis**: Uses actual Chatbot Arena voting data with millions of pairwise comparisons
- **Win rate prediction**: Calculates expected win probabilities between any two models
- **Historical tracking**: Builds time-series snapshots showing ranking evolution
- **Interactive visualizations**: Multiple visualization types including animated bar chart races
- **Scalable**: Efficiently handles 2GB datasets with hundreds of thousands of matches
## Mathematical Foundation
The Elo system is based on the Bradley-Terry model, which models the probability that model A beats model B as:
```
P(A beats B) = 1 / (1 + 10^((R_B - R_A) / 400))
```
After each match, ratings are updated using:
```
R_A_new = R_A + K * (S_A - E_A)
```
Where:
- `R_A` is the current rating of model A
- `K` is the learning rate (K-factor)
- `S_A` is the actual score (1 for win, 0 for loss, 0.5 for tie)
- `E_A` is the expected score (predicted win probability)
## Requirements
- **Disk Space**: At least 3GB free (2GB for data file, 1GB for processing)
- **RAM**: 4GB+ recommended for full dataset analysis
- **Internet**: Stable connection for ~2GB download
- **Python**: 3.12 for the root `ch6` install
## Installation
```bash
# From the repository root: use the shared Chapter 6 environment
uv sync --locked --python 3.12 --extra ch6
# Activate it before changing directories:
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# Windows cmd: .venv\Scripts\activate.bat
# pip fallback when uv is not installed:
# python -m pip install -e ".[ch6]"
cd chapter7/elo-leaderboard
# Single-project compatibility path, still supported during migration:
# python -m pip install -r requirements.txt
```
## Testing
```bash
# From the repository root, include the shared test tooling:
uv sync --locked --python 3.12 --extra ch6 --extra dev
# Activate it before changing directories:
source .venv/bin/activate
# pip fallback when uv is not installed:
# python -m pip install -e ".[ch6,dev]"
cd chapter7/elo-leaderboard
python -m pytest tests
```
## Canonical full-data validation
The accepted Experiment 7-7 run uses the complete public 2024-08-14 Arena
snapshot rather than the synthetic quickstart or a small sample. The 2.0 GB
input is not committed to git; the manifest records its official URL, byte
size, 1,799,991-row count, and SHA-256.
```bash
curl -L \
https://storage.googleapis.com/arena_external_data/public/clean_battle_20240814_public.json \
-o /path/to/arena_data.json
python validation/run_experiment.py \
--input /path/to/arena_data.json \
--output-dir validation/runs/exp7-7-arena-20260731-v1 \
--bootstrap-rounds 20
python validation/validate_evidence.py \
validation/runs/exp7-7-arena-20260731-v1 \
--input /path/to/arena_data.json
```
The retained canonical run accepted 1,670,250 anonymous, deduplicated votes
over 129 models. Chronological online Elo (initial 1000, K=4) and the
Bradley-Terry reconstruction reached Spearman 0.787, Kendall 0.606, and 12/20
top-model overlap. This is the expected broad agreement, not score identity:
online Elo is order-dependent while Bradley-Terry fits all comparisons at
once. Seventeen cumulative monthly snapshots drive the retained D3 animation.
Canonical evidence: [`validation/latest.json`](validation/latest.json).
## 命令行工具 / Command-Line Interface (`cli.py`)
`cli.py` 是本实验统一的 argparse 命令行入口(中文 `--help`),把整条流水线拆成子命令:
**对战 (battle) -> 计算评分 (elo) -> 展示排行榜 (leaderboard)**,并提供 `pipeline` 一步到位。
```bash
python cli.py --help # 查看全部子命令
python cli.py battle --help # 查看某个子命令的参数
# 默认离线端到端演示:模拟对战 -> 在线 Elo -> 最终排行榜表格(无需任何数据/API)
python cli.py # 等价于 python cli.py pipeline
```
### 子命令
| 子命令 | 作用 | 关键参数 |
|--------|------|----------|
| `battle` | 生成两两对战结果 | `--source {simulate,arena,llm}``--num-battles``--tie-prob``--seed``--sample``--output` |
| `elo` | 从对战结果计算评分 | `--method {online-elo,bradley-terry}``--k``--bootstrap``--input``--output` |
| `leaderboard` | 渲染最终排行榜表格 | `--input`(对战或评分文件)、`--method``--bootstrap``--top-n` |
| `pipeline` | 一步跑完 对战 -> Elo -> 排行榜 | 上述参数的并集 |
### 三种对战来源(`--source`
- **`simulate`(默认,纯离线)**:从已知的潜在实力分模拟对战。因为真值已知,可用来**校验**恢复出的排行榜排序是否正确;`--tie-prob` 控制平局比例,用于演练平局处理。
- **`arena`(离线)**:加载真实 Chatbot Arena 投票数据(默认 `arena_data.json`,约 2GB),可用 `--sample N` 抽样。
- **`llm`(需 API)**:用 LLM 做配对评判,并内置**位置偏差消除**——每对交换顺序各评一次,两次判决一致才计胜负、否则记为平局(对应书中 6.4 位置偏差讨论)。仅此来源需要 LLM API Key。
**两种评判后端(`--judge-backend {anthropic,openrouter,auto}`,默认 `auto`**
- `anthropic`:官方 `anthropic` SDK,用 `ANTHROPIC_API_KEY`
- `openrouter`OpenAI 兼容 SDK 指向 `https://openrouter.ai/api/v1`,用 `OPENROUTER_API_KEY`。内部 Claude 名字会自动映射为 OpenRouter id`claude-opus-4-8``anthropic/claude-opus-4.8``claude-haiku-4-5``anthropic/claude-haiku-4.5`);已含 `/` 的 id(如 `openai/gpt-5.6-luna`)原样透传。当直连 Anthropic key 缺失或失效时用它兜底。
- `auto`(默认):有 `ANTHROPIC_API_KEY` 走 anthropic,否则回退 openrouter。注意 `auto` 只看 key 是否存在、不校验有效性;若 `ANTHROPIC_API_KEY` 存在但已失效,请显式 `--judge-backend openrouter`
位置偏差消除与 A/B/tie 解析逻辑与后端无关,两条路径完全一致。
### 分步示例
```bash
# 1) 模拟 5000 场对战(含 10% 平局)
python cli.py battle --source simulate --num-battles 5000 --output battles.json
# 2) 用官方 Bradley-Terry MLE + 100 轮 bootstrap 置信区间计算评分
python cli.py elo --input battles.json --method bradley-terry --bootstrap 100
# 3) 展示前 20 名排行榜(也可直接读评分文件)
python cli.py leaderboard --input battles.json --top-n 20
# 用真实 Arena 数据抽样跑(离线)
python cli.py pipeline --source arena --arena-file arena_data.json --sample 50000 --method bradley-terry --bootstrap 100
# LLM 评判对战(需要 API Key)——官方 Anthropic
export ANTHROPIC_API_KEY=your-anthropic-api-key
python cli.py battle --source llm --candidate-models claude-opus-4-8 claude-haiku-4-5
# LLM 评判对战——通过 OpenRouter 兜底(直连 Anthropic key 缺失/失效时)
export OPENROUTER_API_KEY=your-openrouter-api-key
python cli.py battle --source llm --judge-backend openrouter \
--judge-model claude-opus-4-8 \
--candidate-models anthropic/claude-haiku-4.5 openai/gpt-5.6-luna
```
模拟来源会同时打印真值潜在实力,方便和恢复出的排行榜对照;在线 Elo 与 Bradley-Terry 两种方法都应恢复出与真值一致的排名(分值不必精确对齐,见下文说明)。
## Quick Start
The project implements **two ranking methods** following official Chatbot Arena:
### 1. Bradley-Terry Model (Default - Recommended)
```bash
python main.py
# or explicitly:
python main.py bradley-terry
```
**Use this for**: Official leaderboard, stable rankings, production use
**Key features**:
- ✅ Official Chatbot Arena method
- ✅ Uses sklearn LogisticRegression for Maximum Likelihood Estimation
- ✅ Order-independent (processes all matches simultaneously)
- ✅ Includes 95% confidence intervals via bootstrap (100 samples)
- ✅ More stable and reliable rankings
**Processing time**: ~2-3 minutes (including bootstrap)
### 2. Online Elo (K=4)
```bash
python main.py online-elo
```
**Use this for**: Understanding Elo mechanics, educational purposes, faster computation
**Key features**:
- ✅ K-factor = 4 (official value used by Chatbot Arena)
- ✅ Simple sequential rating updates
- ✅ Order-dependent (processes matches chronologically)
- ✅ Faster computation (~30 seconds)
- ⚠️ Less stable, can vary based on match order
### Method Comparison
| Feature | Bradley-Terry | Online Elo |
|---------|--------------|------------|
| **Stability** | High (MLE fit) | Medium (sequential) |
| **Order dependence** | None | High |
| **Confidence intervals** | Yes (bootstrap) | No |
| **Speed** | Slower (~3 min) | Faster (~30 sec) |
| **Official method** | ✅ Yes | For comparison only |
| **Recommended** | ✅ Production | Educational |
### What Both Methods Do
1. Download Chatbot Arena voting data (~2GB, 5-15 minutes depending on connection)
2. Apply official filters:
- Anonymous votes only (blind evaluation)
- Deduplication (removes top 0.1% redundant prompts)
3. Compute model ratings using selected method
4. Calculate predicted win rates between all model pairs
5. Generate visualizations:
- `leaderboard.png` - Top 20 models ranked by rating
- `rating_distribution.png` - Rating histogram and statistics
- `win_rate_matrix.png` - Predicted win rates (top 30 models)
**Note**: The initial data download is ~2GB and may take several minutes. A progress bar shows download status.
### Quick Demo (Synthetic Data)
To quickly understand Elo mechanics without downloading 2GB:
```bash
python quickstart.py
```
This runs a small demo with synthetic matchups between GPT-4, Claude, Llama, and Gemini.
### Benchmark
To compare both methods:
```bash
python benchmark.py
```
This shows performance and accuracy differences between online Elo and Bradley-Terry approaches.
## Project Structure
```
elo-leaderboard/
├── cli.py # Unified argparse CLI (battle / elo / leaderboard / pipeline)
├── battle_simulator.py # Offline synthetic pairwise-battle generator
├── llm_judge.py # LLM-as-judge battles with position-bias mitigation (needs API)
├── main.py # Main analysis script
├── optimized_elo.py # NumPy + Numba Elo rating system
├── parallel_processing.py # Multi-core parallel processing utilities
├── data_loader.py # Data download and preprocessing
├── leaderboard.py # Leaderboard calculation and analysis
├── visualization.py # Static and interactive visualizations
├── animation.py # Animated bar chart race generator
├── benchmark.py # Performance benchmark tool
├── quickstart.py # Quick demo with synthetic data
├── elo_rating.py # Reference implementation (for comparison)
├── tests/ # Unit and regression tests
├── requirements.txt # Python dependencies
└── README.md # This file
```
## Usage Examples
### Building Elo Leaderboard
```python
from optimized_elo import build_leaderboard_optimized
# Build Elo leaderboard from DataFrame
elo = build_leaderboard_optimized(
df, # DataFrame with columns: model_a, model_b, winner
initial_rating=1000.0,
k_factor=32.0,
show_progress=True
)
# Get leaderboard
leaderboard = elo.get_leaderboard()
for rank, (model, rating, matches, wins) in enumerate(leaderboard[:10], 1):
win_rate = wins / matches * 100 if matches > 0 else 0
print(f"{rank}. {model}: {rating:.1f} ({matches} matches, {win_rate:.1f}% win rate)")
```
### Loading and Filtering Data
```python
from data_loader import load_arena_data, filter_data
# Load data
df = load_arena_data("arena_data.json")
# Filter for blind votes only (reduces bias)
df_filtered = filter_data(
df,
anony_only=True, # Only anonymous votes
language="English", # Specific language
min_turn=1 # Minimum conversation turn
)
```
### Building Historical Leaderboards
```python
from data_loader import get_time_slices
from leaderboard import build_historical_leaderboards, get_rating_history
# Create weekly time slices
time_slices = get_time_slices(df, interval='W')
# Build leaderboard for each time point
historical_leaderboards = build_historical_leaderboards(
df, time_slices, initial_rating=1000.0, k_factor=32.0
)
# Get rating history DataFrame
history_df = get_rating_history(historical_leaderboards)
```
### Creating Visualizations
```python
from visualization import (
plot_leaderboard,
plot_win_rate_matrix,
plot_rating_history,
create_interactive_leaderboard
)
# Static leaderboard chart
plot_leaderboard(leaderboard, top_n=20, save_path="leaderboard.png")
# Win rate heatmap
win_rate_df = calculate_win_rate_matrix_from_data(df)
plot_win_rate_matrix(win_rate_df, top_n=15, save_path="matrix.png")
# Rating evolution
plot_rating_history(history_df, models=["gpt-4", "claude-v1"],
save_path="history.png")
# Interactive chart
fig = create_interactive_leaderboard(history_df, top_n=15)
fig.write_html("interactive.html")
```
### Creating Animated Bar Chart Race
```python
from animation import create_simple_animation
# Generate animated HTML
animation_file = create_simple_animation(
history_df,
output_path="animation.html",
top_n=15
)
# Open animation.html in browser to view
```
## Output Files
After running `main.py`, the following files are generated:
### Static Images (PNG)
- `leaderboard.png` - Current top 20 models ranked by Elo rating
- `rating_distribution.png` - Histogram and box plot of rating distribution
- `win_rate_matrix.png` - Heatmap showing pairwise win rates
- `rating_history.png` - Line chart showing rating evolution over time
### Interactive Visualizations (HTML)
- `interactive_rating_evolution.html` - Interactive chart with zoom/pan
- `interactive_rank_evolution.html` - Interactive rank tracking
- `leaderboard_animation.html` - **Animated bar chart race** showing ranking evolution
## Key Parameters
### Elo System Parameters (Online Elo Method)
- **initial_rating** (default: 1000.0): Starting rating for all models
- **k_factor** (default: 4.0): Learning rate controlling update magnitude
- Official Chatbot Arena uses K=4 for stability
- Higher K-factor (e.g., 32): More volatile, faster adaptation to new data
- Lower K-factor (e.g., 4): More stable, less influenced by recent matches
### Bradley-Terry Parameters
- **SCALE** (400): Elo scale parameter - determines rating point interpretation
- **BASE** (10): Base for logistic function - standard for Elo calculations
- **INIT_RATING** (1000): Initial rating for all models
- **bootstrap_rounds** (100): Number of bootstrap samples for confidence intervals
### Time Slice Intervals
For historical analysis, you can adjust the time granularity:
- `'D'` - Daily snapshots
- `'W'` - Weekly snapshots (recommended)
- `'M'` - Monthly snapshots
### Visualization Parameters
- **top_n**: Number of top models to display (10-20 recommended)
- **Animation speed**: Adjustable in the HTML interface (1x to 10x)
## Data Format
The Chatbot Arena data includes the following fields:
- `model_a`: Identifier for first model
- `model_b`: Identifier for second model
- `winner`: Match outcome ('model_a', 'model_b', or 'tie')
- `tstamp`: Unix timestamp of the vote
- `judge`: User who made the vote
- `turn`: Conversation turn number
- `anony`: Whether vote was anonymous/blind
- `language`: Language of the conversation
## Validation
The implementation validates the Elo predictions against empirical win rates:
```python
from leaderboard import compare_win_rates
# Compare predicted vs actual win rates
comparison = compare_win_rates(elo_system, empirical_win_rates)
mean_error = comparison['error'].mean()
print(f"Mean Absolute Error: {mean_error:.4f}")
```
A low MAE (< 0.05) indicates the Elo model fits the data well.
## Analysis Insights
The project helps identify:
1. **Current Rankings**: Which models are currently strongest
2. **Rating Trends**: How model performance evolves over time
3. **Breakthrough Moments**: When new models enter or shake up rankings
4. **Competitive Dynamics**: Which models are closely matched
5. **Long-term Trajectories**: Models in ascent vs. decline
6. **Rating Stability**: Volatility in model performance
## Performance Architecture
The implementation is designed for high performance on large datasets (2GB+).
### Core Optimizations
#### 1. **NumPy + Numba JIT Compilation**
Uses NumPy arrays and Numba's just-in-time compilation:
- **NumPy arrays** for O(1) integer indexing (vs O(n) dictionary lookups)
- **Numba JIT** compiles hot loops to machine code (50-100x speedup)
- **Pre-allocated arrays** eliminate dynamic memory allocation overhead
- **Integer indices** instead of string model names for cache-friendly access
#### 2. **Multi-Core Parallel Processing**
Parallelizes independent operations across all CPU cores:
- **Historical analysis**: Each time slice processed independently
- **Win rate matrices**: Model pairs computed in parallel chunks
- **Data filtering**: DataFrame operations distributed across cores
```python
from parallel_processing import build_historical_leaderboards_parallel
# Automatically uses all available CPU cores
historical_lb = build_historical_leaderboards_parallel(
df, time_slices, n_jobs=-1
)
```
#### 3. **Memory Optimization**
Reduces memory footprint through intelligent data types:
- Downcasts numeric types (int64 → int32, float64 → float32)
- Converts repetitive strings to categorical types
- Achieves 30-50% memory reduction
```python
from parallel_processing import optimize_dataframe
df = optimize_dataframe(df) # Automatic memory optimization
```
### Performance Characteristics
On typical hardware (4-8 core CPU) with the full 2GB dataset:
| Component | Technique | Impact |
|---|---|---|
| Elo Computation | NumPy + Numba JIT | 50-100x faster |
| Historical Analysis | Multi-core parallel | 4-8x faster |
| Win Rate Matrix | Parallel processing | 4-8x faster |
| Memory Usage | Type optimization | 30-50% reduction |
| **Overall** | **Combined** | **~10-15x speedup** |
**Processing time**: 1-2 minutes for full dataset (hundreds of thousands of matches)
## Advanced Usage
### Custom Analysis
The modular design allows for flexible customization:
### Focused Analysis
```python
from optimized_elo import build_leaderboard_optimized
from data_loader import load_arena_data, filter_data
df = load_arena_data("arena_data.json")
# Analyze only recent data
df_recent = filter_data(df, min_date="2024-01-01")
elo_recent = build_leaderboard_optimized(df_recent)
# Analyze specific model family
gpt_models = [m for m in df['model_a'].unique() if 'gpt' in m.lower()]
df_gpt = df[df['model_a'].isin(gpt_models) & df['model_b'].isin(gpt_models)]
elo_gpt = build_leaderboard_optimized(df_gpt)
```
### Export Results
```python
import pandas as pd
# Export leaderboard to CSV
lb_df = pd.DataFrame(leaderboard, columns=['model', 'rating', 'matches', 'wins'])
lb_df.to_csv('leaderboard.csv', index=False)
# Export rating history
history_df.to_csv('rating_history.csv', index=False)
```
## Troubleshooting
### Data Download Issues
If automatic download fails:
1. Manually download from: https://storage.googleapis.com/arena_external_data/public/clean_battle_20240814_public.json
2. Save as `arena_data.json` in the project directory
3. Run `python main.py` again
### Memory Issues
The dataset is large (~2GB, hundreds of thousands of battles). If you encounter memory issues on systems with limited RAM:
```python
from data_loader import load_arena_data, filter_data, get_time_slices
from optimized_elo import build_leaderboard_optimized
# Load and immediately filter to reduce memory usage
df = load_arena_data("arena_data.json")
# Filter to recent data only
df_filtered = filter_data(df, min_date="2024-01-01", anony_only=True)
# Use monthly instead of weekly intervals for historical analysis
time_slices = get_time_slices(df_filtered, interval='M') # vs 'W' for weekly
# Analyze with smaller top_n for visualizations
elo = build_leaderboard_optimized(df_filtered)
```
The built-in memory optimization reduces footprint by 30-50%, but very large analyses may still require 4GB+ RAM.
### Visualization Issues
- Ensure matplotlib, seaborn, and plotly are installed via the root `ch6` extra or the compatibility `requirements.txt` path.
- For HTML animations, use a modern web browser (Chrome, Firefox, Safari, Edge)
- If plots don't display in Jupyter, use `%matplotlib inline` or save to file
## References
- **Chatbot Arena**: https://chat.lmsys.org/
- **Elo Rating System**: https://en.wikipedia.org/wiki/Elo_rating_system
- **Bradley-Terry Model**: https://en.wikipedia.org/wiki/BradleyTerry_model
- **LMSYS Paper**: "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena"
## Learning Objectives
This experiment demonstrates:
1. **Statistical Modeling**: How pairwise comparisons reveal relative abilities
2. **Online Learning**: Incremental rating updates as new data arrives
3. **Probabilistic Prediction**: Converting rating differences to win probabilities
4. **Data Visualization**: Effective techniques for showing temporal dynamics
5. **Model Evaluation**: Alternative to traditional benchmark approaches
## Technical Details
### Why Elo Computation is Hard to Parallelize
Elo rating computation is **inherently sequential** because each match's rating update depends on the current ratings, which were modified by all previous matches. This is why we can't simply split matches into chunks and process them independently.
However, we can still achieve significant speedups through:
1. **Algorithmic optimization**: NumPy arrays + Numba JIT
2. **Parallelizing independent operations**: Historical analysis, win rate matrices
3. **Memory efficiency**: Better cache utilization
4. **Data structure optimization**: Integer indexing, pre-allocation
### Numba JIT Compilation
The core Elo update loop is compiled to machine code using Numba:
```python
@jit(nopython=True)
def process_elo_updates_vectorized(ratings, model_a_indices, model_b_indices,
outcomes, k_factor, match_counts, win_counts):
for i in range(len(model_a_indices)):
# This loop runs at C speed, not Python speed
# Typical speedup: 50-100x over pure Python
...
```
### Memory Layout
Using NumPy arrays with proper data types:
- `ratings`: float64 array (8 bytes per model)
- `match_counts`: int32 array (4 bytes per model)
- `model_indices`: int32 array (4 bytes per match)
For 500 models and 500K matches: ~10 MB vs ~500 MB for dictionaries.
## Extensions
Potential enhancements:
- Implement Glicko or Glicko-2 rating systems (account for rating uncertainty)
- Add confidence intervals for rating estimates
- Analyze rating by language or task type
- Compare with other ranking methods (e.g., TrueSkill, PageRank)
- Implement time-decay for older matches
- Add statistical significance testing
- Build prediction model for future rankings
- GPU acceleration using CuPy for even larger datasets
- Distributed processing using Dask for multi-machine scaling
## License
This project is part of the AI Agent practical training course materials.
## Contact
For questions or issues, please refer to the course materials or discussion forums.
---
## 中文
该项目围绕**配对比较(pairwise)数据**构建 Elo/Bradley-Terry 排名流程,目标是用公开的模型对战投票数据(重点是 Chatbot Arena)形成可复现的模型排行榜与可视化分析。
### 实验导向背景
Elo 本质上用于“成对对局中的胜率”学习相对能力,最初用于棋类,现被广泛用于语言模型两两对比的排序。
### 关键特性
- 高性能实现:NumPy + Numba JIT + 并行处理。
- 真实数据分析:接入大规模公开投票数据。
- 胜率推断:可预测任意两个模型的胜率。
- 历史追踪:可输出时间序列排行快照。
- 交互可视化:支持静态图与动态动画。
- 可扩展:可承接较大规模比赛集合。
### 数学原理
与 AndroidWorld 风格一致,评分来自 Bradley-Terry
```
P(A 胜过 B) = 1 / (1 + 10^((R_B - R_A) / 400))
```
单步更新:
```
R_A_new = R_A + K * (S_A - E_A)
```
### 安装与运行
```bash
# 在仓库根目录使用统一的第 6 章环境
uv sync --locked --python 3.12 --extra ch6
# 切换目录前先激活环境:
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell.\.venv\Scripts\Activate.ps1
# Windows cmd.venv\Scripts\activate.bat
# 未安装 uv 时可用 pip 兜底:
# python -m pip install -e ".[ch6]"
cd chapter7/elo-leaderboard
# 迁移期间仍支持单项目兼容路径:
# python -m pip install -r requirements.txt
```
### 测试
```bash
# 在仓库根目录安装测试工具:
uv sync --locked --python 3.12 --extra ch6 --extra dev
# 切换目录前先激活环境:
source .venv/bin/activate
# 未安装 uv 时可用 pip 兜底:
# python -m pip install -e ".[ch6,dev]"
cd chapter7/elo-leaderboard
python -m pytest tests
```
### 命令行(`cli.py`
`cli.py` 是统一入口:
- `battle`:生成/采集两两对战
- `elo`:计算评级
- `leaderboard`:出榜
- `pipeline`:端到端一条龙
```bash
python cli.py --help
python cli.py battle --help
python cli.py # 等价于 python cli.py pipeline
```
### 三类对战源
- `simulate`:合成对战(有真值),用于验证是否恢复出正确排序。
- `arena`:离线加载 `arena_data.json`(约 2GB);可用 `--sample` 抽样。
- `llm`:调用 LLM 判断对战,带位置偏差消除;支持 `anthropic``openrouter``auto`
`auto` 会优先使用 Anthropic key,失败时回退 OpenRouter;位置消偏策略与 A/B/tie 判定和后端无关。
### 两种核心评分方法
- Bradley-Terry(推荐):更稳定,适合正式排行。
- Online Elo:更贴近课程里的机制讲解,速度快但对顺序敏感。
### 项目结构
同上方英文学段落中的文件列表。
### 使用示例
核心示例同上英文学:
- `python cli.py battle ...`
- `python cli.py elo ...`
- `python cli.py leaderboard ...`
- `python demo.py` / `python benchmark.py`
### 注意
- `--sample``--pipeline``--top-n` 等参数见命令行帮助。
- 建议先看 CLI 输出再对照 `leaderboard` 与可视化文件确认理解。
+436
View File
@@ -0,0 +1,436 @@
"""
Create animated bar chart race showing leaderboard evolution over time
"""
import pandas as pd
import numpy as np
from typing import List, Tuple
import json
import os
def prepare_animation_data(history_df: pd.DataFrame, top_n: int = 15) -> dict:
"""
Prepare data for D3.js bar chart race animation.
Args:
history_df: DataFrame with columns: date, model, rating, rank
top_n: Number of top models to show at each time point
Returns:
Dictionary with animation data
"""
# Convert date column to Timestamp to support ISO date strings and date objects
if history_df is not None and len(history_df) > 0 and 'date' in history_df.columns:
history_df = history_df.copy()
history_df['date'] = pd.to_datetime(history_df['date'])
# Get all unique dates
dates = sorted(history_df['date'].unique())
if len(dates) == 0:
return {
'frames': [],
'total_frames': 0,
'top_n': top_n,
'start_date': None,
'end_date': None,
}
# For each date, get top N models
frames = []
for date in dates:
date_data = history_df[history_df['date'] == date].nlargest(top_n, 'rating')
frame = {
'date': date.strftime('%Y-%m-%d'),
'timestamp': int(date.timestamp()),
'models': []
}
for rank, row in enumerate(date_data.itertuples(), 1):
frame['models'].append({
'rank': rank,
'name': row.model,
'rating': float(row.rating),
'matches': int(row.matches),
'wins': float(row.wins)
})
frames.append(frame)
animation_data = {
'frames': frames,
'total_frames': len(frames),
'top_n': top_n,
'start_date': dates[0].strftime('%Y-%m-%d'),
'end_date': dates[-1].strftime('%Y-%m-%d')
}
return animation_data
def generate_html_animation(animation_data: dict, output_path: str = "leaderboard_animation.html"):
"""
Generate standalone HTML file with D3.js bar chart race animation.
Args:
animation_data: Dictionary from prepare_animation_data
output_path: Path to save HTML file
"""
html_template = """<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Model Leaderboard Evolution</title>
<script src="https://d3js.org/d3.v7.min.js"></script>
<style>
body {
font-family: 'Segoe UI', Tahoma, Geneva, Verdana, sans-serif;
margin: 0;
padding: 20px;
background: #f5f5f5;
}
#container {
max-width: 1200px;
margin: 0 auto;
background: white;
padding: 30px;
border-radius: 10px;
box-shadow: 0 2px 10px rgba(0,0,0,0.1);
}
h1 {
text-align: center;
color: #333;
margin-bottom: 10px;
}
#date-display {
text-align: center;
font-size: 24px;
font-weight: bold;
color: #666;
margin-bottom: 20px;
}
#chart {
margin: 20px 0;
}
.bar {
fill: steelblue;
cursor: pointer;
transition: fill 0.3s;
}
.bar:hover {
fill: #4682b4;
}
.bar-label {
font-size: 14px;
fill: white;
font-weight: bold;
}
.bar-value {
font-size: 12px;
fill: #333;
}
.rank-label {
font-size: 18px;
fill: #666;
font-weight: bold;
}
#controls {
text-align: center;
margin-top: 30px;
}
button {
padding: 10px 20px;
margin: 0 5px;
font-size: 16px;
cursor: pointer;
border: none;
border-radius: 5px;
background: #4CAF50;
color: white;
transition: background 0.3s;
}
button:hover {
background: #45a049;
}
button:disabled {
background: #ccc;
cursor: not-allowed;
}
#progress-bar {
width: 100%;
height: 5px;
background: #e0e0e0;
margin-top: 20px;
border-radius: 3px;
overflow: hidden;
}
#progress {
height: 100%;
background: #4CAF50;
width: 0%;
transition: width 0.5s;
}
#speed-control {
margin-top: 20px;
text-align: center;
}
#speed-slider {
width: 300px;
margin: 0 10px;
}
.info-box {
background: #f9f9f9;
padding: 15px;
border-radius: 5px;
margin-top: 20px;
font-size: 14px;
color: #666;
}
</style>
</head>
<body>
<div id="container">
<h1>🏆 Model Leaderboard Evolution</h1>
<div id="date-display">Loading...</div>
<div id="chart"></div>
<div id="controls">
<button id="play-btn">▶ Play</button>
<button id="pause-btn" disabled>⏸ Pause</button>
<button id="reset-btn">↺ Reset</button>
</div>
<div id="progress-bar">
<div id="progress"></div>
</div>
<div id="speed-control">
<label>Speed: </label>
<input type="range" id="speed-slider" min="1" max="10" value="5">
<span id="speed-value">5x</span>
</div>
<div class="info-box">
<strong>About:</strong> This animation shows the evolution of model rankings based on Elo ratings
calculated from Chatbot Arena voting data. Each frame represents a snapshot in time,
with models ranked by their current Elo rating. Bars show the rating value,
and the animation reveals how models compete and evolve over time.
</div>
</div>
<script>
const data = """ + json.dumps(animation_data, indent=2) + """;
// Configuration
const margin = {top: 20, right: 100, bottom: 40, left: 50};
const width = 1100 - margin.left - margin.right;
const height = 600 - margin.top - margin.bottom;
const barHeight = height / data.top_n - 5;
// Create SVG
const svg = d3.select("#chart")
.append("svg")
.attr("width", width + margin.left + margin.right)
.attr("height", height + margin.top + margin.bottom)
.append("g")
.attr("transform", `translate(${margin.left},${margin.top})`);
// Scales
const xScale = d3.scaleLinear()
.domain([0, d3.max(data.frames.flatMap(f => f.models.map(m => m.rating)))])
.range([0, width - 200]);
// Color scale
const colorScale = d3.scaleOrdinal(d3.schemeCategory10);
// Animation state
let currentFrame = 0;
let isPlaying = false;
let animationInterval = null;
let animationSpeed = 500; // milliseconds per frame
// Update speed based on slider
d3.select("#speed-slider").on("input", function() {
const speed = +this.value;
animationSpeed = 1000 / speed;
d3.select("#speed-value").text(`${speed}x`);
if (isPlaying) {
stopAnimation();
startAnimation();
}
});
function updateChart(frameIndex) {
const frame = data.frames[frameIndex];
// Update date display
d3.select("#date-display").text(frame.date);
// Update progress bar
const progress = ((frameIndex + 1) / data.total_frames) * 100;
d3.select("#progress").style("width", `${progress}%`);
// Update max value for scale
const maxRating = d3.max(frame.models, d => d.rating);
xScale.domain([0, maxRating * 1.1]);
// Bind data
const bars = svg.selectAll(".bar-group")
.data(frame.models, d => d.name);
// Remove old bars
bars.exit()
.transition()
.duration(animationSpeed * 0.8)
.style("opacity", 0)
.remove();
// Add new bars
const enter = bars.enter()
.append("g")
.attr("class", "bar-group")
.style("opacity", 0);
enter.append("rect")
.attr("class", "bar")
.attr("height", barHeight);
enter.append("text")
.attr("class", "bar-label")
.attr("x", 10)
.attr("y", barHeight / 2)
.attr("dy", "0.35em");
enter.append("text")
.attr("class", "bar-value")
.attr("y", barHeight / 2)
.attr("dy", "0.35em");
enter.append("text")
.attr("class", "rank-label")
.attr("x", -40)
.attr("y", barHeight / 2)
.attr("dy", "0.35em")
.attr("text-anchor", "middle");
// Update all bars
const merged = enter.merge(bars);
merged.transition()
.duration(animationSpeed * 0.8)
.style("opacity", 1)
.attr("transform", (d, i) => `translate(0,${i * (barHeight + 5)})`);
merged.select(".bar")
.transition()
.duration(animationSpeed * 0.8)
.attr("width", d => xScale(d.rating))
.attr("fill", d => colorScale(d.name));
merged.select(".bar-label")
.text(d => d.name);
merged.select(".bar-value")
.transition()
.duration(animationSpeed * 0.8)
.attr("x", d => xScale(d.rating) + 10)
.text(d => `${Math.round(d.rating)} (${d.matches} matches)`);
merged.select(".rank-label")
.text(d => `#${d.rank}`);
}
function startAnimation() {
if (currentFrame >= data.total_frames - 1) {
currentFrame = 0;
}
isPlaying = true;
d3.select("#play-btn").property("disabled", true);
d3.select("#pause-btn").property("disabled", false);
animationInterval = setInterval(() => {
updateChart(currentFrame);
currentFrame++;
if (currentFrame >= data.total_frames) {
stopAnimation();
currentFrame = data.total_frames - 1;
}
}, animationSpeed);
}
function stopAnimation() {
isPlaying = false;
d3.select("#play-btn").property("disabled", false);
d3.select("#pause-btn").property("disabled", true);
if (animationInterval) {
clearInterval(animationInterval);
animationInterval = null;
}
}
function resetAnimation() {
stopAnimation();
currentFrame = 0;
updateChart(currentFrame);
d3.select("#progress").style("width", "0%");
}
// Button handlers
d3.select("#play-btn").on("click", startAnimation);
d3.select("#pause-btn").on("click", stopAnimation);
d3.select("#reset-btn").on("click", resetAnimation);
// Initialize with first frame
updateChart(0);
// Auto-play on load
setTimeout(startAnimation, 1000);
</script>
</body>
</html>"""
# Keep generated evidence friendly to `git diff --check` and deterministic
# across editors that otherwise strip indentation-only lines.
html_template = "\n".join(line.rstrip() for line in html_template.splitlines()) + "\n"
# Write to file
with open(output_path, 'w', encoding='utf-8') as f:
f.write(html_template)
print(f"Generated animation HTML at: {output_path}")
print(f"Open the file in a web browser to view the animation.")
def create_simple_animation(history_df: pd.DataFrame, output_path: str = "leaderboard_animation.html", top_n: int = 15):
"""
Convenience function to create animation in one step.
Args:
history_df: DataFrame with rating history
output_path: Path to save HTML file
top_n: Number of top models to show
"""
print("Preparing animation data...")
animation_data = prepare_animation_data(history_df, top_n)
print(f"Generating HTML animation with {animation_data['total_frames']} frames...")
generate_html_animation(animation_data, output_path)
return output_path
@@ -0,0 +1,77 @@
"""
Synthetic pairwise battle generator (offline).
Generates head-to-head "battle" outcomes from a set of known latent skill
scores, so the whole battles -> Elo -> leaderboard pipeline can be demonstrated
end-to-end without downloading the 2GB Chatbot Arena dataset or calling any API.
Because the ground-truth skills are known, the recovered Elo leaderboard can be
checked against them: the ranking should match, which validates the
implementation. Ties are produced with a configurable probability to exercise
the tie-handling paths in both the online Elo and Bradley-Terry code.
"""
import random
from typing import Dict, List, Optional
# Default roster with plausible latent skills (in Elo points). The exact numbers
# are only used to *generate* battles; the experiment then tries to recover them.
DEFAULT_TRUE_SKILLS: Dict[str, float] = {
"gpt-4": 1250.0,
"claude-3-opus": 1225.0,
"gemini-1.5-pro": 1180.0,
"llama-3-70b": 1120.0,
"mixtral-8x7b": 1075.0,
"gpt-3.5-turbo": 1035.0,
"llama-2-13b": 980.0,
"vicuna-13b": 935.0,
}
def expected_score(rating_a: float, rating_b: float,
base: float = 10.0, scale: float = 400.0) -> float:
"""Bradley-Terry / Elo win probability of A against B."""
return 1.0 / (1.0 + base ** ((rating_b - rating_a) / scale))
def simulate_battles(true_skills: Dict[str, float],
num_battles: int,
tie_prob: float = 0.1,
seed: Optional[int] = None) -> List[dict]:
"""
Simulate `num_battles` random pairwise battles.
For each battle two distinct models are drawn uniformly at random. With
probability `tie_prob` the outcome is a tie; otherwise the winner is sampled
according to the Bradley-Terry win probability implied by the latent skills
(so upsets happen, but stronger models win more often).
Args:
true_skills: Mapping of model name -> latent skill (Elo points).
num_battles: Number of battles to generate.
tie_prob: Probability that a battle ends in a tie.
seed: Optional RNG seed for reproducibility.
Returns:
List of dicts with keys 'model_a', 'model_b', 'winner'
(winner in {'model_a', 'model_b', 'tie'}), matching the Chatbot Arena
schema consumed by the Elo / Bradley-Terry code.
"""
if len(true_skills) < 2:
raise ValueError("Need at least 2 models to simulate battles")
rng = random.Random(seed)
models = list(true_skills.keys())
battles: List[dict] = []
for _ in range(num_battles):
model_a, model_b = rng.sample(models, 2)
if rng.random() < tie_prob:
winner = "tie"
elif rng.random() < expected_score(true_skills[model_a], true_skills[model_b]):
winner = "model_a"
else:
winner = "model_b"
battles.append({"model_a": model_a, "model_b": model_b, "winner": winner})
return battles
+144
View File
@@ -0,0 +1,144 @@
"""
Benchmark script to compare performance of different Elo implementations
"""
import time
import pandas as pd
import numpy as np
from elo_rating import EloRatingSystem
from optimized_elo import build_leaderboard_optimized
from data_loader import load_arena_data, filter_data
def benchmark_basic_elo(df: pd.DataFrame) -> float:
"""Benchmark the basic Elo implementation."""
print("\n" + "="*80)
print("Benchmarking Basic Elo Implementation (Python dict)")
print("="*80)
start_time = time.time()
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
for _, row in df.iterrows():
elo.update_ratings(row['model_a'], row['model_b'], row['winner'])
end_time = time.time()
elapsed = end_time - start_time
leaderboard = elo.get_leaderboard()
print(f"✓ Processed {len(df)} matches in {elapsed:.2f} seconds")
print(f" Speed: {len(df)/elapsed:.0f} matches/second")
print(f" Top 3 models: {[m[0] for m in leaderboard[:3]]}")
return elapsed
def benchmark_optimized_elo(df: pd.DataFrame) -> float:
"""Benchmark the NumPy + Numba optimized implementation."""
print("\n" + "="*80)
print("Benchmarking Optimized Elo Implementation (NumPy + Numba JIT)")
print("="*80)
start_time = time.time()
elo = build_leaderboard_optimized(
df,
initial_rating=1000.0,
k_factor=32.0,
show_progress=False
)
end_time = time.time()
elapsed = end_time - start_time
leaderboard = elo.get_leaderboard()
print(f"✓ Processed {len(df)} matches in {elapsed:.2f} seconds")
print(f" Speed: {len(df)/elapsed:.0f} matches/second")
print(f" Top 3 models: {[m[0] for m in leaderboard[:3]]}")
return elapsed
def main():
"""Run benchmark comparison."""
print("="*80)
print("ELO RATING COMPUTATION BENCHMARK")
print("="*80)
print("\nThis benchmark compares the performance of different Elo implementations")
print("on Chatbot Arena voting data.\n")
# Load data
print("Loading data...")
try:
df = load_arena_data("arena_data.json")
except FileNotFoundError:
print("Error: arena_data.json not found. Please run main.py first to download the data.")
return
# Filter for blind votes
print("Filtering data...")
df_filtered = filter_data(df, anony_only=True, min_turn=1)
# Use a subset for quick benchmarking (can change to full dataset)
sample_size = 50000
if len(df_filtered) > sample_size:
print(f"\nUsing a sample of {sample_size} matches for benchmarking")
print("(To benchmark on full dataset, set sample_size = len(df_filtered))")
df_sample = df_filtered.head(sample_size).copy()
else:
df_sample = df_filtered.copy()
print(f"\nBenchmark dataset: {len(df_sample)} matches")
print(f"Unique models: {len(set(df_sample['model_a'].unique()) | set(df_sample['model_b'].unique()))}")
# Warm up Numba JIT (first run compiles the functions)
print("\n" + "-"*80)
print("Warming up Numba JIT compiler (first run)...")
print("-"*80)
df_tiny = df_sample.head(1000)
build_leaderboard_optimized(df_tiny, show_progress=False)
print("✓ JIT compilation complete")
# Run benchmarks
time_basic = benchmark_basic_elo(df_sample)
time_optimized = benchmark_optimized_elo(df_sample)
# Results summary
print("\n" + "="*80)
print("BENCHMARK RESULTS")
print("="*80)
speedup = time_basic / time_optimized if time_optimized > 0 else 0
print(f"\nBasic Implementation: {time_basic:8.2f} seconds")
print(f"Optimized Implementation: {time_optimized:8.2f} seconds")
print(f"\nSpeedup: {speedup:.1f}x faster")
pct_reduction = (1 - time_optimized / time_basic) * 100 if time_basic > 0 else 0.0
print(f"Time saved: {time_basic - time_optimized:.2f} seconds ({pct_reduction:.1f}% reduction)")
# Extrapolate to full dataset
if len(df_sample) > 0 and len(df_sample) < len(df_filtered):
full_time_basic = time_basic * (len(df_filtered) / len(df_sample))
full_time_optimized = time_optimized * (len(df_filtered) / len(df_sample))
print(f"\nExtrapolated times for full dataset ({len(df_filtered)} matches):")
print(f" Basic: ~{full_time_basic/60:.1f} minutes")
print(f" Optimized: ~{full_time_optimized/60:.1f} minutes")
print(f" Time saved: ~{(full_time_basic - full_time_optimized)/60:.1f} minutes")
print("\n" + "="*80)
print("\nOptimization Techniques Applied:")
print(" • NumPy arrays instead of Python dicts (O(1) integer indexing)")
print(" • Numba JIT compilation (compiles hot loops to machine code)")
print(" • Pre-allocated arrays (no dynamic memory allocation)")
print(" • Integer model indices (no string lookups)")
print(" • Vectorized operations where possible")
print("\nFor the full optimized pipeline with parallel processing, run main_optimized.py")
print("="*80 + "\n")
if __name__ == "__main__":
main()
+238
View File
@@ -0,0 +1,238 @@
"""
Bradley-Terry Model Implementation
Official Chatbot Arena leaderboard calculation method
"""
import math
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
def compute_mle_elo(df: pd.DataFrame,
SCALE: int = 400,
BASE: int = 10,
INIT_RATING: int = 1000,
calibration_model: str | None = None,
calibration_rating: int | None = None) -> pd.Series:
"""
Compute Elo ratings using Bradley-Terry model with Maximum Likelihood Estimation.
This is the official method used by Chatbot Arena for their leaderboard.
It uses sklearn's LogisticRegression to fit a Bradley-Terry model.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
SCALE: Elo scale parameter (default 400)
BASE: Base for logistic function (default 10)
INIT_RATING: Initial rating (default 1000)
calibration_model: Model name to calibrate ratings to
calibration_rating: Target rating for calibration model
Returns:
Series of Elo ratings indexed by model name
"""
# Empty battle frame (e.g. --num-battles 0 or fully filtered input) is valid.
if df is None or len(df) == 0:
return pd.Series(dtype=float)
models = sorted({m for m in (set(df["model_a"]) | set(df["model_b"])) if pd.notna(m)})
if len(models) <= 1:
res_dict = {}
for m in models:
if calibration_model == m and calibration_rating is not None:
res_dict[m] = float(calibration_rating)
else:
res_dict[m] = float(INIT_RATING)
return pd.Series(res_dict, index=pd.Index(models), dtype=float)
# Create pivot tables for wins
ptbl_a_win = pd.pivot_table(
df[df["winner"] == "model_a"],
index="model_a",
columns="model_b",
aggfunc="size",
fill_value=0,
observed=False
)
# Handle ties (including "tie (bothbad)"). Symmetrize only after aligning to
# the full model square; (pivot + pivot.T) on a one-sided A×B pivot zeroes
# every cell via pandas index/column alignment.
if sum(df["winner"].isin(["tie", "tie (bothbad)"])) == 0:
ptbl_tie = pd.DataFrame(0, index=ptbl_a_win.index, columns=ptbl_a_win.columns)
else:
ptbl_tie = pd.pivot_table(
df[df["winner"].isin(["tie", "tie (bothbad)"])],
index="model_a",
columns="model_b",
aggfunc="size",
fill_value=0,
observed=False
)
ptbl_b_win = pd.pivot_table(
df[df["winner"] == "model_b"],
index="model_a",
columns="model_b",
aggfunc="size",
fill_value=0,
observed=False
)
# Align pivots on the full model universe (small samples otherwise leave NaNs).
models = sorted({m for m in (set(df["model_a"]) | set(df["model_b"])) if pd.notna(m)})
ptbl_a_win = ptbl_a_win.reindex(index=models, columns=models, fill_value=0)
ptbl_b_win = ptbl_b_win.reindex(index=models, columns=models, fill_value=0)
ptbl_tie = ptbl_tie.reindex(index=models, columns=models, fill_value=0)
ptbl_tie = ptbl_tie + ptbl_tie.T
# Compute win matrix (A wins * 2 + B wins * 2 + ties)
ptbl_win = (ptbl_a_win * 2 + ptbl_b_win.T * 2 + ptbl_tie).fillna(0)
# Map models to indices
models = pd.Series(np.arange(len(ptbl_win.index)), index=ptbl_win.index)
p = len(models)
X = np.zeros([p * (p - 1) * 2, p])
Y = np.zeros(p * (p - 1) * 2)
cur_row = 0
sample_weights = []
for m_a in ptbl_win.index:
for m_b in ptbl_win.columns:
if m_a == m_b:
continue
# Skip if nan
if math.isnan(ptbl_win.loc[m_a, m_b]) or math.isnan(ptbl_win.loc[m_b, m_a]):
continue
X[cur_row, models[m_a]] = +math.log(BASE)
X[cur_row, models[m_b]] = -math.log(BASE)
Y[cur_row] = 1.0
sample_weights.append(ptbl_win.loc[m_a, m_b])
X[cur_row + 1, models[m_a]] = math.log(BASE)
X[cur_row + 1, models[m_b]] = -math.log(BASE)
Y[cur_row + 1] = 0.0
sample_weights.append(ptbl_win.loc[m_b, m_a])
cur_row += 2
X = X[:cur_row]
Y = Y[:cur_row]
# Fit logistic regression
lr = LogisticRegression(fit_intercept=False, penalty=None, tol=1e-6)
lr.fit(X, Y, sample_weight=sample_weights)
# Convert to Elo scores
elo_scores = SCALE * lr.coef_[0] + INIT_RATING
# Calibrate to reference model if provided
if calibration_model and calibration_model in models.index:
target_rating = INIT_RATING if calibration_rating is None else calibration_rating
elo_scores += target_rating - elo_scores[models[calibration_model]]
return pd.Series(elo_scores, index=models.index).sort_values(ascending=False)
def predict_win_rate(elo_ratings: dict[str, float],
SCALE: int = 400,
BASE: int = 10) -> pd.DataFrame:
"""
Predict win rates between all model pairs using Elo ratings.
Args:
elo_ratings: Dictionary of model names to Elo ratings
SCALE: Elo scale parameter
BASE: Base for logistic function
Returns:
DataFrame with predicted win rates (row vs column)
"""
from collections import defaultdict
names = sorted(elo_ratings)
wins = defaultdict(lambda: defaultdict(lambda: 0))
for a in names:
for b in names:
ea = 1 / (1 + BASE ** ((elo_ratings[b] - elo_ratings[a]) / SCALE))
wins[a][b] = ea
wins[b][a] = 1 - ea
data = {
# np.nan, not np.NAN: the upper-case aliases were removed in NumPy 2.0
# and requirements.txt allows numpy>=1.24 (i.e. 2.x).
a: [wins[a][b] if a != b else np.nan for b in names]
for a in names
}
df = pd.DataFrame(data, index=names)
df.index.name = "model_a"
df.columns.name = "model_b"
return df.T
def get_bootstrap_result(battles: pd.DataFrame,
func_compute_elo,
num_round: int = 100,
random_seed: int = 0) -> pd.DataFrame:
"""
Compute bootstrap confidence intervals for Elo ratings.
Args:
battles: DataFrame with battle data
func_compute_elo: Function to compute Elo ratings
num_round: Number of bootstrap rounds
Returns:
DataFrame with ratings from each bootstrap round
"""
from tqdm import tqdm
rows = []
for i in tqdm(range(num_round), desc="Bootstrap sampling"):
rows.append(
func_compute_elo(
battles.sample(frac=1.0, replace=True, random_state=random_seed + i)
)
)
df = pd.DataFrame(rows)
return df[df.median().sort_values(ascending=False).index]
def compute_bradley_terry_leaderboard(df: pd.DataFrame,
bootstrap_rounds: int = 0) -> pd.DataFrame:
"""
Compute leaderboard using Bradley-Terry model (official Chatbot Arena method).
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
bootstrap_rounds: Number of bootstrap rounds for confidence intervals (0 = no bootstrap)
Returns:
DataFrame with model ratings (and confidence intervals if bootstrap > 0)
"""
print("Computing Bradley-Terry model ratings...")
# Compute MLE Elo ratings
elo_ratings = compute_mle_elo(df)
if bootstrap_rounds > 0:
print(f"Computing {bootstrap_rounds} bootstrap samples for confidence intervals...")
bootstrap_df = get_bootstrap_result(df, compute_mle_elo, bootstrap_rounds)
# Compute confidence intervals
result = pd.DataFrame({
'rating': bootstrap_df.quantile(0.5),
'lower_ci': bootstrap_df.quantile(0.025),
'upper_ci': bootstrap_df.quantile(0.975)
}).sort_values('rating', ascending=False)
else:
result = pd.DataFrame({
'rating': elo_ratings
}).sort_values('rating', ascending=False)
result.index.name = 'model'
return result.reset_index()
+343
View File
@@ -0,0 +1,343 @@
#!/usr/bin/env python3
"""
实验 7-7:从配对比较数据构建模型排行榜 —— 命令行入口
统一的 argparse 命令行工具,把整个流程拆成三个子命令:
battle 运行两两对战,生成对战结果(模拟 / Chatbot Arena 真实数据 / LLM 评判)
elo 从对战结果计算 Elo 或 Bradley-Terry 评分
leaderboard 把对战结果或评分渲染成最终排行榜表格
pipeline 一步跑完 对战 -> Elo -> 排行榜(默认离线可复现)
其中 battle 的 simulate/arena 来源与 elo、leaderboard、pipeline 均为纯离线计算,
无需任何 API;只有 --source llmLLM 评判对战)需要 LLM API Key:优先用官方
AnthropicANTHROPIC_API_KEY),若无则自动回退到 OpenRouterOPENROUTER_API_KEY),
也可用 --judge-backend openrouter 强制走 OpenRouterdirect key 失效时)。
示例:
# 离线一条龙:模拟对战 -> Elo -> 排行榜
python cli.py pipeline
# 分步运行
python cli.py battle --source simulate --num-battles 5000 --output battles.json
python cli.py elo --input battles.json --method bradley-terry --bootstrap 100
python cli.py leaderboard --input battles.json --top-n 20
"""
import argparse
import json
import os
import sys
import warnings
from typing import List, Optional
import pandas as pd
# Bradley-Terry 的 LogisticRegression 在新版 sklearn 会对 penalty=None 抛
# FutureWarningbootstrap 会重复上百次,这里静音以保持排行榜输出整洁。
warnings.filterwarnings("ignore", category=FutureWarning, module="sklearn")
from battle_simulator import DEFAULT_TRUE_SKILLS, simulate_battles
from elo_rating import EloRatingSystem
# --------------------------------------------------------------------------- #
# 通用辅助函数
# --------------------------------------------------------------------------- #
def _load_battles(path: str) -> pd.DataFrame:
"""从 JSON 文件加载对战结果,返回带 model_a/model_b/winner 列的 DataFrame。"""
with open(path, "r", encoding="utf-8") as f:
data = json.load(f)
df = pd.DataFrame(data)
# `[]` from --num-battles 0 has no columns; treat as empty battle frame.
if len(df) == 0:
return pd.DataFrame(columns=["model_a", "model_b", "winner"])
required = {"model_a", "model_b", "winner"}
if not required.issubset(df.columns):
raise ValueError(
f"对战文件 {path} 缺少必要字段 {required},实际字段:{list(df.columns)}"
)
return df
def _save_json(obj, path: str) -> None:
with open(path, "w", encoding="utf-8") as f:
json.dump(obj, f, ensure_ascii=False, indent=2)
def _battle_stats(df: pd.DataFrame) -> dict:
"""统计每个模型的对战场数与胜场(平局按 0.5 计)。"""
matches: dict = {}
wins: dict = {}
for model_a, model_b, winner in zip(df["model_a"], df["model_b"], df["winner"]):
matches[model_a] = matches.get(model_a, 0) + 1
matches[model_b] = matches.get(model_b, 0) + 1
if winner == "model_a":
wins[model_a] = wins.get(model_a, 0) + 1.0
elif winner == "model_b":
wins[model_b] = wins.get(model_b, 0) + 1.0
else: # tie / tie (bothbad)
wins[model_a] = wins.get(model_a, 0) + 0.5
wins[model_b] = wins.get(model_b, 0) + 0.5
return {"matches": matches, "wins": wins}
def _compute_online_elo(df: pd.DataFrame, k: float, init_rating: float) -> pd.DataFrame:
"""在线增量 Elo(按记录顺序处理),返回带 model/rating 列的 DataFrame。"""
elo = EloRatingSystem(initial_rating=init_rating, k_factor=k)
for model_a, model_b, winner in zip(df["model_a"], df["model_b"], df["winner"]):
elo.update_ratings(model_a, model_b, winner)
rows = [(m, r) for m, r, *_ in elo.get_leaderboard()]
return pd.DataFrame(rows, columns=["model", "rating"])
def _compute_bradley_terry(df: pd.DataFrame, bootstrap: int) -> pd.DataFrame:
"""Bradley-Terry MLE 评分(可选 bootstrap 置信区间)。"""
# 延迟导入:Bradley-Terry 依赖 scikit-learn,仅在需要时加载。
from bradley_terry import compute_bradley_terry_leaderboard
return compute_bradley_terry_leaderboard(df, bootstrap_rounds=bootstrap)
def _compute_ratings(df: pd.DataFrame, method: str, k: float,
init_rating: float, bootstrap: int) -> pd.DataFrame:
if method == "bradley-terry":
return _compute_bradley_terry(df, bootstrap)
return _compute_online_elo(df, k, init_rating)
def _print_leaderboard(ratings: pd.DataFrame, df: Optional[pd.DataFrame],
top_n: int, title: str) -> None:
"""打印最终排行榜表格。若评分含置信区间则展示 95% CI 列。"""
has_ci = {"lower_ci", "upper_ci"}.issubset(ratings.columns)
stats = _battle_stats(df) if df is not None else {"matches": {}, "wins": {}}
ratings = ratings.sort_values("rating", ascending=False).reset_index(drop=True)
print("=" * 78)
print(title)
print("=" * 78)
if has_ci:
header = f"{'排名':<6}{'模型':<24}{'Elo':>8} {'95% 置信区间':<20}{'场数':>7}{'胜率':>9}"
else:
header = f"{'排名':<6}{'模型':<24}{'Elo':>8} {'场数':>7}{'胜率':>9}"
print(header)
print("-" * 78)
for idx, row in ratings.head(top_n).iterrows():
model = str(row["model"])
n = stats["matches"].get(model, 0)
w = stats["wins"].get(model, 0.0)
win_rate = (w / n * 100.0) if n else 0.0
if has_ci:
ci = f"[{row['lower_ci']:.0f}, {row['upper_ci']:.0f}]"
print(f"{idx + 1:<6}{model:<24}{row['rating']:>8.1f} "
f"{ci:<20}{n:>7}{win_rate:>8.1f}%")
else:
print(f"{idx + 1:<6}{model:<24}{row['rating']:>8.1f} "
f"{n:>7}{win_rate:>8.1f}%")
print("-" * 78)
print(f"{len(ratings)} 个模型,"
f"评分范围 {ratings['rating'].min():.1f} ~ {ratings['rating'].max():.1f}")
if has_ci:
avg_ci = (ratings["upper_ci"] - ratings["lower_ci"]).mean()
print(f"平均 95% 置信区间宽度:{avg_ci:.1f}")
print()
# --------------------------------------------------------------------------- #
# 子命令实现
# --------------------------------------------------------------------------- #
def _make_battles(args) -> List[dict]:
if args.source == "simulate":
skills = DEFAULT_TRUE_SKILLS
if args.models:
# 用户指定模型名时,围绕 1000 分等距分配潜在实力。
n = len(args.models)
skills = {m: 1000.0 + (n - 1 - 2 * i) * 40.0 for i, m in enumerate(args.models)}
print(f"模拟 {args.num_battles} 场对战({len(skills)} 个模型,"
f"平局概率 {args.tie_prob},随机种子 {args.seed}...")
battles = simulate_battles(skills, args.num_battles,
tie_prob=args.tie_prob, seed=args.seed)
print("真实潜在实力(用于事后对照):")
for m, s in sorted(skills.items(), key=lambda kv: -kv[1]):
print(f" {m:<24}{s:>8.1f}")
return battles
if args.source == "arena":
from data_loader import load_arena_data, filter_data
from parallel_processing import optimize_dataframe
if not os.path.exists(args.arena_file):
print(f"错误:找不到 Chatbot Arena 数据文件 {args.arena_file}", file=sys.stderr)
print("可从以下地址下载并保存为该文件名:", file=sys.stderr)
print("https://storage.googleapis.com/arena_external_data/public/"
"clean_battle_20240814_public.json", file=sys.stderr)
sys.exit(1)
df = load_arena_data(args.arena_file)
df = optimize_dataframe(df)
df = filter_data(df, anony_only=True, use_dedup=True, min_turn=1)
if args.sample and args.sample < len(df):
df = df.sample(n=args.sample, random_state=args.seed).reset_index(drop=True)
print(f"采样 {args.sample} 场对战。")
return df[["model_a", "model_b", "winner"]].to_dict("records")
# source == "llm"
from llm_judge import run_llm_battles
print("运行 LLM 评判对战(顺序交换以消除位置偏差)...")
return run_llm_battles(
candidate_models=args.candidate_models,
judge_model=args.judge_model,
backend=args.judge_backend,
)
def cmd_battle(args) -> None:
battles = _make_battles(args)
_save_json(battles, args.output)
print(f"\n已生成 {len(battles)} 场对战,写入 {args.output}")
def cmd_elo(args) -> None:
df = _load_battles(args.input)
print(f"{args.input} 加载 {len(df)} 场对战,方法:{args.method}")
ratings = _compute_ratings(df, args.method, args.k, args.init_rating, args.bootstrap)
_print_leaderboard(ratings, df, top_n=args.top_n,
title=f"Elo 评分({args.method}")
if args.output:
_save_json(ratings.to_dict("records"), args.output)
print(f"评分已写入 {args.output}")
def cmd_leaderboard(args) -> None:
with open(args.input, "r", encoding="utf-8") as f:
data = json.load(f)
sample = data[0] if isinstance(data, list) and data else {}
if "rating" in sample: # 输入已是评分文件,直接展示。
ratings = pd.DataFrame(data)
_print_leaderboard(ratings, None, top_n=args.top_n, title="模型排行榜")
return
# 否则视为对战文件:先计算评分再展示。
df = _load_battles(args.input)
print(f"{args.input} 加载 {len(df)} 场对战,方法:{args.method}")
ratings = _compute_ratings(df, args.method, args.k, args.init_rating, args.bootstrap)
_print_leaderboard(ratings, df, top_n=args.top_n, title="模型排行榜")
def cmd_pipeline(args) -> None:
print("=" * 78)
print("实验 7-7:对战 -> Elo -> 排行榜(端到端)")
print("=" * 78)
battles = _make_battles(args)
if args.output:
_save_json(battles, args.output)
print(f"对战结果写入 {args.output}")
df = pd.DataFrame(battles)
print(f"\n{args.method} 方法从 {len(df)} 场对战计算评分...")
ratings = _compute_ratings(df, args.method, args.k, args.init_rating, args.bootstrap)
_print_leaderboard(ratings, df, top_n=args.top_n, title="最终排行榜")
# --------------------------------------------------------------------------- #
# 参数解析
# --------------------------------------------------------------------------- #
def _add_source_args(parser: argparse.ArgumentParser) -> None:
parser.add_argument("--source", choices=["simulate", "arena", "llm"],
default="simulate",
help="对战来源:simulate=离线模拟(默认)arena=真实 Chatbot Arena 数据,"
"llm=LLM 评判(需 API)")
parser.add_argument("--models", nargs="+", default=None,
help="simulate:自定义模型名列表(默认使用内置 8 个模型)")
parser.add_argument("--num-battles", type=int, default=3000,
help="simulate:模拟对战场数(默认 3000)")
parser.add_argument("--tie-prob", type=float, default=0.1,
help="simulate:平局概率(默认 0.1")
parser.add_argument("--seed", type=int, default=42,
help="随机种子(默认 42")
parser.add_argument("--arena-file", default="arena_data.json",
help="arenaChatbot Arena 数据文件路径(默认 arena_data.json")
parser.add_argument("--sample", type=int, default=0,
help="arena:随机采样 N 场对战,0 表示全部(默认 0)")
parser.add_argument("--candidate-models", nargs="+", default=None,
help="llm:参与对战的候选模型(默认 Claude 系列)")
parser.add_argument("--judge-model", default="claude-opus-4-8",
help="llm:评判模型(默认 claude-opus-4-8")
parser.add_argument("--judge-backend", choices=["anthropic", "openrouter", "auto"],
default="auto",
help="llm:评判后端。auto=有 ANTHROPIC_API_KEY 用官方 Anthropic"
"否则回退到 OpenRouterOPENROUTER_API_KEY);"
"openrouter=强制走 OpenRouterdirect key 失效时用)")
def _add_rating_args(parser: argparse.ArgumentParser) -> None:
parser.add_argument("--method", choices=["online-elo", "bradley-terry"],
default="online-elo",
help="评分方法:online-elo=在线增量 Elo(默认)"
"bradley-terry=官方 MLE 拟合")
parser.add_argument("--k", type=float, default=4.0,
help="online-eloK 因子/学习率(默认 4.0,官方取值)")
parser.add_argument("--init-rating", type=float, default=1000.0,
help="初始评分(默认 1000")
parser.add_argument("--bootstrap", type=int, default=0,
help="bradley-terrybootstrap 轮数以估计 95%% 置信区间(默认 0=不估计)")
parser.add_argument("--top-n", type=int, default=20,
help="排行榜展示的模型数量(默认 20")
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="cli.py",
description="实验 7-7:从配对比较数据构建模型排行榜(对战 -> Elo -> 排行榜)",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
sub = parser.add_subparsers(dest="command", metavar="{battle,elo,leaderboard,pipeline}")
# battle
p_battle = sub.add_parser("battle", help="运行两两对战,生成对战结果")
_add_source_args(p_battle)
p_battle.add_argument("--output", default="battles.json",
help="对战结果输出文件(默认 battles.json")
p_battle.set_defaults(func=cmd_battle)
# elo
p_elo = sub.add_parser("elo", help="从对战结果计算 Elo / Bradley-Terry 评分")
p_elo.add_argument("--input", default="battles.json",
help="对战结果输入文件(默认 battles.json")
_add_rating_args(p_elo)
p_elo.add_argument("--output", default=None,
help="把评分写入 JSON 文件(可选)")
p_elo.set_defaults(func=cmd_elo)
# leaderboard
p_lb = sub.add_parser("leaderboard", help="显示最终排行榜表格")
p_lb.add_argument("--input", default="battles.json",
help="对战结果或评分输入文件(默认 battles.json")
_add_rating_args(p_lb)
p_lb.set_defaults(func=cmd_leaderboard)
# pipeline
p_pipe = sub.add_parser("pipeline", help="一步跑完 对战 -> Elo -> 排行榜(默认离线)")
_add_source_args(p_pipe)
_add_rating_args(p_pipe)
p_pipe.add_argument("--output", default=None,
help="把对战结果写入 JSON 文件(可选)")
p_pipe.set_defaults(func=cmd_pipeline)
return parser
def main(argv: Optional[List[str]] = None) -> None:
parser = build_parser()
# 无子命令时默认运行离线端到端演示,保留开箱即用体验。
args = parser.parse_args(argv if argv is not None else (sys.argv[1:] or ["pipeline"]))
try:
args.func(args)
except (RuntimeError, FileNotFoundError, ValueError) as exc:
print(f"错误:{exc}", file=sys.stderr)
sys.exit(1)
except Exception as exc: # 例如无效 ANTHROPIC_API_KEY 触发的 anthropic.AuthenticationError
print(f"错误:{type(exc).__name__}: {exc}", file=sys.stderr)
print("(若为 LLM 评审路径,请检查对应 provider 的 API key 是否有效)", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
+223
View File
@@ -0,0 +1,223 @@
"""
Data loading and preprocessing for Chatbot Arena voting data
"""
import pandas as pd
import requests
import os
from typing import Optional
from tqdm import tqdm
def download_arena_data(output_path: str = "arena_data.json", force_download: bool = False) -> str:
"""
Download Chatbot Arena voting data via HTTPS.
Args:
output_path: Path to save downloaded file
force_download: If True, re-download even if file exists
Returns:
Path to downloaded file
"""
if os.path.exists(output_path) and not force_download:
print(f"Data file already exists at {output_path}")
file_size = os.path.getsize(output_path) / (1024 * 1024)
print(f"File size: {file_size:.2f} MB")
return output_path
print("Downloading Chatbot Arena voting data...")
url = "https://storage.googleapis.com/arena_external_data/public/clean_battle_20240814_public.json"
try:
# Stream download with progress bar
response = requests.get(url, stream=True)
response.raise_for_status()
# Get total file size
total_size = int(response.headers.get('content-length', 0))
# Download with progress bar
with open(output_path, 'wb') as f, tqdm(
desc=output_path,
total=total_size,
unit='B',
unit_scale=True,
unit_divisor=1024,
) as pbar:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
pbar.update(len(chunk))
file_size = os.path.getsize(output_path) / (1024 * 1024)
print(f"\nDownloaded data to {output_path} ({file_size:.2f} MB)")
return output_path
except Exception as e:
print(f"Error downloading data: {e}")
print("Please ensure you have internet connection and the URL is accessible.")
raise
def load_arena_data(filepath: str) -> pd.DataFrame:
"""
Load and preprocess Chatbot Arena voting data.
Expected columns:
- model_a: Identifier for first model
- model_b: Identifier for second model
- winner: Which model won ('model_a', 'model_b', or 'tie')
- tstamp: Unix timestamp of the vote
- judge: User who made the vote
- turn: Conversation turn
- anony: Whether vote was anonymous (blind)
- language: Language of the conversation
Args:
filepath: Path to data file
Returns:
Preprocessed DataFrame sorted by timestamp
"""
print(f"Loading data from {filepath}...")
print("Note: This is a large file (~2GB), loading may take 1-2 minutes...")
# Try different file formats
if filepath.endswith('.json'):
df = pd.read_json(filepath)
elif filepath.endswith('.jsonl'):
df = pd.read_json(filepath, lines=True)
elif filepath.endswith('.csv'):
df = pd.read_csv(filepath)
else:
# Try JSON by default
try:
df = pd.read_json(filepath)
except (ValueError, KeyError):
df = pd.read_json(filepath, lines=True)
print(f"Loaded {len(df)} records")
print(f"Columns: {df.columns.tolist()}")
# Sort by timestamp
if 'tstamp' in df.columns:
df = df.sort_values('tstamp', ascending=True).reset_index(drop=True)
print(f"Data spans from {pd.to_datetime(df['tstamp'].min(), unit='s')} to {pd.to_datetime(df['tstamp'].max(), unit='s')}")
# Basic statistics
if 'winner' in df.columns:
print(f"\nOutcome distribution:")
print(df['winner'].value_counts())
if 'model_a' in df.columns and 'model_b' in df.columns:
all_models = set(df['model_a'].unique()) | set(df['model_b'].unique())
print(f"\nTotal unique models: {len(all_models)}")
print(f"Top 10 models by appearance:")
model_counts = pd.concat([df['model_a'], df['model_b']]).value_counts().head(10)
print(model_counts)
return df
def filter_data(df: pd.DataFrame,
min_date: Optional[str] = None,
max_date: Optional[str] = None,
anony_only: bool = True,
language: Optional[str] = None,
min_turn: int = 1,
use_dedup: bool = True) -> pd.DataFrame:
"""
Filter voting data based on various criteria (following official Chatbot Arena method).
Args:
df: Input DataFrame
min_date: Minimum date (YYYY-MM-DD format)
max_date: Maximum date (YYYY-MM-DD format)
anony_only: If True, only include anonymous (blind) votes
language: If specified, filter by language
min_turn: Minimum conversation turn
use_dedup: If True, apply deduplication filter (official Arena method)
Returns:
Filtered DataFrame
"""
filtered = df.copy()
print(f"Before filtering: {len(filtered)} records")
# Filter by anonymous votes only (official method)
if anony_only and 'anony' in filtered.columns:
filtered = filtered[filtered['anony'] == True]
print(f" After anony filter: {len(filtered)} records")
# Apply deduplication (official method removes top 0.1% redundant prompts)
if use_dedup and 'dedup_tag' in filtered.columns:
try:
filtered = filtered[filtered["dedup_tag"].apply(lambda x: x.get("sampled", False) if isinstance(x, dict) else False)]
print(f" After dedup filter: {len(filtered)} records")
except Exception as e:
print(f" Warning: Could not apply dedup filter: {e}")
# Filter by date
if 'tstamp' in filtered.columns:
if min_date:
min_timestamp = pd.to_datetime(min_date).timestamp()
filtered = filtered[filtered['tstamp'] >= min_timestamp]
if max_date:
max_timestamp = pd.to_datetime(max_date).timestamp()
filtered = filtered[filtered['tstamp'] <= max_timestamp]
# Filter by language
if language and 'language' in filtered.columns:
filtered = filtered[filtered['language'] == language]
# Filter by turn
if 'turn' in filtered.columns:
filtered = filtered[filtered['turn'] >= min_turn]
pct = (len(filtered) / len(df) * 100) if len(df) else 0.0
print(f"After filtering: {len(filtered)} records ({pct:.1f}% of original)")
return filtered.reset_index(drop=True)
def get_time_slices(df: pd.DataFrame, interval: str = 'W') -> list:
"""
Split data into time slices for historical analysis.
Args:
df: Input DataFrame with 'tstamp' column
interval: Pandas frequency string ('D' for daily, 'W' for weekly, 'M' for monthly)
Returns:
List of (end_date, dataframe_slice) tuples
"""
if 'tstamp' not in df.columns:
raise ValueError("DataFrame must have 'tstamp' column")
if len(df) == 0:
print(f"Created 0 time slices with interval '{interval}'")
return []
df['datetime'] = pd.to_datetime(df['tstamp'], unit='s')
min_date = df['datetime'].min()
max_date = df['datetime'].max()
# Pandas 2.2+ removed 'M' (month-end); keep the documented monthly alias.
freq = "ME" if interval == "M" else interval
# Generate date ranges
date_ranges = pd.date_range(start=min_date, end=max_date, freq=freq)
slices = []
for end_date in date_ranges:
slice_df = df[df['datetime'] <= end_date].copy()
if len(slice_df) > 0:
slices.append((end_date, slice_df))
# Empty date_ranges when span < interval; also cover trailing gap to max_date.
if len(date_ranges) == 0 or date_ranges[-1] < max_date:
slices.append((max_date, df.copy()))
print(f"Created {len(slices)} time slices with interval '{interval}'")
return slices
+169
View File
@@ -0,0 +1,169 @@
"""
Elo Rating System Implementation
Based on Bradley-Terry model for pairwise comparison
"""
import numpy as np
from typing import Dict, Tuple, Optional
class EloRatingSystem:
"""
Implementation of Elo rating system for model comparison.
The Elo system updates ratings based on pairwise comparison outcomes,
where the rating difference between two models determines expected win probability.
"""
def __init__(self, initial_rating: float = 1000.0, k_factor: float = 4.0):
"""
Initialize Elo rating system.
Args:
initial_rating: Starting rating for all models
k_factor: Learning rate controlling magnitude of rating updates
"""
self.initial_rating = initial_rating
self.k_factor = k_factor
self.ratings: Dict[str, float] = {}
self.match_counts: Dict[str, int] = {}
self.win_counts: Dict[str, float] = {} # ties add 0.5, so this is float
def get_rating(self, model: str) -> float:
"""Get current rating for a model, initializing if necessary."""
if model not in self.ratings:
self.ratings[model] = self.initial_rating
self.match_counts[model] = 0
self.win_counts[model] = 0
return self.ratings[model]
def expected_score(self, rating_a: float, rating_b: float) -> float:
"""
Calculate expected win probability for model A against model B.
Uses logistic function: P(A wins) = 1 / (1 + 10^((R_B - R_A)/400))
Args:
rating_a: Rating of model A
rating_b: Rating of model B
Returns:
Expected probability that A wins (between 0 and 1)
"""
return 1.0 / (1.0 + 10.0 ** ((rating_b - rating_a) / 400.0))
def update_ratings(self, model_a: str, model_b: str, outcome: str) -> Tuple[float, float]:
"""
Update ratings after a match between two models.
Args:
model_a: Identifier for first model
model_b: Identifier for second model
outcome: Match result ('model_a', 'model_b', or 'tie')
Returns:
Tuple of (new_rating_a, new_rating_b)
"""
# Get current ratings
rating_a = self.get_rating(model_a)
rating_b = self.get_rating(model_b)
# Calculate expected scores
expected_a = self.expected_score(rating_a, rating_b)
expected_b = 1.0 - expected_a
# Determine actual scores
if outcome == 'model_a':
score_a, score_b = 1.0, 0.0
self.win_counts[model_a] = self.win_counts.get(model_a, 0) + 1
elif outcome == 'model_b':
score_a, score_b = 0.0, 1.0
self.win_counts[model_b] = self.win_counts.get(model_b, 0) + 1
else: # tie
score_a, score_b = 0.5, 0.5
# A tie counts as half a win for each side, keeping win_counts (and
# the win-rate derived from it) consistent with the 0.5-per-tie
# convention used elsewhere (leaderboard win-rate matrix, CLI stats).
self.win_counts[model_a] = self.win_counts.get(model_a, 0) + 0.5
self.win_counts[model_b] = self.win_counts.get(model_b, 0) + 0.5
# Update ratings using Elo formula
new_rating_a = rating_a + self.k_factor * (score_a - expected_a)
new_rating_b = rating_b + self.k_factor * (score_b - expected_b)
# Store updated ratings
self.ratings[model_a] = new_rating_a
self.ratings[model_b] = new_rating_b
# Update match counts
self.match_counts[model_a] = self.match_counts.get(model_a, 0) + 1
self.match_counts[model_b] = self.match_counts.get(model_b, 0) + 1
return new_rating_a, new_rating_b
def get_leaderboard(self) -> list:
"""
Get current leaderboard sorted by rating.
Returns:
List of tuples (model, rating, matches, wins) sorted by rating descending
"""
leaderboard = []
for model in self.ratings:
leaderboard.append((
model,
self.ratings[model],
self.match_counts.get(model, 0),
self.win_counts.get(model, 0)
))
# Sort by rating descending
leaderboard.sort(key=lambda x: x[1], reverse=True)
return leaderboard
def calculate_win_probability(self, model_a: str, model_b: str) -> float:
"""
Calculate win probability of model_a against model_b based on current ratings.
Args:
model_a: First model identifier
model_b: Second model identifier
Returns:
Probability that model_a wins (between 0 and 1)
"""
rating_a = self.get_rating(model_a)
rating_b = self.get_rating(model_b)
return self.expected_score(rating_a, rating_b)
def get_win_rate_matrix(self) -> Dict[Tuple[str, str], float]:
"""
Calculate pairwise win probability matrix for all models.
Returns:
Dictionary mapping (model_a, model_b) to win probability of model_a
"""
models = sorted(self.ratings.keys())
matrix = {}
for model_a in models:
for model_b in models:
if model_a != model_b:
prob = self.calculate_win_probability(model_a, model_b)
matrix[(model_a, model_b)] = prob
return matrix
def reset(self):
"""Reset all ratings to initial values."""
self.ratings.clear()
self.match_counts.clear()
self.win_counts.clear()
def copy(self) -> 'EloRatingSystem':
"""Create a deep copy of the current rating system."""
new_system = EloRatingSystem(self.initial_rating, self.k_factor)
new_system.ratings = self.ratings.copy()
new_system.match_counts = self.match_counts.copy()
new_system.win_counts = self.win_counts.copy()
return new_system
+24
View File
@@ -0,0 +1,24 @@
# 实验 7-7 环境变量示例
#
# 只有 `--source llm`LLM 评判对战)需要 API Key
# simulate / arena / elo / leaderboard / pipeline 均为纯离线,无需任何 Key。
#
# 用法:把本文件复制为 .env 并填入真实 Key,或直接 `export` 到 shell。
# --- 评判后端 1:官方 Anthropic(默认优先)---
# 有此 Key 时 --judge-backend auto 会走官方 Anthropic SDK。
ANTHROPIC_API_KEY=your-anthropic-api-key
# --- 评判后端 2OpenRouter 兜底 ---
# 当 ANTHROPIC_API_KEY 缺失或失效时使用;用 OpenAI 兼容 SDK 指向 OpenRouter。
# 内部 Claude 名字会自动映射为 OpenRouter id
# claude-opus-4-8 -> anthropic/claude-opus-4.8
# claude-haiku-4-5 -> anthropic/claude-haiku-4.5
# claude-sonnet-4-6 -> anthropic/claude-sonnet-4.6
# 已含 '/' 的 id(如 openai/gpt-5.6-luna)原样透传。
#
# 强制走 OpenRouter
# python cli.py battle --source llm --judge-backend openrouter \
# --judge-model claude-opus-4-8 \
# --candidate-models anthropic/claude-haiku-4.5 openai/gpt-5.6-luna
OPENROUTER_API_KEY=your-openrouter-api-key
+243
View File
@@ -0,0 +1,243 @@
"""
Leaderboard calculation and analysis
"""
import pandas as pd
import numpy as np
from typing import Dict, List, Tuple
from tqdm import tqdm
from elo_rating import EloRatingSystem
def build_leaderboard(df: pd.DataFrame,
initial_rating: float = 1000.0,
k_factor: float = 32.0,
show_progress: bool = True) -> EloRatingSystem:
"""
Build Elo leaderboard from voting data.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
initial_rating: Starting rating for all models
k_factor: Elo learning rate
show_progress: Whether to show progress bar
Returns:
EloRatingSystem with final ratings
"""
elo = EloRatingSystem(initial_rating=initial_rating, k_factor=k_factor)
iterator = tqdm(df.iterrows(), total=len(df), desc="Processing matches") if show_progress else df.iterrows()
for idx, row in iterator:
model_a = row['model_a']
model_b = row['model_b']
winner = row['winner']
elo.update_ratings(model_a, model_b, winner)
return elo
def calculate_win_rate_matrix_from_data(df: pd.DataFrame) -> pd.DataFrame:
"""
Calculate empirical win rate matrix directly from vote data.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
Returns:
DataFrame with win rates (rows beat columns)
"""
# Get all unique models
all_models = sorted(set(df['model_a'].unique()) | set(df['model_b'].unique()))
# Initialize counts
wins = {model: {opponent: 0 for opponent in all_models} for model in all_models}
total = {model: {opponent: 0 for opponent in all_models} for model in all_models}
# Count wins and totals
for _, row in df.iterrows():
model_a = row['model_a']
model_b = row['model_b']
winner = row['winner']
total[model_a][model_b] += 1
total[model_b][model_a] += 1
if winner == 'model_a':
wins[model_a][model_b] += 1
elif winner == 'model_b':
wins[model_b][model_a] += 1
else: # tie
wins[model_a][model_b] += 0.5
wins[model_b][model_a] += 0.5
# Calculate win rates
win_rates = {}
for model in all_models:
win_rates[model] = {}
for opponent in all_models:
if model == opponent:
win_rates[model][opponent] = 0.5
elif total[model][opponent] > 0:
win_rates[model][opponent] = wins[model][opponent] / total[model][opponent]
else:
win_rates[model][opponent] = np.nan
# Convert to DataFrame
win_rate_df = pd.DataFrame(win_rates).T
win_rate_df = win_rate_df[all_models] # Ensure consistent ordering
return win_rate_df
def compare_win_rates(elo_system: EloRatingSystem, empirical_df: pd.DataFrame) -> pd.DataFrame:
"""
Compare predicted win rates from Elo with empirical win rates.
Args:
elo_system: Trained Elo rating system
empirical_df: DataFrame with empirical win rates
Returns:
DataFrame with comparison statistics
"""
models = empirical_df.index.tolist()
comparisons = []
for model_a in models:
for model_b in models:
if model_a != model_b:
empirical = empirical_df.loc[model_a, model_b]
if not np.isnan(empirical):
predicted = elo_system.calculate_win_probability(model_a, model_b)
error = abs(predicted - empirical)
comparisons.append({
'model_a': model_a,
'model_b': model_b,
'empirical': empirical,
'predicted': predicted,
'error': error
})
cols = ['model_a', 'model_b', 'empirical', 'predicted', 'error']
if not comparisons:
return pd.DataFrame(columns=cols)
comparison_df = pd.DataFrame(comparisons, columns=cols)
return comparison_df
def build_historical_leaderboards(df: pd.DataFrame,
time_slices: List[Tuple],
initial_rating: float = 1000.0,
k_factor: float = 32.0) -> List[Tuple]:
"""
Build leaderboard snapshots at different time points.
Args:
df: Full voting DataFrame
time_slices: List of (end_date, slice_df) tuples from get_time_slices
initial_rating: Starting rating
k_factor: Elo learning rate
Returns:
List of (date, leaderboard_data) tuples
"""
historical_leaderboards = []
for end_date, slice_df in tqdm(time_slices, desc="Building historical leaderboards"):
elo = build_leaderboard(slice_df, initial_rating, k_factor, show_progress=False)
leaderboard = elo.get_leaderboard()
# Convert to DataFrame for easier handling
lb_df = pd.DataFrame(leaderboard, columns=['model', 'rating', 'matches', 'wins'])
lb_df['date'] = end_date
lb_df['rank'] = range(1, len(lb_df) + 1)
historical_leaderboards.append((end_date, lb_df))
return historical_leaderboards
def get_rating_history(historical_leaderboards: List[Tuple]) -> pd.DataFrame:
"""
Extract rating history for all models over time.
Args:
historical_leaderboards: List of (date, leaderboard_df) tuples
Returns:
DataFrame with columns: date, model, rating, rank
"""
all_data = []
for date, lb_df in historical_leaderboards:
for _, row in lb_df.iterrows():
all_data.append({
'date': date,
'model': row['model'],
'rating': row['rating'],
'rank': row['rank'],
'matches': row['matches'],
'wins': row['wins']
})
cols = ['date', 'model', 'rating', 'rank', 'matches', 'wins']
if not all_data:
return pd.DataFrame(columns=cols)
history_df = pd.DataFrame(all_data)
return history_df
def analyze_rating_changes(history_df: pd.DataFrame, top_n: int = 20) -> pd.DataFrame:
"""
Analyze rating changes over time for top models.
Args:
history_df: DataFrame from get_rating_history
top_n: Number of top models to analyze
Returns:
DataFrame with change statistics
"""
stats_cols = [
'model', 'final_rating', 'initial_rating', 'rating_change',
'max_rating', 'min_rating', 'volatility', 'total_matches',
]
if history_df is None or len(history_df) == 0:
return pd.DataFrame(columns=stats_cols)
# Get final ratings
final_date = history_df['date'].max()
final_ratings = history_df[history_df['date'] == final_date].nlargest(top_n, 'rating')
top_models = final_ratings['model'].tolist()
# Calculate statistics for each model
stats = []
for model in top_models:
model_data = history_df[history_df['model'] == model].sort_values('date')
if len(model_data) > 0:
initial_rating = model_data.iloc[0]['rating']
final_rating = model_data.iloc[-1]['rating']
max_rating = model_data['rating'].max()
min_rating = model_data['rating'].min()
rating_change = final_rating - initial_rating
volatility = model_data['rating'].std()
stats.append({
'model': model,
'final_rating': final_rating,
'initial_rating': initial_rating,
'rating_change': rating_change,
'max_rating': max_rating,
'min_rating': min_rating,
'volatility': volatility,
'total_matches': model_data.iloc[-1]['matches']
})
if not stats:
return pd.DataFrame(columns=stats_cols)
stats_df = pd.DataFrame(stats).sort_values('final_rating', ascending=False)
return stats_df
+271
View File
@@ -0,0 +1,271 @@
"""
LLM-as-judge pairwise battles with position-bias mitigation.
This is the only battle source that needs network access; `simulate` and
`arena` run fully offline.
Two backends are supported, selected automatically or via ``backend=``:
* ``anthropic`` the official ``anthropic`` SDK, using ``ANTHROPIC_API_KEY``
(the default when that key is present).
* ``openrouter`` the OpenAI-compatible ``openai`` SDK pointed at
``https://openrouter.ai/api/v1`` with ``OPENROUTER_API_KEY``. Internal
Claude ids (e.g. ``claude-opus-4-8``) are mapped to their OpenRouter ids
(``anthropic/claude-opus-4.8``); ids that already contain a ``/`` such as
``openai/gpt-5.6-luna`` are passed through untouched. This lets the judge run
when a direct Anthropic key is missing or invalid.
The two backends are interchangeable: the position-bias swap-and-agree logic and
the A/B/tie response parsing are identical regardless of which one is used.
The book (实验 7-7, 位置偏差 discussion) notes that an LLM judge systematically
favours whichever answer appears in a fixed slot (usually the first). The
standard mitigation, implemented here, is to judge each pair twice with the
answers swapped and only record a winner when both judgements agree; a
disagreement is counted as a tie. This cancels the position bias instead of
letting it leak into the ratings.
The resulting battle list uses the same {'model_a', 'model_b', 'winner'} schema
as the simulated and Chatbot Arena data, so it feeds straight into the Elo /
Bradley-Terry pipeline.
"""
import os
from typing import Dict, List, Optional
try:
from dotenv import load_dotenv
load_dotenv()
except ImportError:
pass
# Default candidate roster and judge (Claude models). Kept small because every
# battle costs several API calls (two responses + two swapped judgements).
DEFAULT_CANDIDATE_MODELS = ["claude-opus-4-8", "claude-haiku-4-5"]
DEFAULT_JUDGE_MODEL = "claude-opus-4-8"
DEFAULT_PROMPTS = [
"用一句话解释什么是 Transformer 的自注意力机制。",
"Write a haiku about distributed systems.",
"给出快速排序的时间复杂度,并简要说明最坏情况。",
]
OPENROUTER_BASE_URL = "https://openrouter.ai/api/v1"
# Map internal Claude ids -> OpenRouter model ids. Any id already containing a
# '/' (e.g. 'openai/gpt-5.6-luna') is treated as a native OpenRouter id and used
# verbatim; unknown ids are also passed through unchanged.
_OPENROUTER_MODEL_MAP = {
"claude-opus-4-8": "anthropic/claude-opus-4.8",
"claude-opus-4-1": "anthropic/claude-opus-4.1",
"claude-sonnet-4-6": "anthropic/claude-sonnet-4.6",
"claude-sonnet-4-5": "anthropic/claude-sonnet-4.5",
"claude-haiku-4-5": "anthropic/claude-haiku-4.5",
}
_JUDGE_SYSTEM = (
"你是一个严格的评委。用户会给你一个问题和两个候选回答(回答 A 和回答 B)。"
"请只根据回答质量判断哪个更好,忽略它们出现的顺序。"
"只输出一个词:A、B 或 tie。"
)
def _to_openrouter_model(model: str) -> str:
"""Translate an internal model id into an OpenRouter model id."""
if "/" in model: # already a native OpenRouter id
return model
return _OPENROUTER_MODEL_MAP.get(model, model)
class JudgeClient:
"""
Thin adapter over either the Anthropic SDK or the OpenAI-compatible
OpenRouter endpoint, exposing a single ``chat()`` method so the rest of the
module is backend-agnostic.
"""
def __init__(self, backend: str, impl):
self.backend = backend
self.impl = impl
def chat(self, model: str, user: str, max_tokens: int,
system: Optional[str] = None) -> str:
"""Send a single-turn chat and return the assistant's text reply."""
if self.backend == "anthropic":
kwargs = {
"model": model,
"max_tokens": max_tokens,
"messages": [{"role": "user", "content": user}],
}
if system is not None:
kwargs["system"] = system
response = self.impl.messages.create(**kwargs)
return "".join(
block.text for block in response.content if block.type == "text"
).strip()
# openrouter (OpenAI-compatible chat.completions)
messages = []
if system is not None:
messages.append({"role": "system", "content": system})
messages.append({"role": "user", "content": user})
response = self.impl.chat.completions.create(
model=_to_openrouter_model(model),
max_tokens=max_tokens,
messages=messages,
)
return (response.choices[0].message.content or "").strip()
def _resolve_backend(backend: str = "auto") -> str:
"""
Resolve the effective backend.
``auto`` -> ``anthropic`` if ANTHROPIC_API_KEY is set, else ``openrouter``
if OPENROUTER_API_KEY is set. Raises if neither key is available.
"""
if backend not in ("anthropic", "openrouter", "auto"):
raise ValueError(
f"Unknown judge backend {backend!r}; expected 'anthropic', "
"'openrouter' or 'auto'."
)
if backend != "auto":
return backend
if os.environ.get("ANTHROPIC_API_KEY"):
return "anthropic"
if os.environ.get("OPENROUTER_API_KEY"):
return "openrouter"
raise RuntimeError(
"No LLM-judge credentials found. Set ANTHROPIC_API_KEY (direct Anthropic) "
"or OPENROUTER_API_KEY (OpenRouter fallback); or use --source simulate / "
"--source arena to run the experiment fully offline."
)
def _get_client(backend: str = "auto") -> JudgeClient:
"""Create a JudgeClient for the resolved backend, with clear errors."""
backend = _resolve_backend(backend)
if backend == "anthropic":
try:
import anthropic
except ImportError as exc: # pragma: no cover - depends on environment
raise RuntimeError(
"The 'anthropic' package is required for the anthropic judge "
"backend. Install it with: pip install anthropic"
) from exc
if not os.environ.get("ANTHROPIC_API_KEY"):
raise RuntimeError(
"ANTHROPIC_API_KEY is not set. Set it, or use "
"--judge-backend openrouter with OPENROUTER_API_KEY, or run "
"--source simulate / --source arena fully offline."
)
return JudgeClient("anthropic", anthropic.Anthropic())
# backend == "openrouter"
try:
import openai
except ImportError as exc: # pragma: no cover - depends on environment
raise RuntimeError(
"The 'openai' package is required for the openrouter judge backend. "
"Install it with: pip install openai"
) from exc
if not os.environ.get("OPENROUTER_API_KEY"):
raise RuntimeError(
"OPENROUTER_API_KEY is not set. Set it, or use --judge-backend "
"anthropic with ANTHROPIC_API_KEY, or run --source simulate / "
"--source arena fully offline."
)
return JudgeClient(
"openrouter",
openai.OpenAI(
base_url=OPENROUTER_BASE_URL,
api_key=os.environ["OPENROUTER_API_KEY"],
),
)
def generate_response(client: JudgeClient, model: str, prompt: str,
max_tokens: int = 1024) -> str:
"""Generate a single model answer for a prompt."""
return client.chat(model, prompt, max_tokens=max_tokens)
def _judge_once(client: JudgeClient, judge_model: str, prompt: str,
answer_first: str, answer_second: str) -> str:
"""Ask the judge which slot is better; returns 'first', 'second' or 'tie'."""
user = (
f"问题:\n{prompt}\n\n"
f"回答 A\n{answer_first}\n\n"
f"回答 B\n{answer_second}\n\n"
"哪个回答更好?只输出 A、B 或 tie。"
)
verdict = client.chat(judge_model, user, max_tokens=8, system=_JUDGE_SYSTEM).lower()
if verdict.startswith("a"):
return "first"
if verdict.startswith("b"):
return "second"
return "tie"
def judge_pair(client: JudgeClient, judge_model: str, prompt: str,
answer_a: str, answer_b: str) -> str:
"""
Judge a pair with position-bias mitigation (swap order, tie on disagreement).
Returns 'model_a', 'model_b', or 'tie'.
"""
# First pass: A in slot 1, B in slot 2.
first_pass = _judge_once(client, judge_model, prompt, answer_a, answer_b)
# Second pass: swap the slots so B is now in slot 1.
second_pass = _judge_once(client, judge_model, prompt, answer_b, answer_a)
# Translate both judgements into "which real model won", then require
# agreement. Slot 1 in the first pass is A; slot 1 in the second pass is B.
winner_first = {"first": "model_a", "second": "model_b", "tie": "tie"}[first_pass]
winner_second = {"first": "model_b", "second": "model_a", "tie": "tie"}[second_pass]
if winner_first == winner_second:
return winner_first
return "tie" # inconsistent under swap -> position bias, count as tie
def run_llm_battles(candidate_models: Optional[List[str]] = None,
prompts: Optional[List[str]] = None,
judge_model: str = DEFAULT_JUDGE_MODEL,
backend: str = "auto") -> List[dict]:
"""
Run LLM-judged battles between every model pair over every prompt.
Args:
candidate_models: Models to compare (default: DEFAULT_CANDIDATE_MODELS).
prompts: Prompts to battle on (default: DEFAULT_PROMPTS).
judge_model: Model used as the judge.
backend: 'anthropic', 'openrouter', or 'auto' (anthropic if
ANTHROPIC_API_KEY else openrouter).
Returns:
List of battle dicts ({'model_a', 'model_b', 'winner'}).
"""
candidate_models = candidate_models or DEFAULT_CANDIDATE_MODELS
prompts = prompts or DEFAULT_PROMPTS
if len(candidate_models) < 2:
raise ValueError("Need at least 2 candidate models for LLM-judge battles")
client = _get_client(backend)
battles: List[dict] = []
for prompt in prompts:
# Cache each model's answer per prompt so it is generated only once.
answers: Dict[str, str] = {
model: generate_response(client, model, prompt) for model in candidate_models
}
for i, model_a in enumerate(candidate_models):
for model_b in candidate_models[i + 1:]:
winner = judge_pair(
client, judge_model, prompt, answers[model_a], answers[model_b]
)
battles.append(
{"model_a": model_a, "model_b": model_b, "winner": winner}
)
return battles
+242
View File
@@ -0,0 +1,242 @@
"""
Main script for Model Leaderboard Calculation
Experiment 7-7: Building Model Leaderboard from Pairwise Comparison Data
Supports two methods (following official Chatbot Arena):
1. Online Elo (K=4) - Simple but order-dependent
2. Bradley-Terry MLE - Official leaderboard method (more stable)
"""
import os
import sys
import pandas as pd
from bradley_terry import (
compute_bradley_terry_leaderboard,
predict_win_rate,
)
from data_loader import download_arena_data, filter_data, load_arena_data
from elo_rating import EloRatingSystem
from parallel_processing import optimize_dataframe
from visualization import plot_leaderboard, plot_rating_distribution, plot_win_rate_matrix
def compute_online_elo_leaderboard(df: pd.DataFrame) -> pd.DataFrame:
"""
Compute leaderboard using online Elo updates (K=4, official value).
This method updates ratings sequentially as matches are processed.
It's simpler but can be unstable and order-dependent.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
Returns:
DataFrame with model ratings
"""
from tqdm import tqdm
print("Computing online Elo ratings (K=4)...")
elo = EloRatingSystem(initial_rating=1000.0, k_factor=4.0)
# Process matches sequentially
for _, row in tqdm(df.iterrows(), total=len(df), desc="Processing matches"):
elo.update_ratings(row['model_a'], row['model_b'], row['winner'])
# Get leaderboard
leaderboard = elo.get_leaderboard()
# Convert to DataFrame
result = pd.DataFrame(leaderboard, columns=['model', 'rating', 'matches', 'wins'])
return result
def main(method: str = 'bradley-terry'):
"""
Run model leaderboard calculation.
Args:
method: 'bradley-terry' (default, official) or 'online-elo' (simple)
"""
print("="*80)
print("Experiment 7-7: Building Model Leaderboard from Pairwise Comparisons")
if method == 'bradley-terry':
print("Method: Bradley-Terry Model with MLE (Official Chatbot Arena)")
else:
print("Method: Online Elo Updates (K=4)")
print("="*80)
print()
# Step 1: Download and load data
print("Step 1: Loading Chatbot Arena voting data...")
print("-" * 80)
data_file = "arena_data.json"
try:
# Download data if not exists
if not os.path.exists(data_file):
data_file = download_arena_data(data_file)
# Load data
df = load_arena_data(data_file)
# Optimize memory usage
df = optimize_dataframe(df)
except Exception as e: # noqa: BLE001 - surface loader/provider diagnostics to CLI users
print(f"Error loading data: {e}")
print("\nNote: If the data download fails, you can manually download the file from:")
print("https://storage.googleapis.com/arena_external_data/public/clean_battle_20240814_public.json")
print("and save it as 'arena_data.json' in the current directory.")
return
print()
# Step 2: Filter data (official Chatbot Arena method)
print("Step 2: Filtering data (following official Arena method)...")
print("-" * 80)
df_filtered = filter_data(
df,
anony_only=True, # Only anonymous/blind votes
use_dedup=True, # Apply deduplication (removes top 0.1% redundant prompts)
min_turn=1
)
print()
# Step 3: Compute ratings using selected method
print(f"Step 3: Computing ratings using {method} method...")
print("-" * 80)
if method == 'bradley-terry':
print("Note: Bradley-Terry model uses sklearn LogisticRegression for MLE.")
print("This is the official Chatbot Arena method - more stable than online Elo.")
print()
# Compute ratings with bootstrap for confidence intervals
leaderboard_df = compute_bradley_terry_leaderboard(df_filtered, bootstrap_rounds=100)
print("\nTop 20 models by Bradley-Terry rating:")
print("-" * 80)
print(f"{'Rank':<6}{'Model':<35}{'Rating':<10}{'95% CI':<20}")
print("-" * 80)
for idx, row in leaderboard_df.head(20).iterrows():
if 'lower_ci' in row and 'upper_ci' in row:
ci_str = f"[{row['lower_ci']:.1f}, {row['upper_ci']:.1f}]"
else:
ci_str = "N/A"
print(f"{idx+1:<6}{row['model']:<35}{row['rating']:7.1f} {ci_str:<20}")
else: # online-elo
print("Note: Online Elo uses K=4 (official value) for stable ratings.")
print("Processes matches sequentially - simpler but can be order-dependent.")
print()
# Compute online Elo ratings
leaderboard_df = compute_online_elo_leaderboard(df_filtered)
print("\nTop 20 models by Online Elo rating:")
print("-" * 80)
print(f"{'Rank':<6}{'Model':<35}{'Rating':<10}{'Matches':<10}{'Win Rate':<10}")
print("-" * 80)
for idx, row in leaderboard_df.head(20).iterrows():
win_rate = row['wins'] / row['matches'] * 100 if row['matches'] > 0 else 0
print(f"{idx+1:<6}{row['model']:<35}{row['rating']:7.1f} {row['matches']:<10}{win_rate:6.1f}%")
print()
# Step 4: Predict win rates using Bradley-Terry model
print("Step 4: Calculating predicted win rates...")
print("-" * 80)
# Get ratings as dictionary
ratings_dict = dict(zip(leaderboard_df['model'], leaderboard_df['rating']))
# Predict win rates
predicted_win_rates = predict_win_rate(ratings_dict)
print(f"Calculated predicted win rates for {len(ratings_dict)} models")
print()
# Step 5: Create visualizations
print("Step 5: Creating visualizations...")
print("-" * 80)
# Convert leaderboard_df to format expected by visualization functions
leaderboard_tuples = [(row['model'], row['rating'], 0, 0) for _, row in leaderboard_df.iterrows()]
plot_leaderboard(leaderboard_tuples, top_n=20, save_path="leaderboard.png")
plot_rating_distribution(leaderboard_tuples, save_path="rating_distribution.png")
# Plot win rate matrix
top_30_models = leaderboard_df.head(30)['model'].tolist()
plot_win_rate_matrix(predicted_win_rates.loc[top_30_models, top_30_models],
top_n=30, save_path="win_rate_matrix.png")
print()
# Summary
print("="*80)
print("Analysis complete!")
print("="*80)
print("\nGenerated files:")
print(" - leaderboard.png : Top 20 models by Bradley-Terry rating")
print(" - rating_distribution.png : Distribution of ratings")
print(" - win_rate_matrix.png : Predicted win rate matrix (top 30 models)")
print()
# Method summary
print("Method:")
print("-" * 80)
if method == 'bradley-terry':
print(" ✓ Bradley-Terry model with Maximum Likelihood Estimation")
print(" ✓ sklearn LogisticRegression for stable rating computation")
print(" ✓ Bootstrap confidence intervals (100 samples)")
else:
print(" ✓ Online Elo with K=4 (official value)")
print(" ✓ Sequential match processing")
print(" ✓ Simple but order-dependent")
print(" ✓ Deduplication filter (removes top 0.1% redundant prompts)")
print(" ✓ Anonymous votes only (blind evaluation)")
print()
# Key insights
print("Key Insights:")
print("-" * 80)
print(f" • Total models analyzed: {len(leaderboard_df)}")
print(f" • Total battles: {len(df_filtered):,}")
print(f" • Rating range: {leaderboard_df['rating'].min():.1f} - {leaderboard_df['rating'].max():.1f}")
print(f" • Top model: {leaderboard_df.iloc[0]['model']} ({leaderboard_df.iloc[0]['rating']:.1f})")
if 'lower_ci' in leaderboard_df.columns:
avg_ci_width = (leaderboard_df['upper_ci'] - leaderboard_df['lower_ci']).mean()
print(f" • Average confidence interval width: {avg_ci_width:.1f} rating points")
print()
print("This implementation matches the official Chatbot Arena leaderboard calculation!")
print("Source: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH")
print()
if __name__ == "__main__":
# Allow method selection via command line argument
import sys
method = 'bradley-terry' # Default to official method
if len(sys.argv) > 1:
if sys.argv[1] in ['bradley-terry', 'bt', 'mle']:
method = 'bradley-terry'
elif sys.argv[1] in ['online-elo', 'elo', 'online']:
method = 'online-elo'
else:
print(f"Unknown method: {sys.argv[1]}")
print("Usage: python main.py [bradley-terry|online-elo]")
print(" bradley-terry (default): Official Arena method, more stable")
print(" online-elo: Simple Elo updates with K=4")
sys.exit(1)
main(method)
+303
View File
@@ -0,0 +1,303 @@
"""
Optimized Elo rating system using NumPy vectorization and Numba JIT
"""
import numpy as np
import pandas as pd
from typing import Dict, Tuple, List
from numba import jit
from tqdm import tqdm
@jit(nopython=True)
def expected_score_fast(rating_a: float, rating_b: float) -> float:
"""
Fast expected score calculation using Numba JIT.
Args:
rating_a: Rating of model A
rating_b: Rating of model B
Returns:
Expected probability that A wins
"""
return 1.0 / (1.0 + 10.0 ** ((rating_b - rating_a) / 400.0))
@jit(nopython=True)
def process_elo_updates_vectorized(ratings: np.ndarray,
model_a_indices: np.ndarray,
model_b_indices: np.ndarray,
outcomes: np.ndarray,
k_factor: float,
match_counts: np.ndarray,
win_counts: np.ndarray) -> np.ndarray:
"""
Process Elo updates using vectorized NumPy operations with Numba JIT.
This is the core hot loop optimized with Numba for maximum performance.
Args:
ratings: Array of current ratings for all models
model_a_indices: Indices of model A for each match
model_b_indices: Indices of model B for each match
outcomes: Match outcomes (1.0 = A wins, 0.0 = B wins, 0.5 = tie)
k_factor: Elo K-factor
match_counts: Array to track match counts per model
win_counts: Array to track win counts per model
Returns:
Updated ratings array
"""
n_matches = len(model_a_indices)
for i in range(n_matches):
idx_a = model_a_indices[i]
idx_b = model_b_indices[i]
outcome = outcomes[i]
# Get current ratings
rating_a = ratings[idx_a]
rating_b = ratings[idx_b]
# Calculate expected scores
expected_a = 1.0 / (1.0 + 10.0 ** ((rating_b - rating_a) / 400.0))
expected_b = 1.0 - expected_a
# Update ratings
ratings[idx_a] += k_factor * (outcome - expected_a)
ratings[idx_b] += k_factor * ((1.0 - outcome) - expected_b)
# Update counts
match_counts[idx_a] += 1
match_counts[idx_b] += 1
win_counts[idx_a] += outcome
win_counts[idx_b] += (1.0 - outcome)
return ratings
@jit(nopython=True)
def calculate_expected_scores_vectorized(ratings_a: np.ndarray,
ratings_b: np.ndarray) -> np.ndarray:
"""
Vectorized calculation of expected scores for multiple matches.
Args:
ratings_a: Array of ratings for model A
ratings_b: Array of ratings for model B
Returns:
Array of expected scores for model A
"""
return 1.0 / (1.0 + np.power(10.0, (ratings_b - ratings_a) / 400.0))
class NumpyEloRatingSystem:
"""
Highly optimized Elo rating system using NumPy arrays and Numba JIT.
Optimizations:
- NumPy arrays for O(1) indexing instead of dictionary lookups
- Numba JIT compilation of hot loops
- Pre-allocated arrays to avoid memory reallocation
- Integer indexing for models instead of string lookups
"""
def __init__(self, initial_rating: float = 1000.0, k_factor: float = 4.0):
"""Initialize NumPy-based Elo system."""
self.initial_rating = initial_rating
self.k_factor = k_factor
# Model name to index mapping
self.model_to_idx: Dict[str, int] = {}
self.idx_to_model: Dict[int, str] = {}
# NumPy arrays for fast access
self.ratings: np.ndarray = None
self.match_counts: np.ndarray = None
self.win_counts: np.ndarray = None
self.n_models = 0
def _prepare_data(self, df: pd.DataFrame):
"""
Prepare NumPy arrays from DataFrame for fast processing.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
"""
print("Preparing data structures...")
# Get all unique models
all_models = sorted(set(df['model_a'].unique()) | set(df['model_b'].unique()))
self.n_models = len(all_models)
print(f"Found {self.n_models} unique models")
# Create model mappings
for idx, model in enumerate(all_models):
self.model_to_idx[model] = idx
self.idx_to_model[idx] = model
# Initialize arrays
self.ratings = np.full(self.n_models, self.initial_rating, dtype=np.float64)
self.match_counts = np.zeros(self.n_models, dtype=np.int32)
self.win_counts = np.zeros(self.n_models, dtype=np.float64)
# Convert DataFrame columns to NumPy arrays with integer indices
print("Converting model names to indices...")
model_a_indices = df['model_a'].map(self.model_to_idx).values.astype(np.int32)
model_b_indices = df['model_b'].map(self.model_to_idx).values.astype(np.int32)
# Convert outcomes to numeric (1.0 for A wins, 0.0 for B wins, 0.5 for tie)
print("Converting outcomes to numeric...")
# 'tie (bothbad)' is a real Arena outcome; unmapped values become NaN and
# would silently poison every rating they touch, so fall back to a tie
# (same as EloRatingSystem.update_ratings).
outcome_map = {'model_a': 1.0, 'model_b': 0.0,
'tie': 0.5, 'tie (bothbad)': 0.5}
outcomes = df['winner'].map(outcome_map).fillna(0.5).values.astype(np.float64)
return model_a_indices, model_b_indices, outcomes
def process_matches_vectorized(self, df: pd.DataFrame, show_progress: bool = True):
"""
Process all matches using vectorized NumPy operations and Numba JIT.
This is the fastest way to compute Elo ratings for large datasets.
Args:
df: DataFrame with columns 'model_a', 'model_b', 'winner'
show_progress: Whether to show progress bar
"""
# Prepare data
model_a_indices, model_b_indices, outcomes = self._prepare_data(df)
print(f"\nProcessing {len(df)} matches with NumPy + Numba JIT...")
# Process all matches using JIT-compiled function
# This is where the magic happens - Numba compiles this to machine code
if show_progress:
# Process in chunks to show progress
chunk_size = 50000
n_chunks = (len(model_a_indices) + chunk_size - 1) // chunk_size
for i in tqdm(range(n_chunks), desc="Processing matches"):
start_idx = i * chunk_size
end_idx = min((i + 1) * chunk_size, len(model_a_indices))
self.ratings = process_elo_updates_vectorized(
self.ratings,
model_a_indices[start_idx:end_idx],
model_b_indices[start_idx:end_idx],
outcomes[start_idx:end_idx],
self.k_factor,
self.match_counts,
self.win_counts
)
else:
self.ratings = process_elo_updates_vectorized(
self.ratings,
model_a_indices,
model_b_indices,
outcomes,
self.k_factor,
self.match_counts,
self.win_counts
)
print("✓ Processing complete!")
def get_leaderboard(self) -> List[Tuple]:
"""
Get sorted leaderboard using NumPy's fast sorting.
Returns:
List of tuples (model, rating, matches, wins)
"""
# Use NumPy's argsort for fast sorting
sorted_indices = np.argsort(-self.ratings) # Negative for descending order
leaderboard = []
for idx in sorted_indices:
model = self.idx_to_model[idx]
rating = float(self.ratings[idx])
matches = int(self.match_counts[idx])
wins = float(self.win_counts[idx])
leaderboard.append((model, rating, matches, wins))
return leaderboard
def calculate_win_probability(self, model_a: str, model_b: str) -> float:
"""
Calculate win probability using fast NumPy operations.
Args:
model_a: First model identifier
model_b: Second model identifier
Returns:
Probability that model_a wins
"""
if model_a not in self.model_to_idx or model_b not in self.model_to_idx:
return 0.5
idx_a = self.model_to_idx[model_a]
idx_b = self.model_to_idx[model_b]
rating_a = self.ratings[idx_a]
rating_b = self.ratings[idx_b]
return expected_score_fast(rating_a, rating_b)
def get_win_rate_matrix(self) -> Dict[Tuple[str, str], float]:
"""
Calculate pairwise win probability matrix using vectorized operations.
Returns:
Dictionary mapping (model_a, model_b) to win probability
"""
matrix = {}
# Vectorized calculation for all pairs
for i in range(self.n_models):
model_a = self.idx_to_model[i]
# Calculate win probabilities against all other models at once
ratings_a = np.full(self.n_models, self.ratings[i])
win_probs = calculate_expected_scores_vectorized(ratings_a, self.ratings)
for j in range(self.n_models):
if i != j:
model_b = self.idx_to_model[j]
matrix[(model_a, model_b)] = float(win_probs[j])
return matrix
def build_leaderboard_optimized(df: pd.DataFrame,
initial_rating: float = 1000.0,
k_factor: float = 4.0,
show_progress: bool = True) -> NumpyEloRatingSystem:
"""
Build Elo leaderboard using highly optimized NumPy + Numba algorithm.
This implementation is significantly faster than the basic version:
- Uses NumPy arrays for O(1) indexing
- Numba JIT compilation for hot loops
- Pre-allocated arrays to avoid memory overhead
- Integer-based model indexing instead of string lookups
Args:
df: DataFrame with match data (columns: model_a, model_b, winner)
initial_rating: Starting rating for all models
k_factor: Elo learning rate (K-factor)
show_progress: Whether to display progress bar
Returns:
NumpyEloRatingSystem with final ratings
"""
elo = NumpyEloRatingSystem(initial_rating=initial_rating, k_factor=k_factor)
elo.process_matches_vectorized(df, show_progress=show_progress)
return elo
@@ -0,0 +1,322 @@
"""
Parallel processing utilities for Elo rating computation
"""
import pandas as pd
import numpy as np
from multiprocessing import Pool, cpu_count
from functools import partial
from typing import List, Tuple
from tqdm import tqdm
from elo_rating import EloRatingSystem
def process_time_slice(args: Tuple) -> Tuple:
"""
Process a single time slice to build leaderboard.
Args:
args: Tuple of (end_date, slice_df, initial_rating, k_factor)
Returns:
Tuple of (end_date, leaderboard_data)
"""
end_date, slice_df, initial_rating, k_factor = args
# Build Elo system for this time slice
elo = EloRatingSystem(initial_rating=initial_rating, k_factor=k_factor)
# Process all matches in this slice
for _, row in slice_df.iterrows():
elo.update_ratings(row['model_a'], row['model_b'], row['winner'])
# Get leaderboard
leaderboard = elo.get_leaderboard()
# Convert to list of dicts for easier handling
lb_data = []
for rank, (model, rating, matches, wins) in enumerate(leaderboard, 1):
lb_data.append({
'model': model,
'rating': rating,
'matches': matches,
'wins': wins,
'rank': rank,
'date': end_date
})
return (end_date, lb_data)
def build_historical_leaderboards_parallel(df: pd.DataFrame,
time_slices: List[Tuple],
initial_rating: float = 1000.0,
k_factor: float = 32.0,
n_jobs: int = -1) -> List[Tuple]:
"""
Build historical leaderboards using parallel processing.
Args:
df: Full voting DataFrame
time_slices: List of (end_date, slice_df) tuples
initial_rating: Starting rating
k_factor: Elo learning rate
n_jobs: Number of parallel jobs (-1 for all cores)
Returns:
List of (date, leaderboard_data) tuples
"""
if n_jobs == -1:
n_jobs = cpu_count()
print(f"Building historical leaderboards using {n_jobs} cores...")
# Prepare arguments for parallel processing
args_list = [
(end_date, slice_df, initial_rating, k_factor)
for end_date, slice_df in time_slices
]
# Process in parallel
with Pool(processes=n_jobs) as pool:
results = list(tqdm(
pool.imap(process_time_slice, args_list),
total=len(args_list),
desc="Processing time slices"
))
# Convert results to expected format
historical_leaderboards = []
for end_date, lb_data in results:
lb_df = pd.DataFrame(lb_data)
historical_leaderboards.append((end_date, lb_df))
# Sort by date
historical_leaderboards.sort(key=lambda x: x[0])
return historical_leaderboards
def calculate_pairwise_win_rates_chunk(args: Tuple) -> List[dict]:
"""
Calculate win rates for a chunk of model pairs.
Args:
args: Tuple of (model_pairs, df)
Returns:
List of win rate dictionaries
"""
model_pairs, df = args
results = []
for model_a, model_b in model_pairs:
# Filter matches between these two models
matches = df[
((df['model_a'] == model_a) & (df['model_b'] == model_b)) |
((df['model_a'] == model_b) & (df['model_b'] == model_a))
]
if len(matches) == 0:
continue
wins_a = 0
total = len(matches)
for _, row in matches.iterrows():
# Arena data has four winner values; any non-win outcome
# ('tie' and 'tie (bothbad)') is worth 0.5, matching the
# serial calculate_win_rate_matrix_from_data.
if row['model_a'] == model_a:
if row['winner'] == 'model_a':
wins_a += 1
elif row['winner'] != 'model_b':
wins_a += 0.5
else: # model_a is model_b in the row
if row['winner'] == 'model_b':
wins_a += 1
elif row['winner'] != 'model_a':
wins_a += 0.5
win_rate = wins_a / total if total > 0 else 0.5
results.append({
'model_a': model_a,
'model_b': model_b,
'win_rate': win_rate,
'total_matches': total
})
return results
def calculate_win_rate_matrix_parallel(df: pd.DataFrame,
models: List[str] = None,
n_jobs: int = -1) -> pd.DataFrame:
"""
Calculate win rate matrix using parallel processing.
Args:
df: DataFrame with match data
models: List of models to include (if None, use all)
n_jobs: Number of parallel jobs
Returns:
DataFrame with win rates
"""
if n_jobs == -1:
n_jobs = cpu_count()
if models is None:
models = sorted(set(df['model_a'].unique()) | set(df['model_b'].unique()))
print(f"Calculating win rate matrix for {len(models)} models using {n_jobs} cores...")
# Generate all model pairs
model_pairs = [(m1, m2) for i, m1 in enumerate(models) for m2 in models[i+1:]]
# Split pairs into chunks for parallel processing
chunk_size = max(1, len(model_pairs) // (n_jobs * 4))
chunks = [model_pairs[i:i+chunk_size] for i in range(0, len(model_pairs), chunk_size)]
# Prepare arguments
args_list = [(chunk, df) for chunk in chunks]
# Process in parallel
with Pool(processes=n_jobs) as pool:
results_chunks = list(tqdm(
pool.imap(calculate_pairwise_win_rates_chunk, args_list),
total=len(args_list),
desc="Calculating win rates"
))
# Flatten results
all_results = [item for chunk in results_chunks for item in chunk]
# Build matrix. Pairs with no data stay NaN (the serial version's
# convention) — 0.5 would misreport "no data" as an even record;
# the diagonal is 0.5 by definition.
win_rates = {model: {opponent: (0.5 if opponent == model else np.nan)
for opponent in models} for model in models}
for result in all_results:
model_a = result['model_a']
model_b = result['model_b']
win_rate = result['win_rate']
win_rates[model_a][model_b] = win_rate
win_rates[model_b][model_a] = 1.0 - win_rate
# Convert to DataFrame
win_rate_df = pd.DataFrame(win_rates).T
win_rate_df = win_rate_df[models]
return win_rate_df
def filter_data_parallel(df: pd.DataFrame,
filters: dict,
n_jobs: int = -1) -> pd.DataFrame:
"""
Filter large DataFrame using parallel processing.
Args:
df: Input DataFrame
filters: Dictionary of filter conditions
n_jobs: Number of parallel jobs
Returns:
Filtered DataFrame
"""
if n_jobs == -1:
n_jobs = min(cpu_count(), 4) # Cap at 4 for filtering
if len(df) == 0:
return df.copy()
n_jobs = max(1, min(n_jobs, len(df)))
# Split DataFrame into chunks
chunk_size = max(1, len(df) // n_jobs)
chunks = [df.iloc[i:i+chunk_size] for i in range(0, len(df), chunk_size)]
def apply_filters(chunk):
filtered = chunk.copy()
# Apply each filter
if 'anony_only' in filters and filters['anony_only'] and 'anony' in filtered.columns:
filtered = filtered[filtered['anony'] == True]
if 'language' in filters and filters['language'] and 'language' in filtered.columns:
filtered = filtered[filtered['language'] == filters['language']]
if 'min_turn' in filters and 'turn' in filtered.columns:
filtered = filtered[filtered['turn'] >= filters['min_turn']]
if 'min_date' in filters and 'tstamp' in filtered.columns:
min_timestamp = pd.to_datetime(filters['min_date']).timestamp()
filtered = filtered[filtered['tstamp'] >= min_timestamp]
if 'max_date' in filters and 'tstamp' in filtered.columns:
max_timestamp = pd.to_datetime(filters['max_date']).timestamp()
filtered = filtered[filtered['tstamp'] <= max_timestamp]
return filtered
# Process chunks in parallel
with Pool(processes=n_jobs) as pool:
filtered_chunks = pool.map(apply_filters, chunks)
# Combine results
result = pd.concat(filtered_chunks, ignore_index=True)
return result
def optimize_dataframe(df: pd.DataFrame) -> pd.DataFrame:
"""
Optimize DataFrame memory usage by downcasting numeric types.
Args:
df: Input DataFrame
Returns:
Optimized DataFrame
"""
print("Optimizing DataFrame memory usage...")
initial_memory = df.memory_usage(deep=True).sum() / 1024**2
# Optimize numeric columns
for col in df.columns:
col_type = df[col].dtype
if col_type == 'int64':
df[col] = pd.to_numeric(df[col], downcast='integer')
elif col_type == 'float64':
df[col] = pd.to_numeric(df[col], downcast='float')
# Convert string columns to category if they have few unique values
for col in df.select_dtypes(include=['object']).columns:
try:
# Check if column contains hashable types (not dict, list, etc.)
# Try to get unique values - will fail if unhashable
num_unique = df[col].nunique()
num_total = len(df[col])
if num_total == 0:
continue
# If less than 50% unique values, convert to category
if num_unique / num_total < 0.5:
df[col] = df[col].astype('category')
except (TypeError, AttributeError):
# Column contains unhashable types (dicts, lists), skip optimization
print(f" Skipping column '{col}' (contains complex data types)")
continue
final_memory = df.memory_usage(deep=True).sum() / 1024**2
reduction = 0.0 if initial_memory == 0 else (1 - final_memory / initial_memory) * 100
print(f"Memory usage reduced from {initial_memory:.2f} MB to {final_memory:.2f} MB ({reduction:.1f}% reduction)")
return df
+80
View File
@@ -0,0 +1,80 @@
"""
Quick start demo - minimal example to get started quickly
"""
from elo_rating import EloRatingSystem
def demo_basic_elo():
"""Demonstrate basic Elo rating calculation with synthetic data."""
print("="*60)
print("Quick Start: Elo Rating System Demo")
print("="*60)
print()
# Initialize Elo system
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
# Simulate some matches
matches = [
("GPT-4", "Claude-v1", "GPT-4"),
("GPT-4", "Llama-2", "GPT-4"),
("Claude-v1", "Llama-2", "Claude-v1"),
("GPT-4", "Claude-v1", "tie"),
("Llama-2", "Gemini", "Gemini"),
("GPT-4", "Gemini", "GPT-4"),
("Claude-v1", "Gemini", "Claude-v1"),
("GPT-4", "Llama-2", "GPT-4"),
("Claude-v1", "Llama-2", "Claude-v1"),
("Gemini", "Llama-2", "Gemini"),
]
print("Processing matches:")
print("-" * 60)
for i, (model_a, model_b, winner) in enumerate(matches, 1):
old_rating_a = elo.get_rating(model_a)
old_rating_b = elo.get_rating(model_b)
# update_ratings expects 'model_a' / 'model_b' / 'tie', not the
# winning model's name (anything unrecognized is scored as a tie).
outcome = ("model_a" if winner == model_a
else "model_b" if winner == model_b else "tie")
new_rating_a, new_rating_b = elo.update_ratings(model_a, model_b, outcome)
print(f"Match {i}: {model_a} vs {model_b} -> {winner} wins")
print(f" {model_a}: {old_rating_a:.1f}{new_rating_a:.1f} ({new_rating_a-old_rating_a:+.1f})")
print(f" {model_b}: {old_rating_b:.1f}{new_rating_b:.1f} ({new_rating_b-old_rating_b:+.1f})")
print()
# Show final leaderboard
print("=" * 60)
print("Final Leaderboard:")
print("=" * 60)
leaderboard = elo.get_leaderboard()
for rank, (model, rating, matches, wins) in enumerate(leaderboard, 1):
win_rate = (wins / matches * 100) if matches > 0 else 0
print(f"{rank}. {model:15s} - Rating: {rating:7.1f} | "
f"Matches: {matches:2d} | Wins: {wins:4.1f} | Win Rate: {win_rate:5.1f}%")
print()
# Show win probability predictions
print("=" * 60)
print("Win Probability Predictions:")
print("=" * 60)
models = [m[0] for m in leaderboard]
for i, model_a in enumerate(models):
for model_b in models[i+1:]:
prob = elo.calculate_win_probability(model_a, model_b)
print(f"{model_a} vs {model_b}: {prob*100:.1f}% - {(1-prob)*100:.1f}%")
print()
print("=" * 60)
print("Demo complete! Check main.py for full analysis with real data.")
print("=" * 60)
if __name__ == "__main__":
demo_basic_elo()
@@ -0,0 +1,9 @@
pandas>=2.0.0
numpy>=1.24.0
matplotlib>=3.7.0
seaborn>=0.12.0
requests>=2.31.0
plotly>=5.14.0
tqdm>=4.65.0
scikit-learn>=1.3.0
python-dotenv>=1.0.0
@@ -0,0 +1,10 @@
"""Helpers for direct execution of tests moved under tests/."""
from pathlib import Path
import sys
def bootstrap_experiment_root() -> None:
experiment_root = Path(__file__).resolve().parents[1]
if str(experiment_root) not in sys.path:
sys.path.insert(0, str(experiment_root))
@@ -0,0 +1,9 @@
"""Test import bootstrap for the elo-leaderboard experiment."""
from pathlib import Path
import sys
EXPERIMENT_ROOT = Path(__file__).resolve().parents[1]
if str(EXPERIMENT_ROOT) not in sys.path:
sys.path.insert(0, str(EXPERIMENT_ROOT))
@@ -0,0 +1,39 @@
"""Empty rating history must not crash analyze_rating_changes / get_rating_history."""
import pandas as pd
from animation import prepare_animation_data
from leaderboard import (
analyze_rating_changes,
build_historical_leaderboards,
get_rating_history,
)
def test_get_rating_history_empty_keeps_columns():
hist = build_historical_leaderboards(
pd.DataFrame(columns=["model_a", "model_b", "winner"]),
[(pd.Timestamp("2020-01-01"), pd.DataFrame(columns=["model_a", "model_b", "winner"]))],
)
rh = get_rating_history(hist)
assert list(rh.columns) == ["date", "model", "rating", "rank", "matches", "wins"]
assert len(rh) == 0
def test_analyze_empty_history_returns_empty_frame():
empty = pd.DataFrame(columns=["date", "model", "rating", "rank", "matches", "wins"])
stats = analyze_rating_changes(empty)
assert len(stats) == 0
assert "model" in stats.columns
def test_analyze_after_empty_historical_leaderboards():
hist = build_historical_leaderboards(
pd.DataFrame(columns=["model_a", "model_b", "winner"]),
[(pd.Timestamp("2020-01-01"), pd.DataFrame(columns=["model_a", "model_b", "winner"]))],
)
rh = get_rating_history(hist)
stats = analyze_rating_changes(rh)
assert len(stats) == 0
anim = prepare_animation_data(rh)
assert anim["frames"] == []
assert anim["total_frames"] == 0
@@ -0,0 +1,11 @@
"""Regression: prepare_animation_data must tolerate empty history."""
import pandas as pd
from animation import prepare_animation_data
def test_empty_history_returns_empty_frames():
df = pd.DataFrame(columns=["date", "model", "rating", "rank", "matches", "wins"])
data = prepare_animation_data(df)
assert data["frames"] == []
assert data["total_frames"] == 0
assert data["start_date"] is None
@@ -0,0 +1,35 @@
"""prepare_animation_data must keep fractional wins from Elo ties."""
import pandas as pd
from animation import prepare_animation_data
def test_tie_half_wins_are_not_truncated():
history = pd.DataFrame(
{
"date": pd.to_datetime(["2024-01-07", "2024-01-07"]),
"model": ["A", "B"],
"rating": [1000.0, 1000.0],
"rank": [1, 2],
"matches": [1, 1],
"wins": [0.5, 0.5],
}
)
data = prepare_animation_data(history, top_n=2)
wins = {m["name"]: m["wins"] for m in data["frames"][0]["models"]}
assert wins["A"] == 0.5
assert wins["B"] == 0.5
def test_whole_wins_still_serialize():
history = pd.DataFrame(
{
"date": pd.to_datetime(["2024-01-07"]),
"model": ["A"],
"rating": [1010.0],
"rank": [1],
"matches": [2],
"wins": [2.0],
}
)
data = prepare_animation_data(history, top_n=1)
assert data["frames"][0]["models"][0]["wins"] == 2.0
@@ -0,0 +1,23 @@
"""
Test suite locking out ZeroDivisionError in benchmark summary print logic
when time_basic is 0.0 or df_sample is empty.
"""
def test_benchmark_pct_reduction_zero_division():
"""
Ensure zero time_basic does not raise ZeroDivisionError during benchmark calculation.
"""
time_basic = 0.0
time_optimized = 0.0
pct_reduction = (1 - time_optimized / time_basic) * 100 if time_basic > 0 else 0.0
assert pct_reduction == 0.0
def test_benchmark_extrapolation_zero_sample():
"""
Ensure empty df_sample does not raise ZeroDivisionError during extrapolation check.
"""
df_sample = []
df_filtered = [1, 2, 3]
should_extrapolate = len(df_sample) > 0 and len(df_sample) < len(df_filtered)
assert not should_extrapolate
@@ -0,0 +1,16 @@
import pandas as pd
from bradley_terry import compute_mle_elo, get_bootstrap_result
def test_bootstrap_is_reproducible():
battles = pd.DataFrame(
[
{"model_a": "a", "model_b": "b", "winner": "model_a"},
{"model_a": "a", "model_b": "b", "winner": "model_b"},
{"model_a": "a", "model_b": "b", "winner": "tie"},
{"model_a": "b", "model_b": "a", "winner": "model_a"},
]
)
first = get_bootstrap_result(battles, compute_mle_elo, num_round=3)
second = get_bootstrap_result(battles, compute_mle_elo, num_round=3)
pd.testing.assert_frame_equal(first, second)
@@ -0,0 +1,11 @@
"""Regression: compute_mle_elo must work on small Arena-shaped battle sets."""
import pandas as pd
from battle_simulator import simulate_battles
from bradley_terry import compute_mle_elo
def test_small_two_model_sample():
df = pd.DataFrame(simulate_battles({"gpt-4": 1200.0, "llama-3": 1000.0}, 10, seed=1))
ratings = compute_mle_elo(df)
assert len(ratings) == 2
assert set(ratings.index) == {"gpt-4", "llama-3"}
@@ -0,0 +1,41 @@
"""Ties must contribute to Bradley-Terry weights (not be zeroed by pivot+T)."""
import pandas as pd
from bradley_terry import compute_mle_elo
def test_all_ties_rates_models_instead_of_sample_weight_error():
df = pd.DataFrame(
[
{"model_a": "A", "model_b": "B", "winner": "tie"},
{"model_a": "A", "model_b": "C", "winner": "tie (bothbad)"},
{"model_a": "B", "model_b": "C", "winner": "tie"},
]
)
ratings = compute_mle_elo(df)
assert set(ratings.index) == {"A", "B", "C"}
# Pure ties -> equal latent skills under BT.
assert abs(float(ratings["A"]) - float(ratings["B"])) < 1e-6
assert abs(float(ratings["A"]) - float(ratings["C"])) < 1e-6
def test_ties_change_ratings_versus_wins_only():
wins_only = pd.DataFrame(
[
{"model_a": "A", "model_b": "B", "winner": "model_a"},
{"model_a": "B", "model_b": "C", "winner": "model_a"},
]
)
with_ties = pd.concat(
[
wins_only,
pd.DataFrame(
[{"model_a": "A", "model_b": "C", "winner": "tie"}] * 8
),
],
ignore_index=True,
)
r1 = compute_mle_elo(wins_only)
r2 = compute_mle_elo(with_ties)
# Extra AC ties pull A and C together relative to the wins-only fit.
assert abs(float(r2["A"]) - float(r2["C"])) < abs(float(r1["A"]) - float(r1["C"]))
@@ -0,0 +1,32 @@
"""Regression test for prepare_animation_data with string or date objects in history_df."""
import pandas as pd
from animation import prepare_animation_data
def test_prepare_animation_data_string_date():
"""prepare_animation_data must handle string dates without raising AttributeError."""
history = pd.DataFrame([
{
"date": "2024-08-01",
"model": "model_a",
"rating": 1050.0,
"rank": 1,
"matches": 10,
"wins": 7.0,
},
{
"date": "2024-08-01",
"model": "model_b",
"rating": 950.0,
"rank": 2,
"matches": 10,
"wins": 3.0,
},
])
data = prepare_animation_data(history, top_n=2)
assert data["total_frames"] == 1
assert data["start_date"] == "2024-08-01"
assert data["end_date"] == "2024-08-01"
assert len(data["frames"]) == 1
assert data["frames"][0]["date"] == "2024-08-01"
assert data["frames"][0]["timestamp"] == 1722470400
@@ -0,0 +1,17 @@
"""Regression test for compare_win_rates when comparisons list is empty."""
import numpy as np
import pandas as pd
from elo_rating import EloRatingSystem
from leaderboard import compare_win_rates
def test_compare_win_rates_empty_has_required_columns():
"""compare_win_rates must return a DataFrame with required columns when no valid comparisons exist."""
elo = EloRatingSystem()
empirical_df = pd.DataFrame(np.nan, index=["model_a", "model_b"], columns=["model_a", "model_b"])
df_comp = compare_win_rates(elo, empirical_df)
assert list(df_comp.columns) == ["model_a", "model_b", "empirical", "predicted", "error"]
assert len(df_comp) == 0
# Accessing columns on empty result must not raise KeyError
assert "error" in df_comp
assert df_comp["error"].empty
+170
View File
@@ -0,0 +1,170 @@
"""
Unit tests for Elo rating system
"""
import math
import pytest
from _bootstrap import bootstrap_experiment_root
bootstrap_experiment_root()
from elo_rating import EloRatingSystem
def test_initial_rating():
"""Test that models start with initial rating."""
elo = EloRatingSystem(initial_rating=1000.0)
assert elo.get_rating("model_a") == 1000.0
assert elo.get_rating("model_b") == 1000.0
def test_expected_score():
"""Test expected score calculation."""
elo = EloRatingSystem()
# Equal ratings should give 50% probability
assert elo.expected_score(1000, 1000) == 0.5
# Higher rated player should have > 50% probability
assert elo.expected_score(1200, 1000) > 0.5
assert elo.expected_score(1000, 1200) < 0.5
# 400 point difference should give ~91% probability
prob = elo.expected_score(1400, 1000)
assert 0.90 < prob < 0.92
def test_rating_update_win():
"""Test rating update when model_a wins."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
new_a, new_b = elo.update_ratings("model_a", "model_b", "model_a")
# Winner should gain rating, loser should lose rating
assert new_a > 1000.0
assert new_b < 1000.0
# Total rating should be conserved (zero-sum)
assert abs((new_a + new_b) - 2000.0) < 0.01
def test_rating_update_tie():
"""Test rating update for a tie."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
new_a, new_b = elo.update_ratings("model_a", "model_b", "tie")
# With equal ratings, tie should not change ratings much
assert abs(new_a - 1000.0) < 0.01
assert abs(new_b - 1000.0) < 0.01
def test_upset_gives_larger_change():
"""Test that unexpected results cause larger rating changes."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
# Give model_a higher rating
elo.ratings["model_a"] = 1200.0
elo.ratings["model_b"] = 1000.0
# If weaker model wins (upset), changes should be larger
new_a_upset, new_b_upset = elo.update_ratings("model_a", "model_b", "model_b")
# Reset
elo.ratings["model_a"] = 1200.0
elo.ratings["model_b"] = 1000.0
# If stronger model wins (expected), changes should be smaller
new_a_expected, new_b_expected = elo.update_ratings("model_a", "model_b", "model_a")
# Upset should cause larger change
change_upset = abs(new_a_upset - 1200.0)
change_expected = abs(new_a_expected - 1200.0)
assert change_upset > change_expected
def test_leaderboard_sorting():
"""Test that leaderboard is sorted by rating."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
# Create some matches to differentiate ratings
elo.update_ratings("model_a", "model_b", "model_a")
elo.update_ratings("model_a", "model_c", "model_a")
elo.update_ratings("model_b", "model_c", "model_b")
leaderboard = elo.get_leaderboard()
# Check descending order
for i in range(len(leaderboard) - 1):
assert leaderboard[i][1] >= leaderboard[i+1][1]
# model_a should be first (won all matches)
assert leaderboard[0][0] == "model_a"
def test_win_probability_symmetry():
"""Test that win probabilities sum to 1."""
elo = EloRatingSystem()
elo.ratings["model_a"] = 1200.0
elo.ratings["model_b"] = 1000.0
prob_a = elo.calculate_win_probability("model_a", "model_b")
prob_b = elo.calculate_win_probability("model_b", "model_a")
# Should sum to 1
assert abs(prob_a + prob_b - 1.0) < 0.001
def test_match_counting():
"""Test that match and win counts are tracked correctly."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
elo.update_ratings("model_a", "model_b", "model_a") # model_a wins
elo.update_ratings("model_a", "model_c", "model_b") # model_a loses (2nd slot wins)
elo.update_ratings("model_a", "model_b", "tie") # tie -> 0.5 each
# model_a played 3 matches
assert elo.match_counts["model_a"] == 3
# model_a won 1 match and tied 1 (1.5 total)
assert elo.win_counts["model_a"] == 1.5
# model_b played 2 matches
assert elo.match_counts["model_b"] == 2
def test_copy():
"""Test that copy creates independent instance."""
elo1 = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
elo1.update_ratings("model_a", "model_b", "model_a")
elo2 = elo1.copy()
# Modify elo2
elo2.update_ratings("model_a", "model_b", "model_b")
# elo1 should be unchanged
assert elo1.ratings["model_a"] != elo2.ratings["model_a"]
def test_reset():
"""Test that reset clears all data."""
elo = EloRatingSystem(initial_rating=1000.0, k_factor=32.0)
elo.update_ratings("model_a", "model_b", "model_a")
elo.update_ratings("model_a", "model_c", "model_a")
assert len(elo.ratings) > 0
elo.reset()
assert len(elo.ratings) == 0
assert len(elo.match_counts) == 0
assert len(elo.win_counts) == 0
if __name__ == "__main__":
# Run tests
pytest.main([__file__, "-v"])
@@ -0,0 +1,21 @@
"""Regression: filter_data_parallel must tolerate n_jobs > len(df)."""
from unittest.mock import MagicMock, patch
import pandas as pd
from parallel_processing import filter_data_parallel
def test_n_jobs_larger_than_rows():
df = pd.DataFrame({"anony": [True, False, True], "turn": [1, 2, 1]})
def map_inline(fn, chunks):
return [fn(c) for c in chunks]
pool = MagicMock()
pool.__enter__.return_value.map.side_effect = map_inline
pool.__exit__.return_value = False
with patch("parallel_processing.Pool", return_value=pool):
out = filter_data_parallel(df, {"anony_only": True}, n_jobs=8)
assert len(out) == 2
@@ -0,0 +1,37 @@
"""
Regression test for filter_data on empty input (实验 7-7 排行榜).
An empty arena data file (e.g. a failed/truncated download saved as `[]`) used to
crash with ZeroDivisionError at the "After filtering" percentage print.
"""
import pandas as pd
import pytest
from _bootstrap import bootstrap_experiment_root
bootstrap_experiment_root()
from data_loader import filter_data
def test_filter_data_tolerates_empty_dataframe():
"""Empty input no longer raises ZeroDivisionError; returns an empty DataFrame."""
empty = pd.DataFrame({"model_a": [], "model_b": [], "winner": []})
result = filter_data(empty)
assert len(result) == 0
def test_filter_data_normal_case_unchanged():
"""Non-empty input still filters and reports normally."""
df = pd.DataFrame({
"model_a": ["a", "b", "a"],
"model_b": ["b", "a", "c"],
"winner": ["model_a", "model_b", "tie"],
"anony": [True, True, False],
})
result = filter_data(df, anony_only=True, use_dedup=False)
assert len(result) == 2 # 非匿名的一条被过滤
if __name__ == "__main__":
pytest.main([__file__, "-v"])
@@ -0,0 +1,34 @@
"""Empty battles JSON array [] must load as an empty battle frame."""
import json
from pathlib import Path
import cli
def test_load_battles_empty_json_array(tmp_path):
path = tmp_path / "battles.json"
path.write_text("[]", encoding="utf-8")
df = cli._load_battles(str(path))
assert list(df.columns) == ["model_a", "model_b", "winner"]
assert len(df) == 0
def test_load_battles_nonempty_still_requires_columns(tmp_path):
path = tmp_path / "bad.json"
path.write_text(json.dumps([{"x": 1}]), encoding="utf-8")
try:
cli._load_battles(str(path))
assert False, "expected ValueError"
except ValueError as e:
assert "model_a" in str(e)
def test_load_battles_normal(tmp_path):
path = tmp_path / "ok.json"
path.write_text(
json.dumps([{"model_a": "A", "model_b": "B", "winner": "model_a"}]),
encoding="utf-8",
)
df = cli._load_battles(str(path))
assert len(df) == 1
assert df.iloc[0]["winner"] == "model_a"
@@ -0,0 +1,13 @@
"""Regression: optimize_dataframe must tolerate empty object columns."""
import pandas as pd
from parallel_processing import optimize_dataframe
def test_optimize_empty_object_columns():
df = pd.DataFrame({
"model_a": pd.Series([], dtype=object),
"model_b": pd.Series([], dtype=object),
"winner": pd.Series([], dtype=object),
})
out = optimize_dataframe(df)
assert len(out) == 0
@@ -0,0 +1,49 @@
"""
Regression test for 'tie (bothbad)' handling in optimized_elo (实验 7-7 排行榜).
Chatbot Arena battle data has four outcomes; 'tie (bothbad)' was missing from
the outcome map, so Series.map produced NaN. NaN then propagated through the
rating updates and spread to every model that later faced an affected one,
leaving the whole leaderboard NaN.
"""
import numpy as np
import pandas as pd
from optimized_elo import NumpyEloRatingSystem
def test_tie_bothbad_does_not_produce_nan_outcomes():
"""'tie (bothbad)' maps to a tie instead of NaN."""
df = pd.DataFrame({
"model_a": ["a", "a"],
"model_b": ["b", "b"],
"winner": ["model_a", "tie (bothbad)"],
})
_, _, outcomes = NumpyEloRatingSystem()._prepare_data(df)
assert not np.isnan(outcomes).any()
assert outcomes[1] == 0.5
def test_tie_bothbad_does_not_poison_the_leaderboard():
"""One 'tie (bothbad)' battle used to NaN every rating, including model c."""
df = pd.DataFrame({
"model_a": ["a", "a", "b"],
"model_b": ["b", "b", "c"],
"winner": ["model_a", "tie (bothbad)", "model_a"],
})
system = NumpyEloRatingSystem()
system.process_matches_vectorized(df, show_progress=False)
ratings = [rating for _, rating, _, _ in system.get_leaderboard()]
assert len(ratings) == 3
assert not any(np.isnan(r) for r in ratings)
def test_unknown_outcome_falls_back_to_tie():
"""An unrecognized label degrades to a tie rather than NaN."""
df = pd.DataFrame({
"model_a": ["a"],
"model_b": ["b"],
"winner": ["something_new"],
})
_, _, outcomes = NumpyEloRatingSystem()._prepare_data(df)
assert outcomes[0] == 0.5
@@ -0,0 +1,9 @@
"""Regression: documented interval='M' must work on modern pandas."""
import pandas as pd
from data_loader import get_time_slices
def test_monthly_interval_alias():
df = pd.DataFrame({"tstamp": [1_700_000_000, 1_710_000_000]})
slices = get_time_slices(df, interval="M")
assert len(slices) >= 1
@@ -0,0 +1,55 @@
"""
Regression: get_time_slices must not IndexError when the tstamp span is
shorter than the requested interval (default weekly).
Chatbot Arena samples, same-second dumps, and single-row demos all produce an
empty pd.date_range for freq='W'; the old code then crashed on date_ranges[-1].
"""
import pandas as pd
from _bootstrap import bootstrap_experiment_root
bootstrap_experiment_root()
from data_loader import get_time_slices
def test_identical_timestamps_return_one_slice():
"""Two battles at the same unix second (weekly interval) -> one slice."""
ts = 1_700_000_000
df = pd.DataFrame({
"tstamp": [ts, ts],
"model_a": ["a", "c"],
"model_b": ["b", "d"],
"winner": ["model_a", "model_b"],
})
slices = get_time_slices(df, interval="W")
assert len(slices) == 1
end_date, slice_df = slices[0]
assert len(slice_df) == 2
assert end_date == pd.to_datetime(ts, unit="s")
def test_empty_dataframe_returns_empty_list():
"""Empty input returns [] instead of NaT ValueError."""
df = pd.DataFrame({"tstamp": pd.Series(dtype="float64")})
assert get_time_slices(df, interval="W") == []
def test_multi_week_span_still_produces_buckets():
"""A span covering multiple weeks still yields intermediate buckets."""
# ~3 weeks apart
df = pd.DataFrame({
"tstamp": [1_700_000_000, 1_700_000_000 + 21 * 86400],
"model_a": ["a", "c"],
"model_b": ["b", "d"],
"winner": ["model_a", "model_b"],
})
slices = get_time_slices(df, interval="W")
assert len(slices) >= 2
assert all(len(s[1]) > 0 for s in slices)
if __name__ == "__main__":
import pytest
pytest.main([__file__, "-v"])
@@ -0,0 +1,20 @@
import json
from pathlib import Path
def test_canonical_manifest_is_hash_complete():
run_dir = Path(__file__).resolve().parents[1] / "validation" / "runs" / "exp7-7-arena-20260731-v1"
manifest_path = run_dir / "manifest.json"
assert manifest_path.exists()
manifest = json.loads(manifest_path.read_text(encoding="utf-8"))
assert manifest["experiment"] == "7-7"
assert manifest["official_complete"] is True
assert all(manifest["gates"].values())
assert set(manifest["artifacts"]) >= {
"summary.json",
"online_elo.json",
"bradley_terry.json",
"win_rate_matrix.json",
"rating_history.json",
"leaderboard_animation.html",
}
@@ -0,0 +1,32 @@
"""Empty battle DataFrame must not crash Bradley-Terry LogisticRegression."""
import pandas as pd
from bradley_terry import compute_bradley_terry_leaderboard, compute_mle_elo
def test_compute_mle_elo_empty_battles():
df = pd.DataFrame(columns=["model_a", "model_b", "winner"])
ratings = compute_mle_elo(df)
assert isinstance(ratings, pd.Series)
assert len(ratings) == 0
def test_compute_bradley_terry_leaderboard_empty():
df = pd.DataFrame(columns=["model_a", "model_b", "winner"])
board = compute_bradley_terry_leaderboard(df)
assert isinstance(board, pd.DataFrame)
assert len(board) == 0
def test_nonempty_still_rates():
df = pd.DataFrame(
[
{"model_a": "A", "model_b": "B", "winner": "model_a"},
{"model_a": "A", "model_b": "B", "winner": "model_a"},
{"model_a": "B", "model_b": "C", "winner": "model_b"},
{"model_a": "A", "model_b": "C", "winner": "model_a"},
]
)
ratings = compute_mle_elo(df)
assert set(ratings.index) >= {"A", "B", "C"}
assert ratings["A"] > ratings["C"]
@@ -0,0 +1,7 @@
{
"experiment": "7-7",
"manifest_sha256": "b67f8b15a088e694ae356322e0ccd676b54ed79a0dc1c787ca1960ef670ba2e8",
"official_complete": true,
"run": "runs/exp7-7-arena-20260731-v1",
"status": "passed"
}
@@ -0,0 +1,389 @@
"""Run the complete, evidence-producing Experiment 7-7 campaign.
The public Arena file is deliberately not copied into git. A canonical run
binds the exact input by URL, size, record count, and SHA-256, then retains all
derived tables, visualizations, the D3 history animation, and a manifest that
hashes every output and the source used to create it.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import os
import platform
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
os.environ.setdefault("MPLBACKEND", "Agg")
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import seaborn as sns
from scipy.stats import kendalltau, spearmanr
HERE = Path(__file__).resolve().parent
PROJECT = HERE.parent
sys.path.insert(0, str(PROJECT))
from animation import create_simple_animation
from bradley_terry import compute_bradley_terry_leaderboard
from optimized_elo import (
NumpyEloRatingSystem,
process_elo_updates_vectorized,
)
DATASET_URL = (
"https://storage.googleapis.com/arena_external_data/public/"
"clean_battle_20240814_public.json"
)
REQUIRED_COLUMNS = ["model_a", "model_b", "winner", "tstamp", "anony", "turn"]
ALLOWED_OUTCOMES = {"model_a", "model_b", "tie", "tie (bothbad)"}
def sha256_file(path: Path, chunk_size: int = 8 * 1024 * 1024) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
while chunk := handle.read(chunk_size):
digest.update(chunk)
return digest.hexdigest()
def write_json(path: Path, value: Any) -> None:
path.write_text(
json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n",
encoding="utf-8",
)
def json_records(frame: pd.DataFrame) -> list[dict[str, Any]]:
return json.loads(frame.to_json(orient="records", date_format="iso"))
def load_and_filter(path: Path, max_records: int) -> tuple[pd.DataFrame, dict[str, Any]]:
started = time.perf_counter()
raw = pd.read_json(path)
missing = sorted(set(REQUIRED_COLUMNS) - set(raw.columns))
if missing:
raise ValueError(f"Arena input is missing columns: {missing}")
source_records = len(raw)
frame = raw[REQUIRED_COLUMNS + (["dedup_tag"] if "dedup_tag" in raw else [])].copy()
del raw
frame = frame[frame["anony"].eq(True) & frame["turn"].ge(1)]
if "dedup_tag" in frame:
sampled = frame["dedup_tag"].map(
lambda value: bool(value.get("sampled", False)) if isinstance(value, dict) else False
)
frame = frame[sampled]
frame = frame[frame["winner"].isin(ALLOWED_OUTCOMES)]
frame = frame.sort_values("tstamp", kind="stable").reset_index(drop=True)
if max_records:
frame = frame.head(max_records).copy()
if frame.empty:
raise ValueError("Arena filtering produced no accepted blind votes")
metadata = {
"source_records": source_records,
"accepted_records": len(frame),
"model_count": len(set(frame["model_a"]) | set(frame["model_b"])),
"start_utc": datetime.fromtimestamp(float(frame["tstamp"].min()), timezone.utc).isoformat(),
"end_utc": datetime.fromtimestamp(float(frame["tstamp"].max()), timezone.utc).isoformat(),
"outcomes": {str(k): int(v) for k, v in frame["winner"].value_counts().items()},
"load_filter_seconds": round(time.perf_counter() - started, 3),
"bounded_test_run": bool(max_records),
}
return frame, metadata
def online_elo_and_history(
frame: pd.DataFrame,
) -> tuple[pd.DataFrame, pd.DataFrame, NumpyEloRatingSystem, float]:
started = time.perf_counter()
system = NumpyEloRatingSystem(initial_rating=1000.0, k_factor=4.0)
model_a, model_b, outcomes = system._prepare_data(frame)
months = (
pd.to_datetime(frame["tstamp"], unit="s", utc=True)
.dt.tz_localize(None)
.dt.to_period("M")
)
boundaries = np.flatnonzero(months.to_numpy()[1:] != months.to_numpy()[:-1]) + 1
boundaries = np.append(boundaries, len(frame))
start = 0
history_rows: list[dict[str, Any]] = []
for stop in boundaries:
process_elo_updates_vectorized(
system.ratings,
model_a[start:stop],
model_b[start:stop],
outcomes[start:stop],
system.k_factor,
system.match_counts,
system.win_counts,
)
snapshot = system.get_leaderboard()
date = pd.to_datetime(float(frame.iloc[stop - 1]["tstamp"]), unit="s", utc=True)
for rank, (model, rating, matches, wins) in enumerate(snapshot, 1):
history_rows.append(
{
"date": date.tz_localize(None),
"model": model,
"rating": rating,
"rank": rank,
"matches": matches,
"wins": wins,
}
)
start = int(stop)
leaderboard = pd.DataFrame(
system.get_leaderboard(), columns=["model", "rating", "matches", "wins"]
)
leaderboard.insert(0, "rank", range(1, len(leaderboard) + 1))
history = pd.DataFrame(history_rows)
return leaderboard, history, system, round(time.perf_counter() - started, 3)
def rank_comparison(online: pd.DataFrame, official_method: pd.DataFrame) -> dict[str, Any]:
online_rank = online.set_index("model")["rank"]
official = official_method.sort_values("rating", ascending=False).reset_index(drop=True)
official["rank"] = np.arange(1, len(official) + 1)
official_rank = official.set_index("model")["rank"]
common = sorted(set(online_rank.index) & set(official_rank.index))
rho = spearmanr(online_rank.loc[common], official_rank.loc[common]).statistic
tau = kendalltau(online_rank.loc[common], official_rank.loc[common]).statistic
online_top = online.nsmallest(20, "rank")["model"].tolist()
official_top = official.nsmallest(20, "rank")["model"].tolist()
return {
"comparison_target": "Bradley-Terry MLE reconstruction used by Chatbot Arena",
"claim_boundary": (
"This is a same-snapshot reconstruction of the official method, not a scrape of "
"the mutable live leaderboard. Scores need not match the live service."
),
"common_models": len(common),
"spearman_rank_correlation": round(float(rho), 6),
"kendall_rank_correlation": round(float(tau), 6),
"top_20_overlap": len(set(online_top) & set(official_top)),
"online_top_20": online_top,
"official_method_top_20": official_top,
}
def empirical_matrix(frame: pd.DataFrame, models: list[str]) -> pd.DataFrame:
subset = frame[frame["model_a"].isin(models) & frame["model_b"].isin(models)].copy()
rows: list[tuple[str, str, float]] = []
for a, b, winner in subset[["model_a", "model_b", "winner"]].itertuples(index=False):
score = 1.0 if winner == "model_a" else 0.0 if winner == "model_b" else 0.5
rows.append((a, b, score))
rows.append((b, a, 1.0 - score))
scored = pd.DataFrame(rows, columns=["model", "opponent", "score"])
matrix = scored.pivot_table(index="model", columns="opponent", values="score", aggfunc="mean")
matrix = matrix.reindex(index=models, columns=models)
np.fill_diagonal(matrix.values, 0.5)
return matrix
def plot_artifacts(
out: Path,
online: pd.DataFrame,
history: pd.DataFrame,
empirical: pd.DataFrame,
) -> None:
top = online.head(20).sort_values("rating")
fig, ax = plt.subplots(figsize=(11, 8))
ax.barh(top["model"], top["rating"], color="#3b82f6")
ax.set_title("Experiment 7-7: Online Elo leaderboard")
ax.set_xlabel("Elo rating (K=4, chronological)")
fig.tight_layout()
fig.savefig(out / "leaderboard.png", dpi=180)
plt.close(fig)
fig, ax = plt.subplots(figsize=(13, 11))
sns.heatmap(empirical, cmap="RdYlGn", center=0.5, vmin=0, vmax=1, ax=ax)
ax.set_title("Empirical pairwise win rate — final online-Elo top 20")
fig.tight_layout()
fig.savefig(out / "win_rate_matrix.png", dpi=180)
plt.close(fig)
top_models = online.head(10)["model"].tolist()
fig, ax = plt.subplots(figsize=(13, 7))
for model in top_models:
values = history[history["model"].eq(model)].sort_values("date")
ax.plot(values["date"], values["rating"], label=model, linewidth=1.8)
ax.set_title("Monthly online-Elo evolution — final top 10")
ax.set_ylabel("Elo rating")
ax.legend(fontsize=7, ncol=2)
fig.autofmt_xdate()
fig.tight_layout()
fig.savefig(out / "rating_history.png", dpi=180)
plt.close(fig)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--input", type=Path, required=True, help="Downloaded public Arena JSON")
parser.add_argument("--output-dir", type=Path, required=True)
parser.add_argument("--bootstrap-rounds", type=int, default=20)
parser.add_argument("--max-records", type=int, default=0, help="Noncanonical bounded test only")
return parser.parse_args()
def main() -> None:
args = parse_args()
args.output_dir.mkdir(parents=True, exist_ok=True)
input_path = args.input.resolve()
if not input_path.is_file():
raise SystemExit(f"Arena input not found: {input_path}")
run_started = time.perf_counter()
input_hash = sha256_file(input_path)
frame, dataset = load_and_filter(input_path, args.max_records)
online, history, online_system, online_seconds = online_elo_and_history(frame)
bt_started = time.perf_counter()
official_method = compute_bradley_terry_leaderboard(
frame[["model_a", "model_b", "winner"]],
bootstrap_rounds=args.bootstrap_rounds,
)
bt_seconds = round(time.perf_counter() - bt_started, 3)
official_method = official_method.sort_values("rating", ascending=False).reset_index(drop=True)
official_method.insert(0, "rank", range(1, len(official_method) + 1))
comparison = rank_comparison(online, official_method)
top_models = online.head(20)["model"].tolist()
empirical = empirical_matrix(frame, top_models)
predicted = pd.DataFrame(
{
opponent: {
model: online_system.calculate_win_probability(model, opponent)
for model in top_models
}
for opponent in top_models
}
).reindex(index=top_models, columns=top_models)
write_json(args.output_dir / "online_elo.json", json_records(online))
write_json(args.output_dir / "bradley_terry.json", json_records(official_method))
write_json(
args.output_dir / "win_rate_matrix.json",
{
"models": top_models,
"empirical": empirical.where(pd.notna(empirical), None).to_dict(orient="index"),
"online_elo_predicted": predicted.to_dict(orient="index"),
},
)
write_json(args.output_dir / "rating_history.json", json_records(history))
plot_artifacts(args.output_dir, online, history, empirical)
create_simple_animation(history, str(args.output_dir / "leaderboard_animation.html"), top_n=15)
gates = {
"official_public_arena_snapshot_hashed": not args.max_records,
"millions_of_blind_votes_loaded": dataset["source_records"] >= 1_000_000,
"chronological_online_elo_k4_completed": len(online) == dataset["model_count"],
"bradley_terry_official_method_completed": len(official_method) == dataset["model_count"],
"online_vs_official_method_rank_agreement_observed": (
comparison["spearman_rank_correlation"] >= 0.70
and comparison["top_20_overlap"] >= 10
),
"pairwise_empirical_and_predicted_matrix_saved": len(empirical) == 20,
"monthly_history_saved": history["date"].nunique() >= 2,
"d3_animation_saved": (args.output_dir / "leaderboard_animation.html").is_file(),
"static_visualizations_saved": all(
(args.output_dir / name).is_file()
for name in ["leaderboard.png", "win_rate_matrix.png", "rating_history.png"]
),
}
accepted = all(gates.values())
summary = {
"schema_version": 1,
"experiment": "7-7",
"status": "passed" if accepted else "noncanonical_test",
"official_complete": accepted,
"generated_at_utc": datetime.now(timezone.utc).isoformat(),
"dataset": {
"url": DATASET_URL,
"path_recorded_as": input_path.name,
"bytes": input_path.stat().st_size,
"sha256": input_hash,
**dataset,
},
"protocol": {
"online_elo": "initial=1000, K=4, stable chronological order",
"official_method": "Bradley-Terry maximum-likelihood reconstruction",
"history_interval": "monthly cumulative snapshots",
"bootstrap_rounds": args.bootstrap_rounds,
"bootstrap_random_seed": 0,
},
"results": {
"online_top_20": json_records(online.head(20)),
"official_method_top_20": json_records(official_method.head(20)),
"rank_comparison": comparison,
},
"timing_seconds": {
"online_and_history": online_seconds,
"bradley_terry": bt_seconds,
"total": round(time.perf_counter() - run_started, 3),
},
"gates": gates,
}
write_json(args.output_dir / "summary.json", summary)
artifact_names = [
"online_elo.json",
"bradley_terry.json",
"win_rate_matrix.json",
"rating_history.json",
"leaderboard.png",
"win_rate_matrix.png",
"rating_history.png",
"leaderboard_animation.html",
"summary.json",
]
source_names = [
"animation.py",
"bradley_terry.py",
"optimized_elo.py",
"validation/run_experiment.py",
"validation/validate_evidence.py",
]
manifest = {
"schema_version": 1,
"experiment": "7-7",
"status": summary["status"],
"official_complete": accepted,
"input": {
"url": DATASET_URL,
"filename": input_path.name,
"bytes": input_path.stat().st_size,
"sha256": input_hash,
},
"artifacts": {
name: {"bytes": (args.output_dir / name).stat().st_size, "sha256": sha256_file(args.output_dir / name)}
for name in artifact_names
},
"sources": {
name: sha256_file(PROJECT / name)
for name in source_names
},
"runtime": {
"python": sys.version.split()[0],
"platform": platform.platform(),
"pandas": pd.__version__,
"numpy": np.__version__,
},
"gates": gates,
}
write_json(args.output_dir / "manifest.json", manifest)
print(json.dumps({"status": manifest["status"], "output": str(args.output_dir)}, indent=2))
if not accepted:
raise SystemExit(2)
if __name__ == "__main__":
main()
@@ -0,0 +1,905 @@
[
{
"lower_ci": 1199.4356201692,
"model": "chatgpt-4o-latest",
"rank": 1,
"rating": 1202.8722324768,
"upper_ci": 1207.3805726125
},
{
"lower_ci": 1182.767114327,
"model": "gemini-1.5-pro-exp-0801",
"rank": 2,
"rating": 1187.3727762743,
"upper_ci": 1191.7564879767
},
{
"lower_ci": 1173.1270556478,
"model": "gpt-4o-2024-05-13",
"rank": 3,
"rating": 1174.5475812185,
"upper_ci": 1177.0403252386
},
{
"lower_ci": 1161.1978031325,
"model": "gpt-4o-mini-2024-07-18",
"rank": 4,
"rating": 1162.9712796813,
"upper_ci": 1166.568334134
},
{
"lower_ci": 1156.0224735749,
"model": "claude-3-5-sonnet-20240620",
"rank": 5,
"rating": 1159.6692177384,
"upper_ci": 1161.8458744744
},
{
"lower_ci": 1153.6447817793,
"model": "gemini-advanced-0514",
"rank": 6,
"rating": 1155.6305094344,
"upper_ci": 1157.1030226933
},
{
"lower_ci": 1150.4614200682,
"model": "llama-3.1-405b-instruct",
"rank": 7,
"rating": 1153.3950109278,
"upper_ci": 1157.0848375454
},
{
"lower_ci": 1146.348878889,
"model": "gpt-4o-2024-08-06",
"rank": 8,
"rating": 1150.7428051987,
"upper_ci": 1156.3245456986
},
{
"lower_ci": 1146.7772301944,
"model": "gemini-1.5-pro-api-0514",
"rank": 9,
"rating": 1149.1265569601,
"upper_ci": 1151.5010539806
},
{
"lower_ci": 1143.6682527667,
"model": "gemini-1.5-pro-api-0409-preview",
"rank": 10,
"rating": 1146.1756037638,
"upper_ci": 1147.8895040945
},
{
"lower_ci": 1142.3990049162,
"model": "gpt-4-turbo-2024-04-09",
"rank": 11,
"rating": 1144.9187340632,
"upper_ci": 1146.8362809167
},
{
"lower_ci": 1138.1303324947,
"model": "gpt-4-1106-preview",
"rank": 12,
"rating": 1139.5004505155,
"upper_ci": 1140.97644635
},
{
"lower_ci": 1136.3852091186,
"model": "mistral-large-2407",
"rank": 13,
"rating": 1139.3566231011,
"upper_ci": 1143.8960727747
},
{
"lower_ci": 1135.0624903381,
"model": "athene-70b-0725",
"rank": 14,
"rating": 1138.2016223593,
"upper_ci": 1141.397916447
},
{
"lower_ci": 1135.0278139027,
"model": "claude-3-opus-20240229",
"rank": 15,
"rating": 1136.5739464324,
"upper_ci": 1137.5663039742
},
{
"lower_ci": 1131.7219814044,
"model": "llama-3.1-70b-instruct",
"rank": 16,
"rating": 1134.8942150397,
"upper_ci": 1139.3028524197
},
{
"lower_ci": 1132.1800230236,
"model": "gpt-4-0125-preview",
"rank": 17,
"rating": 1133.974390087,
"upper_ci": 1135.7552442263
},
{
"lower_ci": 1127.3300289872,
"model": "yi-large-preview",
"rank": 18,
"rating": 1128.1692229165,
"upper_ci": 1131.371743138
},
{
"lower_ci": 1111.850158929,
"model": "reka-core-20240722",
"rank": 19,
"rating": 1117.1464671073,
"upper_ci": 1121.9537941313
},
{
"lower_ci": 1114.0244965884,
"model": "gemini-1.5-flash-api-0514",
"rank": 20,
"rating": 1116.6161359294,
"upper_ci": 1118.5223962538
},
{
"lower_ci": 1103.6704090106,
"model": "deepseek-v2-api-0628",
"rank": 21,
"rating": 1107.2808888587,
"upper_ci": 1112.0417779712
},
{
"lower_ci": 1103.9432951297,
"model": "gemma-2-27b-it",
"rank": 22,
"rating": 1105.9132519337,
"upper_ci": 1108.0293102426
},
{
"lower_ci": 1099.3459906048,
"model": "deepseek-coder-v2-0724",
"rank": 23,
"rating": 1105.3954927337,
"upper_ci": 1113.0353361228
},
{
"lower_ci": 1099.9821845459,
"model": "yi-large",
"rank": 24,
"rating": 1102.4018397329,
"upper_ci": 1104.8525831946
},
{
"lower_ci": 1094.5198632757,
"model": "nemotron-4-340b-instruct",
"rank": 25,
"rating": 1098.0765768899,
"upper_ci": 1100.3082488796
},
{
"lower_ci": 1092.3225621341,
"model": "bard-jan-24-gemini-pro",
"rank": 26,
"rating": 1096.589845604,
"upper_ci": 1104.5333605578
},
{
"lower_ci": 1089.9500461084,
"model": "glm-4-0520",
"rank": 27,
"rating": 1095.5341569364,
"upper_ci": 1100.1005750187
},
{
"lower_ci": 1093.6731280724,
"model": "llama-3-70b-instruct",
"rank": 28,
"rating": 1094.9281802329,
"upper_ci": 1095.7283536945
},
{
"lower_ci": 1088.0897564085,
"model": "claude-3-sonnet-20240229",
"rank": 29,
"rating": 1090.1434308802,
"upper_ci": 1091.5131310821
},
{
"lower_ci": 1086.224736458,
"model": "reka-core-20240501",
"rank": 30,
"rating": 1088.4321613308,
"upper_ci": 1090.272779249
},
{
"lower_ci": 1082.4844014079,
"model": "reka-flash-20240722",
"rank": 31,
"rating": 1088.2114217704,
"upper_ci": 1094.1409244278
},
{
"lower_ci": 1076.2505646464,
"model": "command-r-plus",
"rank": 32,
"rating": 1078.6392855764,
"upper_ci": 1080.5241685262
},
{
"lower_ci": 1072.978332193,
"model": "gemma-2-9b-it",
"rank": 33,
"rating": 1076.1200490762,
"upper_ci": 1078.284506667
},
{
"lower_ci": 1073.5327980749,
"model": "qwen2-72b-instruct",
"rank": 34,
"rating": 1075.9397511888,
"upper_ci": 1077.9102798056
},
{
"lower_ci": 1073.0022184434,
"model": "gpt-4-0314",
"rank": 35,
"rating": 1074.4688060969,
"upper_ci": 1077.3841437021
},
{
"lower_ci": 1069.6849684012,
"model": "qwen-max-0428",
"rank": 36,
"rating": 1072.288830872,
"upper_ci": 1075.2788486953
},
{
"lower_ci": 1067.4064473159,
"model": "glm-4-0116",
"rank": 37,
"rating": 1072.1854778301,
"upper_ci": 1078.0278828583
},
{
"lower_ci": 1065.2785446605,
"model": "claude-3-haiku-20240307",
"rank": 38,
"rating": 1067.2876085746,
"upper_ci": 1068.5180263057
},
{
"lower_ci": 1063.4499078695,
"model": "deepseek-coder-v2",
"rank": 39,
"rating": 1066.5634197273,
"upper_ci": 1071.0372811761
},
{
"lower_ci": 1052.4060997176,
"model": "llama-3.1-8b-instruct",
"rank": 40,
"rating": 1057.2617029177,
"upper_ci": 1062.4851344029
},
{
"lower_ci": 1050.6516231153,
"model": "reka-flash-preview-20240611",
"rank": 41,
"rating": 1053.4286082128,
"upper_ci": 1056.9047978837
},
{
"lower_ci": 1049.878634752,
"model": "gpt-4-0613",
"rank": 42,
"rating": 1051.2288727429,
"upper_ci": 1053.3859970522
},
{
"lower_ci": 1047.658821981,
"model": "qwen1.5-110b-chat",
"rank": 43,
"rating": 1050.1472801797,
"upper_ci": 1054.1870464263
},
{
"lower_ci": 1043.4416581432,
"model": "yi-1.5-34b-chat",
"rank": 44,
"rating": 1046.5798307425,
"upper_ci": 1048.4870877901
},
{
"lower_ci": 1043.9717227179,
"model": "mistral-large-2402",
"rank": 45,
"rating": 1045.8549746499,
"upper_ci": 1048.6558688014
},
{
"lower_ci": 1039.8435107486,
"model": "reka-flash-21b-20240226-online",
"rank": 46,
"rating": 1044.4541273285,
"upper_ci": 1047.5722059799
},
{
"lower_ci": 1038.3457593513,
"model": "llama-3-8b-instruct",
"rank": 47,
"rating": 1040.8943517232,
"upper_ci": 1042.1727089946
},
{
"lower_ci": 1034.3532315209,
"model": "command-r",
"rank": 48,
"rating": 1037.5954245649,
"upper_ci": 1039.801976874
},
{
"lower_ci": 1033.8210405932,
"model": "claude-1",
"rank": 49,
"rating": 1037.2264989242,
"upper_ci": 1039.8445566211
},
{
"lower_ci": 1032.4181357042,
"model": "reka-flash-21b-20240226",
"rank": 50,
"rating": 1036.2083492306,
"upper_ci": 1038.9143450059
},
{
"lower_ci": 1033.5592570692,
"model": "mistral-medium",
"rank": 51,
"rating": 1035.9830773208,
"upper_ci": 1038.0506828269
},
{
"lower_ci": 1033.4716682328,
"model": "mixtral-8x22b-instruct-v0.1",
"rank": 52,
"rating": 1035.6187089038,
"upper_ci": 1037.3385699205
},
{
"lower_ci": 1032.3748800585,
"model": "qwen1.5-72b-chat",
"rank": 53,
"rating": 1035.5734020095,
"upper_ci": 1037.8584457516
},
{
"lower_ci": 1016.6859168483,
"model": "claude-2.0",
"rank": 54,
"rating": 1020.4032512576,
"upper_ci": 1023.0900087939
},
{
"lower_ci": 1016.4818650296,
"model": "gemini-pro-dev-api",
"rank": 55,
"rating": 1018.8216738996,
"upper_ci": 1022.0575603351
},
{
"lower_ci": 1013.9644553596,
"model": "gemma-2-2b-it",
"rank": 56,
"rating": 1018.7496539694,
"upper_ci": 1022.6257555787
},
{
"lower_ci": 1009.8189712109,
"model": "zephyr-orpo-141b-A35b-v0.1",
"rank": 57,
"rating": 1017.4768858149,
"upper_ci": 1023.4231756618
},
{
"lower_ci": 1011.7444717478,
"model": "qwen1.5-32b-chat",
"rank": 58,
"rating": 1014.2192047782,
"upper_ci": 1017.4848181067
},
{
"lower_ci": 1009.527805507,
"model": "mistral-next",
"rank": 59,
"rating": 1013.8190347615,
"upper_ci": 1016.756252395
},
{
"lower_ci": 1008.0043265533,
"model": "phi-3-medium-4k-instruct",
"rank": 60,
"rating": 1011.6064259691,
"upper_ci": 1014.8977711664
},
{
"lower_ci": 1003.3664596337,
"model": "claude-2.1",
"rank": 61,
"rating": 1006.9894951694,
"upper_ci": 1008.7578243389
},
{
"lower_ci": 1003.5743630599,
"model": "starling-lm-7b-beta",
"rank": 62,
"rating": 1006.9783218473,
"upper_ci": 1009.6647795891
},
{
"lower_ci": 1003.3560004715,
"model": "gpt-3.5-turbo-0613",
"rank": 63,
"rating": 1004.9730864576,
"upper_ci": 1008.7646121872
},
{
"lower_ci": 1000.3011942689,
"model": "mixtral-8x7b-instruct-v0.1",
"rank": 64,
"rating": 1002.12545233,
"upper_ci": 1004.4555563028
},
{
"lower_ci": 996.8355186436,
"model": "yi-34b-chat",
"rank": 65,
"rating": 1000.1263930743,
"upper_ci": 1002.9767969377
},
{
"lower_ci": 996.5513388826,
"model": "claude-instant-1",
"rank": 66,
"rating": 999.4536204497,
"upper_ci": 1003.0420312018
},
{
"lower_ci": 991.5815494729,
"model": "gemini-pro",
"rank": 67,
"rating": 998.3909699055,
"upper_ci": 1005.6665144072
},
{
"lower_ci": 994.5272838271,
"model": "qwen1.5-14b-chat",
"rank": 68,
"rating": 997.6714948516,
"upper_ci": 1000.3276398407
},
{
"lower_ci": 988.7882364736,
"model": "gpt-3.5-turbo-0314",
"rank": 69,
"rating": 997.4790766208,
"upper_ci": 1005.2920277339
},
{
"lower_ci": 989.2933547328,
"model": "wizardlm-70b",
"rank": 70,
"rating": 995.5069328559,
"upper_ci": 1000.0203225061
},
{
"lower_ci": 992.3677866278,
"model": "gpt-3.5-turbo-0125",
"rank": 71,
"rating": 994.221194786,
"upper_ci": 996.0922832004
},
{
"lower_ci": 989.517752069,
"model": "dbrx-instruct-preview",
"rank": 72,
"rating": 991.6558051925,
"upper_ci": 994.3414735308
},
{
"lower_ci": 988.3935167102,
"model": "phi-3-small-8k-instruct",
"rank": 73,
"rating": 990.3926872181,
"upper_ci": 993.4532484928
},
{
"lower_ci": 980.2003929739,
"model": "tulu-2-dpo-70b",
"rank": 74,
"rating": 986.8619167507,
"upper_ci": 995.6082960114
},
{
"lower_ci": 979.2393563337,
"model": "llama-2-70b-chat",
"rank": 75,
"rating": 982.0382072414,
"upper_ci": 983.1882888775
},
{
"lower_ci": 976.2082709482,
"model": "openchat-3.5-0106",
"rank": 76,
"rating": 980.6486425506,
"upper_ci": 984.7448454329
},
{
"lower_ci": 976.21123294,
"model": "vicuna-33b",
"rank": 77,
"rating": 978.8419077515,
"upper_ci": 983.0972163635
},
{
"lower_ci": 976.387496134,
"model": "snowflake-arctic-instruct",
"rank": 78,
"rating": 978.4458039182,
"upper_ci": 981.5060581663
},
{
"lower_ci": 969.5806601345,
"model": "starling-lm-7b-alpha",
"rank": 79,
"rating": 976.025808172,
"upper_ci": 980.086762811
},
{
"lower_ci": 965.8992419629,
"model": "nous-hermes-2-mixtral-8x7b-dpo",
"rank": 80,
"rating": 975.4699172429,
"upper_ci": 981.307231701
},
{
"lower_ci": 969.0894947313,
"model": "gemma-1.1-7b-it",
"rank": 81,
"rating": 971.847848073,
"upper_ci": 975.4703266199
},
{
"lower_ci": 966.6172052307,
"model": "llama2-70b-steerlm-chat",
"rank": 82,
"rating": 971.6537111318,
"upper_ci": 979.8268878882
},
{
"lower_ci": 961.617441776,
"model": "pplx-70b-online",
"rank": 83,
"rating": 967.4367908935,
"upper_ci": 972.8126406202
},
{
"lower_ci": 959.5692952959,
"model": "openchat-3.5",
"rank": 84,
"rating": 966.0510063396,
"upper_ci": 971.8466170437
},
{
"lower_ci": 956.7287406383,
"model": "deepseek-llm-67b-chat",
"rank": 85,
"rating": 965.1147523375,
"upper_ci": 971.4330625559
},
{
"lower_ci": 956.3934469962,
"model": "openhermes-2.5-mistral-7b",
"rank": 86,
"rating": 963.454411119,
"upper_ci": 968.2145476373
},
{
"lower_ci": 957.6795493747,
"model": "mistral-7b-instruct-v0.2",
"rank": 87,
"rating": 961.195768793,
"upper_ci": 964.3273838331
},
{
"lower_ci": 955.9914563912,
"model": "qwen1.5-7b-chat",
"rank": 88,
"rating": 959.7566395303,
"upper_ci": 963.8434392791
},
{
"lower_ci": 953.2396686903,
"model": "phi-3-mini-4k-instruct-june-2024",
"rank": 89,
"rating": 958.8907631692,
"upper_ci": 964.0199135738
},
{
"lower_ci": 953.2564762147,
"model": "gpt-3.5-turbo-1106",
"rank": 90,
"rating": 955.8557887224,
"upper_ci": 959.2854635955
},
{
"lower_ci": 951.2318268117,
"model": "phi-3-mini-4k-instruct",
"rank": 91,
"rating": 955.5964299136,
"upper_ci": 958.4767635868
},
{
"lower_ci": 939.9598426545,
"model": "dolphin-2.2.1-mistral-7b",
"rank": 92,
"rating": 952.2051167375,
"upper_ci": 962.2676381777
},
{
"lower_ci": 945.2431328588,
"model": "solar-10.7b-instruct-v1.0",
"rank": 93,
"rating": 951.1224233106,
"upper_ci": 958.6524028214
},
{
"lower_ci": 948.6812080723,
"model": "llama-2-13b-chat",
"rank": 94,
"rating": 951.0465511917,
"upper_ci": 953.7276982632
},
{
"lower_ci": 940.2487365654,
"model": "wizardlm-13b",
"rank": 95,
"rating": 945.6170139554,
"upper_ci": 952.3153173674
},
{
"lower_ci": 937.3497020159,
"model": "zephyr-7b-beta",
"rank": 96,
"rating": 941.7856240916,
"upper_ci": 945.7193614882
},
{
"lower_ci": 925.3514960917,
"model": "mpt-30b-chat",
"rank": 97,
"rating": 933.4637447927,
"upper_ci": 942.1936901157
},
{
"lower_ci": 928.7552755121,
"model": "pplx-7b-online",
"rank": 98,
"rating": 932.9740226673,
"upper_ci": 940.2890736975
},
{
"lower_ci": 924.8240945839,
"model": "codellama-34b-instruct",
"rank": 99,
"rating": 931.25075507,
"upper_ci": 936.6721012017
},
{
"lower_ci": 919.9794338309,
"model": "zephyr-7b-alpha",
"rank": 100,
"rating": 930.8024303718,
"upper_ci": 940.3904190626
},
{
"lower_ci": 927.6228696967,
"model": "vicuna-13b",
"rank": 101,
"rating": 930.129806663,
"upper_ci": 933.8873904378
},
{
"lower_ci": 916.3079081804,
"model": "codellama-70b-instruct",
"rank": 102,
"rating": 929.7441814853,
"upper_ci": 939.4638589466
},
{
"lower_ci": 919.0275736817,
"model": "gemma-7b-it",
"rank": 103,
"rating": 925.6063141152,
"upper_ci": 931.2321796245
},
{
"lower_ci": 921.4335696267,
"model": "llama-2-7b-chat",
"rank": 104,
"rating": 925.1226871516,
"upper_ci": 927.4211774328
},
{
"lower_ci": 919.914948524,
"model": "phi-3-mini-128k-instruct",
"rank": 105,
"rating": 925.0941128748,
"upper_ci": 928.953409783
},
{
"lower_ci": 917.0649637172,
"model": "qwen-14b-chat",
"rank": 106,
"rating": 923.0565707958,
"upper_ci": 927.5147584562
},
{
"lower_ci": 910.6316722798,
"model": "falcon-180b-chat",
"rank": 107,
"rating": 922.8642075646,
"upper_ci": 938.170426621
},
{
"lower_ci": 912.8204058608,
"model": "guanaco-33b",
"rank": 108,
"rating": 921.0201740806,
"upper_ci": 929.2566479601
},
{
"lower_ci": 903.2450471973,
"model": "gemma-1.1-2b-it",
"rank": 109,
"rating": 908.7494526802,
"upper_ci": 913.4741035764
},
{
"lower_ci": 899.9846491667,
"model": "stripedhyena-nous-7b",
"rank": 110,
"rating": 905.1534113535,
"upper_ci": 912.9382997405
},
{
"lower_ci": 899.9839053334,
"model": "olmo-7b-instruct",
"rank": 111,
"rating": 903.629872537,
"upper_ci": 908.2570306538
},
{
"lower_ci": 890.8398870858,
"model": "mistral-7b-instruct",
"rank": 112,
"rating": 897.2720855909,
"upper_ci": 902.9292885859
},
{
"lower_ci": 887.0047037055,
"model": "vicuna-7b",
"rank": 113,
"rating": 893.825818387,
"upper_ci": 899.7593423485
},
{
"lower_ci": 887.0036385995,
"model": "palm-2",
"rank": 114,
"rating": 891.3072880052,
"upper_ci": 897.1963496269
},
{
"lower_ci": 872.1578125039,
"model": "gemma-2b-it",
"rank": 115,
"rating": 880.4830688196,
"upper_ci": 884.0952002573
},
{
"lower_ci": 868.4676861275,
"model": "qwen1.5-4b-chat",
"rank": 116,
"rating": 877.0664166393,
"upper_ci": 883.3511645246
},
{
"lower_ci": 846.7609834499,
"model": "koala-13b",
"rank": 117,
"rating": 852.5741703315,
"upper_ci": 859.5370102373
},
{
"lower_ci": 833.5366986653,
"model": "chatglm3-6b",
"rank": 118,
"rating": 843.5467457427,
"upper_ci": 853.6979083485
},
{
"lower_ci": 808.672593832,
"model": "gpt4all-13b-snoozy",
"rank": 119,
"rating": 821.0385947815,
"upper_ci": 835.6716204673
},
{
"lower_ci": 804.1197078456,
"model": "chatglm2-6b",
"rank": 120,
"rating": 815.3737962537,
"upper_ci": 821.0889058237
},
{
"lower_ci": 805.0790145038,
"model": "mpt-7b-chat",
"rank": 121,
"rating": 813.8798502443,
"upper_ci": 823.374846381
},
{
"lower_ci": 803.3055604626,
"model": "RWKV-4-Raven-14B",
"rank": 122,
"rating": 810.0683748343,
"upper_ci": 819.7203248755
},
{
"lower_ci": 781.8369679949,
"model": "alpaca-13b",
"rank": 123,
"rating": 789.0910768147,
"upper_ci": 797.3012759671
},
{
"lower_ci": 778.7113024193,
"model": "oasst-pythia-12b",
"rank": 124,
"rating": 782.6327874047,
"upper_ci": 789.0013974908
},
{
"lower_ci": 760.5962982346,
"model": "chatglm-6b",
"rank": 125,
"rating": 767.7699631784,
"upper_ci": 776.2247501319
},
{
"lower_ci": 749.711207841,
"model": "fastchat-t5-3b",
"rank": 126,
"rating": 756.8214872887,
"upper_ci": 762.6617679978
},
{
"lower_ci": 723.1589751426,
"model": "stablelm-tuned-alpha-7b",
"rank": 127,
"rating": 727.6986691281,
"upper_ci": 737.7996673488
},
{
"lower_ci": 701.8523006848,
"model": "dolly-v2-12b",
"rank": 128,
"rating": 710.1023148692,
"upper_ci": 717.2721178625
},
{
"lower_ci": 679.2645194103,
"model": "llama-13b",
"rank": 129,
"rating": 689.1631346499,
"upper_ci": 694.5590720343
}
]

Some files were not shown because too many files have changed in this diff Show More