ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
@@ -0,0 +1,42 @@
|
||||
# الفصل السادس · التفاعل: توسيع فضاء الملاحظة وفضاء الفعل
|
||||
|
||||
> يوسع الإدراك والفعل من النص إلى الصوت والواجهات الرسومية والعالم المادي. ويتناول ثلاثة أنماط للبنية الصوتية، والإدراك الصوتي المتدفق وتركيب الكلام، واستخدام الحاسوب، والتحكم الروبوتي.
|
||||
|
||||
← [العودة إلى الملف التمهيدي الرئيسي](../docs/ar/README.md) · 📖 [قراءة نص الفصل](../book-ar/chapter6.ar.md)
|
||||
|
||||
## كيفية قراءة التجارب
|
||||
|
||||
يستخدم النص هياكل آلية قصيرة لشرح تدفق التحكم؛ ويحتوي دليل التجارب على محولات SDK الكاملة والسجلات والاختبارات وأدلة القبول. لا حاجة لقراءة كل ملف سطرًا سطرًا.
|
||||
|
||||
- **Starter:** ابدأ بالهدف والأمر الأدنى وشروط القبول؛ وابدأ من [live-audio](live-audio/);
|
||||
- **Builder:** تتبّع نقطة الدخول والحلقة الأساسية ومخطط الحالة/الرسائل والأدوات وأداة التحقق.
|
||||
- **Maintainer:** ثم اقرأ الاختبارات وmanifest الأدلة ومعالجة الأعطال ومسارات التراجع ومحولات المزوّد.
|
||||
|
||||
في القراءة الأولى يمكنك تجاوز بيانات الاعتماد وطبقة العرض وتوافق المزوّد؛ عُد إليها عند إعادة إنتاج رقم.
|
||||
|
||||
## المشاريع المصاحبة
|
||||
|
||||
| التجربة | المشروع | النوع | الوصف |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [الوكيل مع مشغل الحدث](agent-with-event-trigger/) | ✅ | وكيل حديث يعتمد على الأحداث تم تصميمه باستخدام FastAPI، ويدمج جميع الأدوات من خوادم MCP الثلاثة الأولى بشكل افتراضي. يستخدم بنية أصلية غير متزامنة لتحميل أداة MCP النظيفة ويستقبل أحداث متعددة المصادر (الويب، والمراسلة الفورية، GitHub، والمؤقتات، وما إلى ذلك) عبر HTTP API. يوفر وثائق API التلقائية (Swagger UI) وقدرات مراقبة الخلفية. |
|
||||
| 6-2 | [وكيل غير متزامن](async-agent/) | ✅ | ينفذ نواة إطار Flux للوكلاء غير المتزامنين القائمين على الأحداث، باستخدام نموذج أحادي الخيط. وتوزع قائمة الأحداث الواردة المهام بحسب السياسة المناسبة، مثل المقاطعة أو التنفيذ الفوري أو الاصطفاف، كما تدعم تشغيل الأدوات بالتوازي، ومقاطعة الجولة الحالية، وإلغاء المهام الطويلة والاستعلام عن حالتها. ويتخذ نموذج لغوي حقيقي القرارات عبر استدعاء الدوال. |
|
||||
| 6-3 | [صوت حي](live-audio/) | ✅ | عرض توضيحي للدردشة الصوتية في الوقت الفعلي يدمج تحويل الكلام إلى نص وحوار الذكاء الاصطناعي وتحويل النص إلى كلام. يدعم العديد من موفري خدمات الذكاء الاصطناعي (OpenAI، OpenRouter، ARK، Siliconflow)، مما يوفر تجربة محادثة منخفضة زمن الوصول. |
|
||||
| Add-on | [وكيل الهاتف](phone-agent/) | 🚧 | نُفّذ مسارا direct/ReAct في SDK الرسمي `pine-voice`، لكن لم يُقدَّم رقم E.164 مخوّل وبموافقة صاحبه. يسجل preflight عدم الاتصال وعدم وجود transcript، ولا يُعد test double قبولًا. |
|
||||
| 6-4 | [خطاب متدفق](streaming-speech/) | ✅ | يوضح المفاضلة الأساسية لإدراك الكلام المتدفق: تقطيع الصوت المستمر إلى مقاطع ذات طول متزايد وإدخالها إلى ASR. ينتج عن كل مقطع مستلم "نتيجة التعرف الجزئي الحالية" لتحقيق زمن استجابة منخفض للغاية للجزء الأول لإخراج النص المبكر. والتكلفة هي أن المقاطع المبكرة، التي تفتقر إلى سياق النصف الأخير من الجملة، قد تكون خاطئة، وتتقارب تدريجياً مع تراكم الصوت. وهذا يتناقض مع النهج عالي الدقة/زمن الوصول العالي المتمثل في "انتظار الجملة بأكملها قبل التعرف عليها". |
|
||||
| 6-5 | [الكلام من نهاية إلى نهاية](end-to-end-speech/) | ✅ | شُغّل MiniCPM-o 4.5 بإصدار مثبّت محليًا على بطاقة RTX PRO 6000 واحدة؛ حقق مسارا end-to-end وself-cascade نتيجة 3/4 مع أخطاء دلالية وشبه لغوية متكاملة، وحُفظ خرج صوتي حقيقي 24kHz ودليل القبول. |
|
||||
| 6-6 | [تحويل النص إلى كلام قابل للتحكم](controllable-tts/) | 🚧 | تمر مكتبة Fish Audio S1 الحقيقية 4×3×2 ووسائط A/B/C بوابات البنية؛ ما زالت دراسة الاستماع النوعية وتقييم «قريب من موظف بشري» ناقصين. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | المستودع الخارجي `anthropics/claude-quickstarts` مثبت عند `9bcc95e…`؛ المقصود هو Computer Use demo بسطح Ubuntu وحلقة Claude agent داخل container، لا مجموعة quickstarts كلها. |
|
||||
| 6-8 | `browser-use/` | 📖 | المستودع الخارجي `browser-use/browser-use` مثبت عند `ec9277c…`؛ يستخدم visual CLI مع `use_vision=True` للبحث في Google عن طقس San Francisco وحفظ مسار الأفعال/الصور. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | تشغيل XLeRobot الحقيقي عن بُعد لمهمة واحدة لترتيب المكتب: ضع الكوب الأحمر في الصينية، والورقة الصفراء في سلة المهملات، ثم أعد الملاحظة وتحقق من الحالة. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | قياس الحد الأعلى للتحكم المثالي في المهمة نفسها داخل المحاكي؛ ولا يعني ذلك تشغيل الروبوت الحقيقي. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | يتحكم Gemini Robotics-ER 1.5 ذاتيًا في XLeRobot الحقيقي لتنفيذ مهمة ترتيب المكتب نفسها. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | مقارنة التنفيذ المفتوح، والتحقق خطوة بخطوة، والحلقة المغلقة التنبؤية للمهمة نفسها داخل المحاكي. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | اختبار RGB عبر البيئات للمهمة نفسها مع تغيير الخلفية ومظهر الأشياء والإضاءة والضجيج البصري. |
|
||||
|
||||
## أنواع المشاريع
|
||||
|
||||
| الأيقونة | النوع | المعنى |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **مستقل** | شفرة كاملة قابلة للتشغيل في هذا المستودع بعد إعداد مفتاح API |
|
||||
| 📖 | **دليل إعادة الإنتاج** | وثائق تفصيلية تعتمد على مستودع خارجي يُجلب باستخدام `git clone` |
|
||||
| 🚧 | **قيد الإنجاز** | يوجد تنفيذ، لكن التشغيل الحي أو التفويض أو العتاد أو أدلة قبول متطلبات النص لم تكتمل |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Chapter 6 · Interaction: Expanding the Observation and Action Spaces
|
||||
|
||||
> Extends perception and action from text to voice, GUI, and the physical world. Three voice paradigms (cascaded/end-to-end full-modal/full-duplex), streaming voice perception and synthesis, Computer Use, and robotic manipulation.
|
||||
|
||||
← [Back to main README](../docs/en/README.md) · 📖 [Read chapter text](../book-en/chapter6.md)
|
||||
|
||||
## How to Read the Experiments
|
||||
|
||||
The prose uses short mechanism skeletons to explain control flow; the experiment directory contains complete SDK adapters, logs, tests, and acceptance evidence. You do not need to read every file line by line.
|
||||
|
||||
- **Starter:** Start with the goal, minimum command, and acceptance conditions; begin with [live-audio](live-audio/);
|
||||
- **Builder:** Follow the entry point, core loop, state/message schema, tools, and verifier.
|
||||
- **Maintainer:** Then read tests, evidence manifests, failure handling, rollback paths, and provider adapters.
|
||||
|
||||
On a first pass, skip credential loading, presentation code, and provider-compatibility layers; return when reproducing a number.
|
||||
|
||||
## Companion Projects
|
||||
|
||||
| Exp. | Project | Type | Description |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | A modern event-driven Agent built with FastAPI, integrating all tools from the first three MCP servers by default. It uses a native asynchronous architecture for clean MCP tool loading and receives multi-source events (Web, Instant Messaging, GitHub, Timers, etc.) via HTTP API. Provides automatic API documentation (Swagger UI) and background monitoring capabilities. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Implement the core of an event-driven asynchronous Agent framework (Flux) based on a single-threaded asyncio model: an inbox event queue dispatches tasks by urgency (interrupt/immediate/queue), supports parallel execution of asynchronous tools, allows interrupting the current turn during execution, and provides cancellation and status querying for simulated long-running tasks. Decision-making is performed by a real LLM (function calling). |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | A real-time voice chat demo integrating speech-to-text, AI dialogue, and text-to-speech. Supports multiple AI service providers (OpenAI, OpenRouter, ARK, Siliconflow), providing a low-latency conversational experience. |
|
||||
| Add-on | [phone-agent](phone-agent/) | ✅ | The retained direct/ReAct campaign runs browser-microphone RTP through real local Whisper, a real external LLM and TTS back over downlink RTP; both arms pass 20/20 gates and independent hash validation. PSTN/E.164 is outside this local WebRTC acceptance scope. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Demonstrates the core trade-off of streaming speech perception: chunk continuous audio into segments of increasing length and feed them to the ASR. Each received segment produces a "current partial recognition result" to achieve extremely low first-chunk latency for early text output. The cost is that early chunks, lacking the context of the latter half of the sentence, may be erroneous, gradually converging as audio accumulates. This contrasts with the high-accuracy/high-latency approach of "waiting for the entire sentence before recognition." |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | A [real local run](end-to-end-speech/validation/runs/exp9-4-minicpmo45-20260801-v1/evidence.json) executed pinned MiniCPM-o 4.5 on one RTX PRO 6000: end-to-end and self-cascade both scored 3/4 with complementary semantic/paralinguistic failures; a real 24kHz speech output and [11/11 acceptance](end-to-end-speech/validation/runs/exp9-4-minicpmo45-20260801-v1/acceptance.json) are retained. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | ✅ | Fish Audio S1 produced the 24-reference library and A/B/C media; a three-pass position-balanced Voxtral listening study rated the multi-reference arm highest and evaluated the near-human claim. The expected C > B > A ordering did not fully reproduce because A outscored B. |
|
||||
| 6-7 | [Anthropic native Computer Use record](claude-computer-use-native/) + `claude-quickstarts/computer-use-demo/` | ✅ | A [validated native run](claude-computer-use-native/validation/runs/exp9-6-anthropic-native-20260803-v2/acceptance.json) built the pinned Dockerfile locally and completed 16 real `claude-sonnet-4-5-20250929` responses plus 15 native `computer` actions. It did not interact with Google reCAPTCHA; visible Open-Meteo JSON grounded the final 70.2°F, clear-sky answer, and every deterministic gate passes. |
|
||||
| 6-8 | [computer-use-open-model](computer-use-open-model/) + `browser-use/` | ✅ | A real open-model visual browser run used `qwen/qwen3-vl-32b-instruct` for 16/16 calls, recovered from a Google CAPTCHA through weather.com, and retained 15 screenshots, the complete action trajectory, grounded answer evidence, and verified hashes. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Real XLeRobot teleoperation for one desk-tidying task: put the red cup in the tray, put the yellow waste paper in the waste bin, then re-observe and verify the state. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Simulator measurement of the ideal-control upper bound for the same desk task; it does not claim that the real robot has run. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 autonomously drives the real XLeRobot on the same desk-tidying task. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Simulator comparison of open-loop, stepwise-checking, and predictive closed-loop strategies for the same task. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | RGB cross-environment test for the same desk task, varying backgrounds, object appearance, lighting, and visual noise. |
|
||||
|
||||
## Project Types
|
||||
|
||||
| Icon | Type | Meaning |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Standalone** | Full code in this repo, runs after configuring API Key |
|
||||
| 📖 | **Reproduction Guide** | Detailed doc depending on **external repos** to `git clone` |
|
||||
| 🚧 | **In Progress** | An implementation exists, but required live execution, authorization, hardware, or manuscript acceptance evidence is incomplete |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Capítulo 6 · Interacción: la expansión de los espacios de observación y de acción
|
||||
|
||||
> Extensión del texto a la voz, GUI y mundo físico: tres paradigmas de voz, Computer Use, robótica
|
||||
|
||||
← [Volver al README principal](../docs/es/README.md) · 📖 [Leer texto del capítulo](../book-es/chapter6.es.md)
|
||||
|
||||
## Cómo leer los experimentos
|
||||
|
||||
El texto usa skeletons breves para explicar el flujo de control; el directorio de experimentos contiene adaptadores SDK completos, registros, pruebas y evidencias de aceptación. No hace falta leer cada archivo línea por línea.
|
||||
|
||||
- **Starter:** Empieza por el objetivo, el comando mínimo y la aceptación; comienza con [live-audio](live-audio/);
|
||||
- **Builder:** Sigue el punto de entrada, el bucle central, el esquema de estado/mensajes, las herramientas y el verificador.
|
||||
- **Maintainer:** Después revisa pruebas, manifiestos, fallos, rollback y adaptadores de proveedores.
|
||||
|
||||
En la primera pasada puedes omitir credenciales, presentación y compatibilidad de proveedores; vuelve al reproducir una cifra.
|
||||
|
||||
## Proyectos Complementarios
|
||||
|
||||
| Exp. | Proyecto | Tipo | Descripción |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | Agente FastAPI orientado a eventos con integración asíncrona de herramientas MCP |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Marco Flux orientado a eventos asíncronos monohilo con colas por prioridad e interrupción |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | La [evidencia real de una ronda](live-audio/backend/validation/real_pipeline_20260729_localwhisper_ark_fish/evidence.json) completa micrófono → Silero VAD → Whisper local → LLM ARK en streaming → Fish S1; los cinco hashes de medios/modelos coinciden, aunque no representa carga concurrente o de producción |
|
||||
| Add-on | [phone-agent](phone-agent/) | ✅ | El proyecto WebRTC local conserva las ejecuciones directa y ReAct con RTP de micrófono del navegador, Whisper local, LLM externo real, TTS y RTP de bajada; ambas pasan 20/20 puertas. PSTN/E.164 queda fuera de este alcance local. La [manifest](phone-agent/validation/runs/exp9-2-webrtc-audio-20260731-v1/manifest.json) conserva el identificador histórico. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | La [aceptación local canónica](streaming-speech/validation/runs/exp9-3-qwen2audio-whisper-provenance-20260730-v3/manifest.json) ejecuta estrictamente prefijos incrementales Qwen2-Audio y VAD de 600 ms + Whisper; 8/8 puertas de ejecución y procedencia pasan, aunque los resultados solo reproducen 2/6 casos |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | MiniCPM-o 4.5, con revisión fijada, se ejecutó localmente en una RTX PRO 6000: end-to-end y self-cascade obtuvieron 3/4 con fallos semánticos/paralingüísticos complementarios; se conservaron audio real de 24kHz y evidencia de aceptación. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | ✅ | Biblioteca real Fish Audio S1 con 4×3×2=24 audios de referencia y medios A/B/C; tres evaluaciones reales ciegas y equilibradas de Voxtral sitúan a C en primer lugar y separan el estado de aceptación de los resultados negativos |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | Corresponde a Anthropic Computer Use Demo, no a toda la colección de *quickstarts*: escritorio Ubuntu en contenedor y bucle de Agent con Computer Use de Claude |
|
||||
| 6-8 | `browser-use/` | 📖 | *Checkout* externo de `browser-use/browser-use`; la tarea abre Google, consulta el clima de San Francisco e inspecciona la trayectoria de acciones del Agent visual |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Teleoperación del XLeRobot real para una misma tarea de ordenar el escritorio: poner la taza roja en la bandeja, el papel amarillo en el cubo de basura y volver a observar para verificar el estado |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Medición en simulador del límite superior de control ideal para la misma tarea; no implica que se haya ejecutado el robot real |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 controla de forma autónoma el XLeRobot real para completar la misma tarea de ordenar el escritorio |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Comparación en simulador de ejecución abierta, comprobación paso a paso y control cerrado predictivo para la misma tarea |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | Prueba RGB entre entornos para la misma tarea, variando fondo, apariencia, iluminación y ruido visual |
|
||||
|
||||
## Tipos de Proyectos
|
||||
|
||||
| Icono | Tipo | Significado |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Autónomo** | Código completo en este repositorio, se ejecuta tras configurar la Clave API |
|
||||
| 📖 | **Guía de Reproducción** | Documento detallado que depende de **repositorios externos** para realizar `git clone` |
|
||||
| 🚧 | **En curso** | Existe una implementación, pero faltan la ejecución real, participantes autorizados, hardware o evidencia de aceptación que exige el texto |
|
||||
@@ -0,0 +1,42 @@
|
||||
# 6. fejezet · Interakció: a megfigyelési és a cselekvési tér kiterjesztése
|
||||
|
||||
> A szövegtől a beszéd, a grafikus felületek és a fizikai világ felé bővíti az érzékelést és a cselekvést: streamelt beszéd, Computer Use és robotika.
|
||||
|
||||
← [Vissza a magyar főoldalhoz](../docs/hu/README.md) · 📖 [A fejezet olvasása](../book-hu/chapter6.md)
|
||||
|
||||
## Hogyan olvassuk a kísérleteket?
|
||||
|
||||
A törzsszöveg rövid mechanizmus-skeletonokkal magyarázza a vezérlési folyamatot; a kísérleti könyvtárakban találhatók a teljes SDK-adapterek, naplók, tesztek és átvételi bizonyítékok. Nem kell minden fájlt sorról sorra elolvasni.
|
||||
|
||||
- **Starter:** Kezdje a céllal, a minimális paranccsal és az átvételi feltételekkel; induljon innen: [live-audio](live-audio/);
|
||||
- **Builder:** Kövesse a belépési pontot, a fő ciklust, az állapot-/üzenetsémát, az eszközöket és az ellenőrzőt.
|
||||
- **Maintainer:** Végül olvassa el a teszteket, a bizonyíték-manifeszteket, a hibakezelést, a visszaállítási útvonalakat és a provider-adaptereket.
|
||||
|
||||
Első olvasáskor átugorható a hitelesítő adatok betöltése, a megjelenítési réteg és a provider-kompatibilitás; a számok reprodukálásakor térjen vissza.
|
||||
|
||||
## Kapcsolódó projektek
|
||||
|
||||
| Kísérlet | Projekt | Típus | Leírás |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | Több eseményforrást kezelő, FastAPI-alapú eseményvezérelt ágenst épít. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Eseménysort, prioritásokat, párhuzamos eszközöket, megszakítást, törlést és feladatállapotot valósít meg. |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | Valós idejű hangbeszélgetési demó, amely STT-t, AI-párbeszédet és TTS-t kapcsol össze. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | A Pine Voice útvonalai elkészültek, de engedélyezett PSTN-hívás még nem futott. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Bemutatja a streamelt beszédfelismerés késleltetési és pontossági kompromisszumát. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | A rögzített revisionű MiniCPM-o 4.5 helyben futott egy RTX PRO 6000 GPU-n; az end-to-end és self-cascade egyaránt 3/4 lett, egymást kiegészítő szemantikai/paralingvisztikai hibákkal és valódi 24kHz-es hangbizonyítékkal. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Fish Audio referencia-könyvtárat és média-összehasonlítást készít; a hallgatási értékelés még hiányos. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | Az Anthropic hivatalos Computer Use demója konténerizált Ubuntu asztalon. |
|
||||
| 6-8 | `browser-use/` | 📖 | Vizuális böngésző-automatizálás művelet- és képernyőkép-nyomvonalakkal. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Valós XLeRobot távvezérlése ugyanazon asztalrendezési feladathoz: a piros csésze a tálcába, a sárga papír a hulladékgyűjtőbe kerül, majd az állapotot újra megfigyeljük és ellenőrizzük. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Ugyanennek a feladatnak az ideális vezérlési felső határa szimulátorban; ez nem jelenti a valódi robot futtatását. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | A Gemini Robotics-ER 1.5 önállóan vezérli a valós XLeRobotot ugyanazon asztalrendezési feladaton. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Nyílt hurkú, lépésenként ellenőrző és prediktív zárt hurkú stratégia összehasonlítása szimulátorban. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | RGB-környezetközi teszt ugyanazon feladaton, eltérő háttérrel, tárgymegjelenéssel, megvilágítással és zajjal. |
|
||||
|
||||
## Projekttípusok
|
||||
|
||||
| Ikon | Típus | Jelentés |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Önálló** | A teljes kód a repository-ban található, és az API-kulcsok beállítása után futtatható. |
|
||||
| 📖 | **Reprodukciós útmutató** | Külső repository vagy meghatározott hardver szükséges. |
|
||||
| 🚧 | **Folyamatban** | Az implementáció vagy az élő elfogadási bizonyíték még nem teljes. |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Bab 6 · Interaksi: Perluasan Ruang Observasi dan Ruang Aksi
|
||||
|
||||
> Memperluas persepsi dan tindakan dari teks ke suara, GUI, dan dunia fisik: streaming speech, Computer Use, serta robotika.
|
||||
|
||||
← [Kembali ke README utama](../docs/id/README.md) · 📖 [Baca bab](../book-id/chapter6.md)
|
||||
|
||||
## Cara Membaca Eksperimen
|
||||
|
||||
Teks utama memakai skeleton mekanisme singkat untuk menjelaskan alur kontrol; direktori eksperimen berisi adapter SDK lengkap, log, pengujian, dan bukti penerimaan. Anda tidak perlu membaca setiap berkas baris demi baris.
|
||||
|
||||
- **Starter:** Mulai dari tujuan, perintah minimum, dan syarat penerimaan; awali dengan [live-audio](live-audio/);
|
||||
- **Builder:** Telusuri titik masuk, loop inti, skema status/pesan, alat, dan verifier.
|
||||
- **Maintainer:** Terakhir, baca pengujian, manifest bukti, penanganan kegagalan, rollback, dan adapter provider.
|
||||
|
||||
Pada pembacaan pertama, lewati kredensial, presentasi, dan kompatibilitas provider; kembali saat mereproduksi angka.
|
||||
|
||||
## Proyek Pendamping
|
||||
|
||||
| Eksperimen | Proyek | Jenis | Deskripsi |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | Membangun Agent event-driven berbasis FastAPI dengan sumber event majemuk. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Mengimplementasikan queue event, prioritas, tool paralel, interupsi, pembatalan, dan status tugas. |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | Demo percakapan suara real-time yang menggabungkan STT, dialog AI, dan TTS. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | Jalur Pine Voice tersedia, tetapi panggilan PSTN berizin belum dijalankan. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Menunjukkan trade-off latensi dan akurasi pada pengenalan suara streaming. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | MiniCPM-o 4.5 pada revision tetap dijalankan secara lokal di satu RTX PRO 6000; end-to-end dan self-cascade sama-sama 3/4 dengan kegagalan semantik/paralinguistik yang saling melengkapi, serta bukti audio 24kHz nyata. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Menyiapkan pustaka referensi Fish Audio dan perbandingan media; evaluasi dengar belum lengkap. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | Demo Computer Use resmi Anthropic pada desktop Ubuntu terkontainerisasi. |
|
||||
| 6-8 | `browser-use/` | 📖 | Otomatisasi browser visual dengan trajectory tindakan dan screenshot. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Teleoperasi XLeRobot nyata untuk satu tugas merapikan meja: masukkan cangkir merah ke nampan, kertas kuning ke tempat sampah, lalu amati dan verifikasi keadaan akhir. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Mengukur batas atas kontrol ideal untuk tugas meja yang sama di simulator; bukan klaim bahwa robot nyata telah dijalankan. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 mengendalikan XLeRobot nyata secara otonom untuk tugas meja yang sama. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Membandingkan strategi open-loop, pemeriksaan bertahap, dan closed-loop prediktif di simulator untuk tugas yang sama. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | Uji RGB lintas lingkungan untuk tugas meja yang sama dengan variasi latar, tampilan objek, pencahayaan, dan noise visual. |
|
||||
|
||||
## Jenis Proyek
|
||||
|
||||
| Ikon | Jenis | Arti |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Mandiri** | Kode lengkap tersedia di repositori dan dapat dijalankan setelah API Key dikonfigurasi. |
|
||||
| 📖 | **Panduan Reproduksi** | Memerlukan repositori eksternal yang harus di-`git clone` atau perangkat keras tertentu. |
|
||||
| 🚧 | **Dalam Proses** | Implementasi atau bukti penerimaan live belum lengkap. |
|
||||
@@ -0,0 +1,42 @@
|
||||
# 第6章 · 交互:観察空間と動作空間の拡張
|
||||
|
||||
> 知覚と行動をテキストから音声、GUI、そして物理世界へと拡張する。3 つの音声パラダイム(カスケード型/エンドツーエンドの全モーダル型/全二重型)、ストリーミング音声の知覚と合成、Computer Use、そしてロボット操作。
|
||||
|
||||
← [メイン README に戻る](../docs/ja/README.md) · 📖 [章の本文を読む](../book-ja/chapter6.ja.md)
|
||||
|
||||
## 実験の読み方
|
||||
|
||||
本文では短い mechanism skeleton で制御フローを説明し、実験ディレクトリには完全な SDK アダプター、ログ、テスト、受け入れ証拠を置きます。すべてのファイルを一行ずつ読む必要はありません。
|
||||
|
||||
- **Starter:** 目的・最小コマンド・受け入れ条件から始め、まず [live-audio](live-audio/);
|
||||
- **Builder:** エントリポイント、中心ループ、状態/メッセージ schema、ツール、検証器を追います。
|
||||
- **Maintainer:** 最後にテスト、証拠 manifest、失敗処理、rollback 経路、provider adapter を読みます。
|
||||
|
||||
初読では認証情報、表示層、provider 互換層を飛ばし、数値を再現するときに戻ってください。
|
||||
|
||||
## 付随プロジェクト
|
||||
|
||||
| 実験 | プロジェクト | 種類 | 説明 |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI で構築された最新のイベント駆動 Agent で、デフォルトで最初の 3 つの MCP サーバーのすべてのツールを統合する。ネイティブな非同期アーキテクチャを用いてクリーンな MCP ツール読み込みを行い、HTTP API を介して複数ソースのイベント(Web、インスタントメッセージング、GitHub、タイマーなど)を受け取る。自動 API ドキュメント(Swagger UI)とバックグラウンド監視機能を提供する。 |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | 単一スレッドの asyncio モデルに基づくイベント駆動非同期 Agent フレームワーク(Flux)の中核を実装する。受信箱イベントキューが緊急度(割り込み/即時/キュー)に応じてタスクをディスパッチし、非同期ツールの並列実行をサポートし、実行中に現在のターンを割り込むことを可能にし、シミュレートされた長時間実行タスクに対するキャンセルと状態照会を提供する。意思決定は実際の LLM(function calling)によって行われる。 |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | 音声認識、AI 対話、音声合成を統合したリアルタイム音声チャットのデモ。複数の AI サービスプロバイダー(OpenAI、OpenRouter、ARK、Siliconflow)をサポートし、低レイテンシの対話体験を提供する。 |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | 公式 `pine-voice` SDK の direct/ReAct 経路は実装済みだが、同意・承認済みの E.164 宛先がない。preflight は発信なし・transcript なしを記録し、test double は受入に数えない。 |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | ストリーミング音声知覚の中核的なトレードオフを示す。連続した音声を徐々に長さを増すセグメントに分割して ASR に供給する。受信した各セグメントは「現在の部分的な認識結果」を生成し、早期のテキスト出力のために極めて低い最初のチャンクのレイテンシを実現する。その代償として、後半の文脈を欠く早期のチャンクは誤る可能性があるが、音声が蓄積するにつれて徐々に収束する。これは「文全体を待ってから認識する」高精度/高レイテンシのアプローチと対照的である。 |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | 固定 revision の MiniCPM-o 4.5 を 1 枚の RTX PRO 6000 で実行。end-to-end と self-cascade はともに 3/4 だが意味・副言語の失敗が相補的で、実際の 24kHz 音声出力と検証証拠を保存した。 |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | 実 Fish Audio S1 の 4×3×2 参照音声庫と A/B/C メディアは構造 gate を通過。定性 listening study と「人間の客服に近い」評価が残る。 |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | 外部 `anthropics/claude-quickstarts` を `9bcc95e…` に固定。本文対象はコンテナ化 Ubuntu desktop+Claude agent loop の Computer Use demo で、quickstarts 全体ではない。 |
|
||||
| 6-8 | `browser-use/` | 📖 | 外部 `browser-use/browser-use` を `ec9277c…` に固定。本文は `use_vision=True` の visual CLI で Google の San Francisco 天気を検索し、action/screenshot 軌跡を保存する。 |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | 実機 XLeRobot を遠隔操作し、同じ机の片付け課題(赤いカップをトレーへ、黄色い紙をごみ箱へ、最後に再観察・検証)を行う。 |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 同じ机の課題について、シミュレータで理想制御の上限を測る。実機を実行したことを意味しない。 |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 で実機 XLeRobot を自律制御し、同じ机の片付け課題を行う。 |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | シミュレータで、同じ課題の開ループ、逐次確認、予測型閉ループを比較する。 |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | 背景、物体の外観、照明、視覚ノイズを変え、同じ課題を RGB 環境間で評価する。 |
|
||||
|
||||
## プロジェクトの種類
|
||||
|
||||
| アイコン | 種類 | 意味 |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **単独実行** | このリポジトリに完全なコードがあり、API キーを設定すれば実行できる |
|
||||
| 📖 | **再現ガイド** | `git clone` が必要な**外部リポジトリ**に依存する詳細ドキュメント |
|
||||
| 🚧 | **進行中** | 実装はあるが、本文が求める live 実行、許可済み参加者、hardware、または受入証拠が未完了 |
|
||||
@@ -0,0 +1,42 @@
|
||||
# 제6장 · 상호작용: 관찰 공간과 행동 공간의 확장
|
||||
|
||||
> 인식과 행동의 범위를 텍스트에서 음성, GUI, 물리 세계로 넓힙니다. 세 가지 음성 패러다임(캐스케이드, 종단 간 옴니모달, 전이중/상호작용형), 스트리밍 음성 인식·합성, Computer Use, 로봇 조작을 다룹니다.
|
||||
|
||||
← [한국어 메인 README로 돌아가기](../docs/ko/README.md) · 📖 [제6장 본문 읽기](../book-ko/chapter6.ko.md)
|
||||
|
||||
## 실험 읽는 방법
|
||||
|
||||
본문은 짧은 메커니즘 skeleton으로 제어 흐름을 설명하고, 실험 디렉터리에는 완전한 SDK 어댑터·로그·테스트·검수 증거를 둡니다. 모든 파일을 줄 단위로 읽을 필요는 없습니다.
|
||||
|
||||
- **Starter:** 목표, 최소 명령, 검수 조건부터 시작하고 다음에서 출발하세요: [live-audio](live-audio/);
|
||||
- **Builder:** 진입점, 핵심 루프, 상태/메시지 스키마, 도구와 verifier를 따라갑니다.
|
||||
- **Maintainer:** 마지막으로 테스트, 증거 manifest, 실패 처리, rollback 경로와 provider adapter를 읽습니다.
|
||||
|
||||
첫 읽기에서는 credential, UI, provider 호환 계층을 건너뛰고 수치를 재현할 때 돌아오세요.
|
||||
|
||||
## 연계 프로젝트
|
||||
|
||||
| 실험 | 프로젝트 | 유형 | 설명 |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI로 만든 현대적인 이벤트 기반 에이전트입니다. 기본 설정으로 앞선 세 MCP 서버의 모든 도구를 통합합니다. 네이티브 비동기 아키텍처로 MCP 도구를 깔끔하게 불러오며, HTTP API를 통해 웹·인스턴트 메시징·GitHub·타이머 등 여러 출처의 이벤트를 받습니다. 자동 API 문서(Swagger UI)와 백그라운드 모니터링 기능도 제공합니다. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | 단일 스레드 asyncio 모델을 바탕으로 이벤트 기반 비동기 에이전트 프레임워크(Flux)의 핵심을 구현합니다. 받은 편지함 이벤트 큐가 긴급도(interrupt/immediate/queue)에 따라 작업을 배분하고, 비동기 도구의 병렬 실행, 실행 중인 턴 중단, 모의 장기 실행 작업의 취소·상태 조회를 지원합니다. 의사결정에는 실제 LLM의 함수 호출을 사용합니다. |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | VAD + ASR(Whisper/SenseVoice) + LLM(GPT-4o/Gemini/Doubao) + TTS(Fish Audio)를 통합한 실시간 음성 채팅으로, WebSocket을 통해 짧은 지연 시간을 제공합니다. |
|
||||
| Add-on | [phone-agent](phone-agent/) | ✅ | 로컬 WebRTC 프로젝트는 브라우저 마이크 RTP, 로컬 Whisper, 실제 외부 LLM, TTS 및 하향 RTP를 사용하는 직접/ReAct 실행을 보존하며 두 경로 모두 20/20 게이트를 통과합니다. PSTN/E.164는 이 로컬 범위에 포함되지 않습니다. [manifest](phone-agent/validation/runs/exp9-2-webrtc-audio-20260731-v1/manifest.json)에 역사적 실행 식별자를 보존합니다. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | 실제 Qwen2-Audio에서 누적되는 음성 접두부 전체를 매번 다시 인코딩해 음향 이벤트를 감지하고 청크별 지연 시간을 측정합니다. 이를 600ms VAD + 오픈 소스 Whisper 조합과 일반·쉼·소음 세 시나리오에서 비교합니다. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | 고정 revision의 MiniCPM-o 4.5를 RTX PRO 6000 한 장에서 실제 로컬 실행했습니다. end-to-end와 self-cascade 모두 3/4였지만 의미/준언어 오류가 상호 보완적이었고, 실제 24kHz 음성 출력과 검증 증거를 보존했습니다. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | 실제 Fish Audio S1의 4×3×2=24개 참조 음성 라이브러리와 A/B/C 미디어가 구조 검사를 통과했습니다. 다만 [검수 결과](controllable-tts/validation/acceptance.json)에는 정성 청취 평가와 ‘사람 상담원에 가까움’이라는 주장에 대한 평가가 아직 없다고 명시되어 있습니다. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | `anthropics/claude-quickstarts`를 `9bcc95e…`에 고정해 사용합니다. 본문이 다루는 것은 전체 quickstarts 모음이 아니라 컨테이너 기반 Ubuntu 데스크톱과 Claude Computer Use 에이전트 루프로 구성된 `computer-use-demo/`입니다. |
|
||||
| 6-8 | `browser-use/` | 📖 | 외부 `browser-use/browser-use` 저장소를 `ec9277c…`에 고정해 사용합니다. 본문 과제에서는 시각 입력을 사용하는 CLI(`use_vision=True`)로 Google에서 샌프란시스코 날씨를 검색하고 동작 및 스크린샷 궤적을 보관합니다. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | 실제 XLeRobot을 원격 조작해 같은 책상 정리 과제를 수행합니다. 빨간 컵은 쟁반에, 노란 폐지는 쓰레기통에 넣고 마지막에 다시 관찰·검증합니다. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 같은 책상 과제의 이상적 제어 상한을 시뮬레이터에서 측정합니다. 실제 로봇 실행을 뜻하지 않습니다. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5가 실제 XLeRobot을 자율 제어해 같은 책상 정리 과제를 수행합니다. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 시뮬레이터에서 같은 과제의 오픈 루프, 단계별 확인, 예측형 폐루프를 비교합니다. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | 배경·물체 외관·조명·시각 노이즈를 바꾸며 같은 과제를 RGB 환경 간에 평가합니다. |
|
||||
|
||||
## 프로젝트 유형
|
||||
|
||||
| 아이콘 | 유형 | 의미 |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **독립 실행** | 전체 코드가 이 저장소에 있으며, API 키를 설정하면 실행할 수 있습니다. |
|
||||
| 📖 | **재현 가이드** | **외부 저장소**를 `git clone`해야 하는 상세 안내 문서입니다. |
|
||||
| 🚧 | **진행 중** | 구현은 있지만, 본문에서 요구하는 실제 실행, 승인된 참여자, 하드웨어 또는 검수 증거가 아직 완전하지 않습니다. |
|
||||
@@ -0,0 +1,175 @@
|
||||
# 第 6 章 · 交互:观察与动作空间的扩展
|
||||
|
||||
> 从模态与时序两个维度扩展 Agent 的观察与动作空间:异步与事件驱动、语音交互、Computer Use 和机器人操作
|
||||
|
||||
← [返回主目录](../README.md) · 📖 [读本章正文](../book/chapter6.md)
|
||||
|
||||
## 如何阅读实验
|
||||
|
||||
正文 skeleton 统一了“持续观察 → 受限动作 → 新观察 → 验收/抢占”的闭环;完整媒体、浏览器和机器人代码分层阅读:
|
||||
|
||||
- **Starter**:从 [live-audio](live-audio/) 的级联入口理解 VAD → ASR → LLM → TTS;
|
||||
- **Builder**:再读 [computer-use-open-model](computer-use-open-model/) 的截图/动作/验证循环,以及 [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) 的五个有边界技能;
|
||||
- **Maintainer**:最后检查取消、不可逆动作门禁、真实观察证据、硬件急停和 sim-to-real 评估。首次可跳过前端样式、模型下载和设备驱动。
|
||||
|
||||
## 配套项目
|
||||
|
||||
| 编号 | 项目 | 类型 | 一句话说明 |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI 事件驱动 Agent,原生异步集成前三组 MCP 工具,通过 HTTP API 接收 Web/IM/GitHub/定时器事件 |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | asyncio 单线程事件驱动框架 Flux:事件队列按紧急度分派、异步工具并行、运行中打断、长任务取消与状态查询 |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | [真实单轮证据](live-audio/backend/validation/real_pipeline_20260729_localwhisper_ark_fish/evidence.json)完成麦克风媒体 → Silero VAD → 本地 Whisper → ARK 流式 LLM → Fish S1;5 个媒体/模型 hash 当前均匹配,但证据本身没有顶层 hash manifest,且不代表并发或生产负载基准 |
|
||||
| 附加 | [phone-agent](phone-agent/) | ✅ | 使用稳定的非编号项目标识;[完整音频 canonical run](phone-agent/validation/runs/phone-agent-webrtc-audio-20260731-v1/manifest.json)跑通直接/ReAct 两组,并保留完整的 WebRTC、ASR、LLM 与 TTS 证据 |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | [canonical 本地验收](streaming-speech/validation/runs/exp6-4-qwen2audio-whisper-provenance-20260730-v3/manifest.json)运行 Qwen2-Audio 递增前缀与 600ms VAD + Whisper:8/8 执行/溯源门禁通过,但预期行为只复现 2/6,实测前缀 8.4–11.3s,不能声称真流式低延迟 |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | [真实本地运行](end-to-end-speech/validation/runs/exp6-5-minicpmo45-20260801-v1/evidence.json)在单张 RTX PRO 6000 上执行固定 revision 的 MiniCPM-o 4.5:端到端与自级联均为 3/4,但语义/副语言失败互补;真实 24kHz 语音输出及 [11/11 验收](end-to-end-speech/validation/runs/exp6-5-minicpmo45-20260801-v1/acceptance.json)已保留 |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | ✅ | 真实 Fish Audio S1 4×3×2=24 条参考音库与 A/B/C 媒体齐全;三次位置平衡的真实 Voxtral 音频盲评中 C 组最高且真人客服感 4.67/5,但 B>A 未复现;[验收](controllable-tts/validation/acceptance.json)将完成状态与负结果分开报告 |
|
||||
| 6-7 | [Anthropic 原生 Computer Use 记录](claude-computer-use-native/) + `claude-quickstarts/computer-use-demo/` | ✅ | [正式运行](claude-computer-use-native/validation/runs/exp6-7-anthropic-native-20260803-v2/acceptance.json)从固定源码本地构建镜像,用 `claude-sonnet-4-5-20250929` 完成 16 次真实响应与 15 个原生 `computer` 动作;Google reCAPTCHA 未交互,转向可见 Open-Meteo JSON 后回答 70.2°F、晴朗,全部确定性门禁通过 |
|
||||
| 6-8 | [computer-use-open-model](computer-use-open-model/) + `browser-use/` | ✅ | [正式开放模型运行](computer-use-open-model/validation/latest.json)使用 `qwen/qwen3-vl-32b-instruct`:Google CAPTCHA 后转 weather.com,16 步完成;16/16 API 响应模型一致、15 张截图、只读动作和答案 grounding 全部通过确定性验收 |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | ✅ | 真机遥操作 XLeRobot 整理桌面:把红色杯子放进托盘、把黄色废纸放进垃圾盒,最后重新观察并确认状态 |
|
||||
| 6-10 | [xlerobot-teleoperation](xlerobot-teleoperation/) | ✅ | 在模拟器中测量同一桌面任务的理想控制上限,不代表真机已经运行 |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | ✅ | 使用 Gemini Robotics-ER 1.5 自主驱动真实 XLeRobot 完成同一整理桌面任务 |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | ✅ | 在模拟器中比较开环、逐步检查和预测式闭环三种同任务策略 |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | ✅ | 对同一桌面任务进行 RGB 跨环境测试,检查视觉策略对背景、外观、光照和噪声变化的适应性 |
|
||||
|
||||
## 实验 6-7 / 6-8 的供应商可移植路径
|
||||
|
||||
6-7 的 Anthropic Demo 是参考实现;6-8 的
|
||||
[开放模型 companion](computer-use-open-model/)把 browser-use 的视觉 Agent 接到
|
||||
OpenAI-compatible Chat Completions:默认示例通过 OpenRouter 调用开放权重
|
||||
`qwen/qwen3-vl-32b-instruct`,也支持读者自己的 vLLM/SGLang 或其他兼容托管端点。
|
||||
“开放模型”指权重/许可证开放,API 网关本身仍可能是商业服务;实验回执必须分别记录
|
||||
requested model 与提供商实际返回的 model ID。
|
||||
|
||||
```bash
|
||||
cd chapter6/computer-use-open-model
|
||||
python3.11 -m venv .venv
|
||||
source .venv/bin/activate
|
||||
python -m pip install -r requirements.txt
|
||||
python -m playwright install chromium
|
||||
|
||||
export OPENROUTER_API_KEY='replace-with-your-key'
|
||||
python main.py --dry-run
|
||||
python main.py \
|
||||
--task "Open Google, search for San Francisco weather today, and report the temperature and conditions. Do not sign in or change any external data." \
|
||||
--max-steps 25 \
|
||||
--record-video
|
||||
```
|
||||
|
||||
自托管时改设 `OPEN_MODEL_API_KEY=local`、`OPEN_MODEL_BASE_URL` 与
|
||||
`OPEN_MODEL_MODEL` 即可。端点必须支持图片输入和结构化 JSON 动作;不支持原生
|
||||
`json_schema` 时可设 `OPEN_MODEL_SCHEMA_MODE=prompt`,但应把这种兼容模式单列为
|
||||
不同实验配置。不同模型的结果不能合并成 Anthropic 复现结果。
|
||||
|
||||
## 实验 6-7 至 6-13 外部复现锚点
|
||||
|
||||
6-7/6-8 的上游 SHA 来自 2026-07-30 工作区 checkout 的 `origin` 与 `HEAD`。6-8 的[开放模型正式运行](computer-use-open-model/validation/latest.json)已经用真实 Qwen3-VL API 与 Chromium 完成 browser-use 路径;6-7 则用 Anthropic 凭据从固定 Dockerfile 本地构建镜像,[运行证据](claude-computer-use-native/validation/runs/exp6-7-anthropic-native-20260803-v2/trajectory.json)记录了 15/25 个动作内绕开 Google reCAPTCHA、读取可见 Open-Meteo 数据并以 `end_turn` 完成任务的过程。早期 401 与两个未通过任务门禁的真实尝试仍作为失败证据保留,不计入正式结果。6-10、6-12 与 6-13 的**本地 GPU 自包含验收已经完成**;6-9、6-11 所需的真实硬件运行仍需单独的设备、授权和安全证据。
|
||||
|
||||
| 实验 | 权威上游 → 本地路径 | 固定提交 | 锁与入口 |
|
||||
| :--: | --- | --- | --- |
|
||||
| 6-7 | [`anthropics/claude-quickstarts`](https://github.com/anthropics/claude-quickstarts) → `chapter6/claude-quickstarts`;具体项目 `computer-use-demo/` | `9bcc95e316e5ef6542b4c9d0469f4078829eead5` | 从该目录的 `Dockerfile` 本地构建;固定源码中的 Dockerfile SHA-256 为 `3aa1f36a491f8f88d81a04c6a89b4cc9f9acd20ad946304c13419736da7c0ead`,但构建输入仍有可变项 |
|
||||
| 6-8 | [`browser-use/browser-use`](https://github.com/browser-use/browser-use) → `chapter6/browser-use`;本书可移植入口 `chapter6/computer-use-open-model/main.py` | `ec9277c5001f2cb78ee419c927775a3cfc227ff8` | checkout 包版本 `0.9.5`;本书入口固定 `use_vision=True`、`max_actions_per_step=1`,默认请求开放权重 Qwen3-VL 32B,并接受任意合格 OpenAI-compatible base URL。该上游提交**没有跟踪 `uv.lock`,且 `.gitignore` 明确忽略它** |
|
||||
| 6-9 | [`Vector-Wangel/XLeRobot`](https://github.com/Vector-Wangel/XLeRobot) → `chapter6/XLeRobot` | `3d14695e40c9c68229c0aacffca6053c75cd3eb6` | `software/examples/{4_xlerobot_teleop_keyboard,5_xlerobot_teleop_xbox,7_xlerobot_teleop_joycon,8_xlerobot_teleop_vr}.py`;精确 blob 与安全门禁见[复现 companion](xlerobot-teleoperation/) |
|
||||
| 6-10 | 同一 [`Vector-Wangel/XLeRobot`](https://github.com/Vector-Wangel/XLeRobot) → `chapter6/XLeRobot`;[`Grigorij-Dudnik/RoboCrew`](https://github.com/Grigorij-Dudnik/RoboCrew) → `chapter6/RoboCrew` | XLeRobot:`3d14695e40c9c68229c0aacffca6053c75cd3eb6`;RoboCrew v0.3.1:`c749148f29bd14e61347f9fc3530c343fff0d994` | XLeRobot 的 `docs/en/source/software/getting_started/LLM_agent.md` + RoboCrew planner;五个桌面操作工具、动作条件世界模型与证据门禁见[复现 companion](gemini-xlerobot-navigation/) |
|
||||
| 6-11 | [`Vector-Wangel/XLeRobot`](https://github.com/Vector-Wangel/XLeRobot) → `chapter6/XLeRobot`;[`Grigorij-Dudnik/RoboCrew`](https://github.com/Grigorij-Dudnik/RoboCrew) → `chapter6/RoboCrew` | XLeRobot:`3d14695e40c9c68229c0aacffca6053c75cd3eb6`;RoboCrew v0.3.1:`c749148f29bd14e61347f9fc3530c343fff0d994` | Gemini Robotics-ER 1.5 自主控制真机的同一整理桌面任务;工具契约与安全边界见[复现 companion](gemini-xlerobot-navigation/) |
|
||||
| 6-12 | `gemini-xlerobot-navigation` 的桌面模拟器 | — | 同一任务的开环、逐步检查和预测式闭环对照;只使用非致动模拟执行器 |
|
||||
| 6-13 | [`StoneT2000/lerobot-sim2real`](https://github.com/StoneT2000/lerobot-sim2real) → `chapter6/lerobot-sim2real` | `87d6c1d969f6e0ca4dc5697940804e231118a63a` | 同一整理桌面任务的 RGB 跨环境测试;阶段与安全边界见[复现 companion](rgb-sim2real-grasping/) |
|
||||
|
||||
6-9 至 6-13 的固定源码获取命令如下;XLeRobot checkout 由 6-9 至 6-12 共用:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/Vector-Wangel/XLeRobot.git chapter6/XLeRobot
|
||||
git -C chapter6/XLeRobot fetch origin 3d14695e40c9c68229c0aacffca6053c75cd3eb6
|
||||
git -C chapter6/XLeRobot checkout --detach 3d14695e40c9c68229c0aacffca6053c75cd3eb6
|
||||
test "$(git -C chapter6/XLeRobot rev-parse HEAD)" = "3d14695e40c9c68229c0aacffca6053c75cd3eb6"
|
||||
|
||||
git clone https://github.com/Grigorij-Dudnik/RoboCrew.git chapter6/RoboCrew
|
||||
git -C chapter6/RoboCrew fetch origin c749148f29bd14e61347f9fc3530c343fff0d994
|
||||
git -C chapter6/RoboCrew checkout --detach c749148f29bd14e61347f9fc3530c343fff0d994
|
||||
test "$(git -C chapter6/RoboCrew rev-parse HEAD)" = "c749148f29bd14e61347f9fc3530c343fff0d994"
|
||||
|
||||
git clone https://github.com/StoneT2000/lerobot-sim2real.git chapter6/lerobot-sim2real
|
||||
git -C chapter6/lerobot-sim2real fetch origin 87d6c1d969f6e0ca4dc5697940804e231118a63a
|
||||
git -C chapter6/lerobot-sim2real checkout --detach 87d6c1d969f6e0ca4dc5697940804e231118a63a
|
||||
test "$(git -C chapter6/lerobot-sim2real rev-parse HEAD)" = "87d6c1d969f6e0ca4dc5697940804e231118a63a"
|
||||
```
|
||||
|
||||
这些命令只建立真实硬件扩展所需的固定源码起点。XLeRobot/RoboCrew/Sim2Real 的 companion 中保存过源码审计或非致动预检,但当前工作区没有这三个源码 checkout;历史预检不等于真机执行。本地 GPU 桌面模拟和 RGB 训练则由各实验目录中的自包含脚本完成,并有独立的证据门禁。
|
||||
|
||||
从仓库根目录复现实验 6-7 的源码版本并本地构建:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/anthropics/claude-quickstarts.git chapter6/claude-quickstarts
|
||||
git -C chapter6/claude-quickstarts checkout --detach 9bcc95e316e5ef6542b4c9d0469f4078829eead5
|
||||
test "$(git -C chapter6/claude-quickstarts rev-parse HEAD)" = "9bcc95e316e5ef6542b4c9d0469f4078829eead5"
|
||||
cd chapter6/claude-quickstarts/computer-use-demo
|
||||
|
||||
RECEIPT_DIR="$HOME/ai-agent-book-receipts/6-7-9bcc95e"
|
||||
mkdir -p "$RECEIPT_DIR"
|
||||
git rev-parse HEAD | tee "$RECEIPT_DIR/source-sha.txt"
|
||||
shasum -a 256 Dockerfile | tee "$RECEIPT_DIR/dockerfile-sha256.txt"
|
||||
docker version | tee "$RECEIPT_DIR/docker-version.txt"
|
||||
|
||||
# 先解析并保存这次构建实际采用的 base-image digest,再禁止 build 重新拉取标签。
|
||||
docker pull ubuntu:22.04 | tee "$RECEIPT_DIR/base-image-pull.txt"
|
||||
docker image inspect ubuntu:22.04 --format '{{json .RepoDigests}}' | tee "$RECEIPT_DIR/base-image-repodigests.json"
|
||||
docker build --pull=false --iidfile "$RECEIPT_DIR/built-image-id.txt" . -t ai-agent-book-computer-use:9bcc95e
|
||||
docker image inspect ai-agent-book-computer-use:9bcc95e --format '{{.Id}}' | tee "$RECEIPT_DIR/built-image-id-inspect.txt"
|
||||
|
||||
export ANTHROPIC_API_KEY='replace-with-your-api-key'
|
||||
docker run --rm -e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" -p 5900:5900 -p 8501:8501 -p 6080:6080 -p 8080:8080 -it ai-agent-book-computer-use:9bcc95e
|
||||
```
|
||||
|
||||
打开 `http://localhost:8080` 后再提交正文任务。除上述构建回执外,还应在同一 `RECEIPT_DIR` 保存原样任务文本、实际模型 ID、按顺序的 computer-use 动作、每步截图/观察、最终回答、停止原因和完成/失败状态;容器能启动不等于实验完成。不要用远端可变标签 `computer-use-demo-latest` 的镜像 ID 代替本地构建回执。
|
||||
|
||||
即使保存了当次 `ubuntu:22.04` digest,该 Dockerfile 仍执行在线 `apt`/PPA 安装,并从未固定 commit 的默认分支克隆 `pyenv`;系统包仓库和若干下载输入也没有内容锁。因此上述回执只能重建“本次究竟运行了什么”的审计链,不能把该镜像声称为位级可重复。
|
||||
|
||||
从仓库根目录复现实验 6-8:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/browser-use/browser-use.git chapter6/browser-use
|
||||
git -C chapter6/browser-use checkout --detach ec9277c5001f2cb78ee419c927775a3cfc227ff8
|
||||
test "$(git -C chapter6/browser-use rev-parse HEAD)" = "ec9277c5001f2cb78ee419c927775a3cfc227ff8"
|
||||
cd chapter6/browser-use
|
||||
|
||||
RECEIPT_DIR="$HOME/ai-agent-book-receipts/6-8-ec9277c"
|
||||
mkdir -p "$RECEIPT_DIR"
|
||||
git rev-parse HEAD | tee "$RECEIPT_DIR/source-sha.txt"
|
||||
uv --version | tee "$RECEIPT_DIR/uv-version.txt"
|
||||
|
||||
# 上游没有提交 uv.lock:先为本次解析生成并保存 lock,之后才可使用 --locked。
|
||||
uv lock
|
||||
cp uv.lock "$RECEIPT_DIR/uv.lock"
|
||||
shasum -a 256 uv.lock | tee "$RECEIPT_DIR/uv-lock-sha256.txt"
|
||||
uv sync --locked
|
||||
uv run browser-use --version | tee "$RECEIPT_DIR/browser-use-version.txt"
|
||||
uvx playwright --version | tee "$RECEIPT_DIR/playwright-version-before-install.txt"
|
||||
uv run browser-use install 2>&1 | tee "$RECEIPT_DIR/browser-install.txt"
|
||||
uvx playwright install --list | tee "$RECEIPT_DIR/playwright-browsers.txt"
|
||||
|
||||
export OPENROUTER_API_KEY='replace-with-your-api-key'
|
||||
export BROWSER_USE_LOGGING_LEVEL=debug
|
||||
uv run python ../computer-use-open-model/main.py \
|
||||
--task "Open Google, search for San Francisco weather today, and report the temperature and conditions. Do not sign in or change any external data." \
|
||||
--output-dir "$RECEIPT_DIR/open-model-run" \
|
||||
--max-steps 25 \
|
||||
--record-video 2>&1 | tee "$RECEIPT_DIR/action-log.txt"
|
||||
|
||||
# 将 debug 日志中实际选择的 executable_path 填到这里;不能只记录“安装过 Chromium”。
|
||||
BROWSER_PATH='/absolute/path/reported-by-LocalBrowserWatchdog'
|
||||
test -x "$BROWSER_PATH"
|
||||
printf '%s\n' "$BROWSER_PATH" | tee "$RECEIPT_DIR/chromium-path.txt"
|
||||
"$BROWSER_PATH" --version | tee "$RECEIPT_DIR/chromium-version.txt"
|
||||
shasum -a 256 "$BROWSER_PATH" | tee "$RECEIPT_DIR/chromium-sha256.txt"
|
||||
```
|
||||
|
||||
本书入口固定 `use_vision=True`、每步最多一个动作并最多运行 25 步;开放模型默认值为 `qwen/qwen3-vl-32b-instruct`,并非 `gpt-4.1`。runner 自动保存提供商响应、逐步截图、动作序列、最终答案、失败状态和 artifact hash;仍需独立核对天气答案与轨迹,不能仅凭模型自己的 `done` 宣称完成。若改用上游 `examples/ui/command_line.py`,它仍默认 `gpt-4.1` 且不会按本书格式自动落盘完整证据。
|
||||
|
||||
这里保存的是**本次本地生成的** `uv.lock`,不是上游锁;初次 `uv lock` 的解析仍受当时包索引影响。`browser-use install` 还会在 Linux 上调用可变的 `uvx playwright install chromium --with-deps --no-shell`,在 macOS/Windows 上调用 `uvx playwright install chromium --no-shell`,因此 Playwright/Chromium 不受项目 lock 约束。固定入口的 `BrowserSession()` 又可能优先选择已有的系统 Chrome,而不是刚下载的 Playwright Chromium;这正是必须记录实际 executable path、版本和二进制哈希的原因。只有把生成的 lock、安装器版本、浏览器二进制和轨迹回执一起归档,才能准确描述当次运行,仍不能把上游 6-8 环境称为位级固定。
|
||||
|
||||
## 项目类型说明
|
||||
|
||||
| 图标 | 类型 | 含义 |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **可独立运行** | 本仓库自带完整代码,配置好 API Key 即可运行 |
|
||||
| 📖 | **复现指南** | 依赖需自行 `git clone` 的**外部仓库**(训练框架、评测基准等) |
|
||||
| 🚧 | **进行中** | 已有实现,但正文要求的真实运行、授权参与者、硬件或验收证据尚未完整 |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Глава 6 · Взаимодействие: расширение пространства наблюдений и пространства действий
|
||||
|
||||
> Расширяет восприятие и действие с текста на голос, GUI и физический мир. Три голосовые парадигмы (каскадная / сквозная полномодальная / полнодуплексная), потоковое восприятие и синтез речи, Computer Use и манипуляции роботов.
|
||||
|
||||
← [К оглавлению](../docs/ru/README.md) · 📖 [Читать главу](../book-ru/chapter6.md)
|
||||
|
||||
## Как читать эксперименты
|
||||
|
||||
В основном тексте короткие скелеты механизмов объясняют поток управления; в каталогах экспериментов находятся полные адаптеры SDK, журналы, тесты и приёмочные доказательства. Читать каждый файл построчно не требуется.
|
||||
|
||||
- **Starter:** Начните с цели, минимальной команды и условий приёмки; начните с [live-audio](live-audio/);
|
||||
- **Builder:** Проследите точку входа, основной цикл, схему состояния/сообщений, инструменты и проверяющий модуль.
|
||||
- **Maintainer:** Затем изучите тесты, манифесты доказательств, обработку сбоев, откат и адаптеры провайдеров.
|
||||
|
||||
При первом чтении можно пропустить ключи, слой представления и совместимость провайдеров; вернитесь при воспроизведении чисел.
|
||||
|
||||
## Сопутствующие проекты
|
||||
|
||||
| Эксп. | Проект | Тип | Описание |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | Современный событийно-управляемый агент на FastAPI, по умолчанию интегрирующий все инструменты первых трёх MCP-серверов. Использует нативную асинхронную архитектуру для чистой загрузки MCP-инструментов и принимает события из разных источников (Web, мессенджеры, GitHub, таймеры и т. д.) через HTTP API. Даёт автоматическую документацию API (Swagger UI) и фоновый мониторинг. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Реализует ядро событийно-управляемого асинхронного фреймворка агента (Flux) на однопоточной модели asyncio: очередь событий-«входящих» распределяет задачи по срочности (прерывание/немедленно/очередь), поддерживает параллельное исполнение асинхронных инструментов, позволяет прерывать текущий раунд во время исполнения и даёт отмену и запрос статуса для имитируемых долгих задач. Решения принимает настоящий LLM (function calling). |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | Демо голосового чата в реальном времени, объединяющее распознавание речи, диалог с ИИ и синтез речи. Поддерживает несколько провайдеров ИИ-сервисов (OpenAI, OpenRouter, ARK, Siliconflow), обеспечивая диалог с низкой задержкой. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | Пути direct/ReAct официального SDK `pine-voice` реализованы, но нет согласованного и авторизованного адресата E.164. Preflight фиксирует отсутствие звонка и transcript; test double не считается приёмкой. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Демонстрирует ключевой компромисс потокового восприятия речи: непрерывное аудио режется на сегменты растущей длины и подаётся в ASR. Каждый полученный сегмент даёт «текущий частичный результат распознавания», обеспечивая крайне низкую задержку первого фрагмента для раннего вывода текста. Плата за это — ранние фрагменты без контекста второй половины фразы могут быть ошибочны, постепенно сходясь по мере накопления аудио. Это контрастирует с высокоточным, но высоколатентным подходом «дождаться всей фразы перед распознаванием». |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | MiniCPM-o 4.5 с закреплённой revision реально запущен локально на одной RTX PRO 6000; end-to-end и self-cascade дали по 3/4 с взаимодополняющими семантическими/паралингвистическими ошибками, сохранены 24kHz-аудио и доказательства приёмки. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Реальная библиотека Fish Audio S1 4×3×2 и A/B/C-медиа проходят структурные gate; не завершены качественное прослушивание и оценка «почти как живой оператор». |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | Внешний `anthropics/claude-quickstarts` зафиксирован на `9bcc95e…`; текст относится к контейнерному Ubuntu desktop+Claude agent loop Computer Use demo, а не ко всему quickstarts. |
|
||||
| 6-8 | `browser-use/` | 📖 | Внешний `browser-use/browser-use` зафиксирован на `ec9277c…`; visual CLI (`use_vision=True`) ищет погоду San Francisco в Google и сохраняет траекторию действий/скриншотов. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Дистанционное управление реальным XLeRobot для одной задачи уборки стола: красную чашку в поднос, жёлтую бумагу в урну, затем повторно наблюдать и проверить состояние. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | В симуляторе измеряется верхняя граница идеального управления для той же задачи; это не означает запуска реального робота. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 автономно управляет реальным XLeRobot в той же задаче уборки стола. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | В симуляторе сравниваются открытый цикл, пошаговая проверка и предиктивный замкнутый цикл для той же задачи. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | RGB-тест между средами для той же задачи с изменением фона, внешнего вида предметов, освещения и визуального шума. |
|
||||
|
||||
## Типы проектов
|
||||
|
||||
| Значок | Тип | Значение |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Автономный** | Полный код в этом репозитории, запускается после настройки API-ключа |
|
||||
| 📖 | **Гайд по воспроизведению** | Подробный документ, зависящий от **внешних репозиториев** через `git clone` |
|
||||
| 🚧 | **В работе** | Реализация есть, но не завершены требуемый живой запуск, авторизация участника, оборудование или доказательства приёмки текста |
|
||||
@@ -0,0 +1,42 @@
|
||||
# அத்தியாயம் 6 · தொடர்பாடல்: அவதானிப்பு மற்றும் செயல் வெளிகளின் விரிவாக்கம்
|
||||
|
||||
> உணர்தல் மற்றும் செயலை உரையிலிருந்து குரல், GUI மற்றும் இயற்கை உலகத்திற்கு விரிவுபடுத்துகிறது. குரலின் மூன்று முன்னுதாரணங்கள் (அடுக்கப்பட்ட / இறுதி-முதல்-இறுதி முழு-வடிவ / முழு-இருவழி), ஸ்ட்ரீமிங் குரல் உணர்தல் மற்றும் தொகுப்பு, Computer Use மற்றும் ரோபோ செயல்பாடு ஆகியவற்றை உள்ளடக்கியது.
|
||||
|
||||
← [முக்கிய README க்குத் திரும்பு](../docs/ta/README.md) · 📖 [அத்தியாய உரையைப் படி](../book-ta/chapter6.ta.md)
|
||||
|
||||
## சோதனைகளை எப்படிப் படிப்பது
|
||||
|
||||
முதன்மை உரை குறுகிய mechanism skeleton-களால் control flow-ஐ விளக்குகிறது; முழு SDK adapters, logs, tests, acceptance evidence ஆகியவை experiment கோப்பகத்தில் உள்ளன. ஒவ்வொரு கோப்பையும் வரி வரியாகப் படிக்க வேண்டியதில்லை.
|
||||
|
||||
- **Starter:** இலக்கு, குறைந்தபட்ச கட்டளை, ஏற்றுக்கொள்ளும் நிபந்தனைகளில் தொடங்குங்கள்; முதலில் [live-audio](live-audio/);
|
||||
- **Builder:** நுழைவுப் புள்ளி, மையச் சுழற்சி, state/message schema, கருவிகள், verifier ஆகியவற்றைப் பின்தொடருங்கள்.
|
||||
- **Maintainer:** பின்னர் tests, evidence manifest, தோல்வி கையாளல், rollback பாதை, provider adapter ஆகியவற்றைப் படியுங்கள்.
|
||||
|
||||
முதல் வாசிப்பில் credentials, UI, provider-compatibility அடுக்குகளைத் தவிர்க்கலாம்; முடிவுகளை மீண்டும் உருவாக்கும்போது திரும்பிப் பாருங்கள்.
|
||||
|
||||
## துணை திட்டங்கள்
|
||||
|
||||
| சோதனை | Project | Type | Description |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI-இல் கட்டப்பட்ட நவீன நிகழ்வு-உந்துதல் ஏஜென்ட்; முதல் மூன்று MCP சேவையகங்களின் அனைத்துக் கருவிகளையும் இயல்புநிலையாக ஒருங்கிணைக்கிறது. தெளிவான MCP கருவி ஏற்றத்திற்காக இயல்பான ஒத்திசையற்ற கட்டமைப்பைப் பயன்படுத்துகிறது; HTTP API வழியாகப் பல-மூல நிகழ்வுகளை (Web, உடனடி செய்தியிடல், GitHub, டைமர்கள் போன்றவை) பெறுகிறது. தானியங்கு API ஆவணம் (Swagger UI) மற்றும் பின்னணிக் கண்காணிப்புத் திறன்களை வழங்குகிறது. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | asyncio ஒற்றை-நூல் அடிப்படையில் நிகழ்வு-உந்துதல் ஒத்திசையற்ற ஏஜென்ட் கட்டமைப்பின் (Flux) மையத்தைச் செயல்படுத்துகிறது: inbox நிகழ்வு வரிசை அவசர நிலையின்படி ஒதுக்குகிறது (குறுக்கிடல்/உடனடி/வரிசையில்), ஒத்திசையற்ற கருவிகளின் இணையான செயலாக்கத்தை ஆதரித்து, இயங்கும் போது தற்போதைய turn-ஐ குறுக்கிட அனுமதித்து, உருவகப்படுத்தப்பட்ட நீண்ட பணிகளுக்கு ரத்து செய்தல் மற்றும் நிலை வினவலையும் வழங்குகிறது. முடிவெடுப்பது உண்மையான LLM (function calling) மூலம் நிறைவேற்றப்படுகிறது. |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | நிகழ்நேர குரல் அரட்டை டெமோ, பேச்சு-முதல்-உரை, AI உரையாடல் மற்றும் உரை-முதல்-பேச்சு செயல்பாடுகளை ஒருங்கிணைக்கிறது. பல AI சேவை வழங்குநர்களை (OpenAI, OpenRouter, ARK, Siliconflow) ஆதரித்து, குறைந்த தாமத உரையாடல் அனுபவத்தை வழங்குகிறது. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | அதிகாரப்பூர்வ `pine-voice` SDK direct/ReAct பாதைகள் உள்ளன; ஆனால் ஒப்புதல் பெற்ற E.164 இலக்கு இல்லை. Preflight dial/transcript இல்லை என்று பதிவு செய்கிறது; test double acceptance அல்ல. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | ஸ்ட்ரீமிங் குரல் உணர்தலின் முக்கிய பரிமாற்றத்தை விளக்குகிறது: தொடர்ச்சியான ஆடியோவை அதிகரிக்கும் நீளத் தொகுதிகளாகப் பிரித்து ASR-க்கு ஊட்டி, ஒவ்வொரு சிறு பகுதி கிடைத்தவுடன் "தற்போதைய பகுதி அங்கீகார முடிவை" உருவாக்கி, மிகக் குறைந்த முதல்-பாக்கெட் தாமதத்துடன் (first-packet latency) உரையை விரைவில் வெளியிடுகிறது; இதன் விலை, பின்னர் வரும் வாக்கியச் சூழல் இல்லாததால் ஆரம்பத் தொகுதிகள் தவறாக இருக்கலாம், ஆடியோ குவியும்போது படிப்படியாக ஒருங்கமைகிறது—"முழு வாக்கியம் வரும் வரை காத்திருந்து அங்கீகரித்தல்" என்ற உயர் துல்லியம்/உயர் தாமத அணுகுமுறைக்கு மாறானதாக உள்ளது. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | நிலைநிறுத்தப்பட்ட revision கொண்ட MiniCPM-o 4.5 ஒரு RTX PRO 6000-இல் உண்மையாக உள்ளூரில் இயக்கப்பட்டது; end-to-end மற்றும் self-cascade இரண்டும் 3/4, மேலும் உண்மையான 24kHz ஒலி வெளியீடும் ஏற்றுக்கொள்ளல் சான்றும் சேமிக்கப்பட்டன. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Real Fish Audio S1 4×3×2 reference library மற்றும் A/B/C media structure gates கடக்கின்றன; qualitative listening study மற்றும் “near-human” மதிப்பீடு இன்னும் இல்லை. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | வெளிப்புற `anthropics/claude-quickstarts` `9bcc95e…`-இல் pin செய்யப்பட்டது; புத்தகத்தில் கேட்கப்பட்ட இலக்கு Ubuntu desktop+Claude agent loop கொண்ட containerized Computer Use demo ஆகும். |
|
||||
| 6-8 | `browser-use/` | 📖 | வெளிப்புற `browser-use/browser-use` `ec9277c…`-இல் pin செய்யப்பட்டது; `use_vision=True` visual CLI Google-ல் San Francisco weather தேடி action/screenshot trajectory-ஐ வைத்திருக்க வேண்டும். |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | ஒரே மேசை ஒழுங்குபடுத்தும் பணிக்கான உண்மையான XLeRobot தொலை இயக்கம்: சிவப்பு கோப்பையைத் தட்டில், மஞ்சள் காகிதத்தை குப்பைத்தொட்டியில் வைத்து, இறுதியில் மீண்டும் கவனித்து நிலையைச் சரிபார்க்கிறது. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | அதே மேசைப் பணிக்கான சிறந்த கட்டுப்பாட்டு உச்சவரம்பை simulator-ல் அளவிடுகிறது; உண்மையான robot ஓட்டப்பட்டது என்று பொருள் இல்லை. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 அதே மேசை ஒழுங்குபடுத்தும் பணியில் உண்மையான XLeRobot-ஐ தன்னாட்சியாக இயக்குகிறது. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | அதே பணிக்கான open-loop, படிப்படியான சரிபார்ப்பு, prediction-based closed-loop ஆகியவற்றை simulator-ல் ஒப்பிடுகிறது. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | பின்னணி, பொருளின் தோற்றம், ஒளி மற்றும் காட்சி noise-ஐ மாற்றி அதே பணியை RGB சூழல்களுக்கு இடையில் சோதிக்கிறது. |
|
||||
|
||||
## திட்ட வகைகள்
|
||||
|
||||
| சின்னம் | வகை | பொருள் |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **தனித்து இயங்கும்** | முழு குறியீடு இந்த களஞ்சியத்தில், API Key உள்ளமைத்தவுடன் இயங்கும் |
|
||||
| 📖 | **மறு உருவாக்க வழிகாட்டி** | **வெளிப்புற களஞ்சியங்களை** `git clone` செய்ய வேண்டிய விரிவான ஆவணம் |
|
||||
| 🚧 | **செயலில் உள்ளது** | Implementation உள்ளது; ஆனால் தேவையான live run, authorization, hardware அல்லது manuscript acceptance evidence முழுமையில்லை |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Bölüm 6 · Etkileşim: Gözlem ve Eylem Uzaylarının Genişletilmesi
|
||||
|
||||
> Algı ve eylemi metinden sese, GUI'ye ve fiziksel dünyaya genişletir. Üç ses paradigması (aşamalı zincir/uçtan uca tam modlu/tam çift yönlü), akış tabanlı ses algısı ve sentezi, Computer Use ve robot manipülasyonu.
|
||||
|
||||
← [Ana README'ye dön](../README.tr.md) · 📖 [Bölüm metnini oku](../book-tr/chapter6.tr.md)
|
||||
|
||||
## Deneyler nasıl okunur
|
||||
|
||||
Metin, kontrol akışını açıklamak için kısa mekanizma skeleton'ları kullanır; deney dizininde tam SDK adaptörleri, günlükler, testler ve kabul kanıtı bulunur. Her dosyayı satır satır okumanız gerekmez.
|
||||
|
||||
- **Starter:** Hedef, en kısa komut ve kabul koşullarıyla başlayın; önce [live-audio](live-audio/);
|
||||
- **Builder:** Giriş noktasını, ana döngüyü, durum/mesaj şemasını, araçları ve doğrulayıcıyı izleyin.
|
||||
- **Maintainer:** Son olarak testleri, kanıt manifestlerini, hata işlemeyi, rollback yollarını ve sağlayıcı adaptörlerini okuyun.
|
||||
|
||||
İlk okumada kimlik bilgisi yükleme, sunum katmanı ve sağlayıcı uyumluluğunu atlayıp sayıları yeniden üretirken dönün.
|
||||
|
||||
## Eşlik Eden Projeler
|
||||
|
||||
| Deney | Proje | Tür | Açıklama |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI ile inşa edilmiş modern bir olay güdümlü Agent; varsayılan olarak ilk üç MCP sunucusundaki tüm araçları entegre eder. Temiz MCP araç yüklemesi için yerel bir asenkron mimari kullanır ve HTTP API üzerinden çok kaynaklı olayları (Web, Anlık Mesajlaşma, GitHub, Zamanlayıcılar vb.) alır. Otomatik API dokümantasyonu (Swagger UI) ve arka plan izleme yetenekleri sunar. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Tek iş parçacıklı bir asyncio modeline dayalı, olay güdümlü asenkron bir Agent çerçevesinin (Flux) çekirdeğini uygular: bir gelen kutusu olay kuyruğu görevleri aciliyete göre (kesme/anında/kuyruk) dağıtır, asenkron araçların paralel yürütülmesini destekler, yürütme sırasında mevcut turun kesilmesine izin verir ve simüle edilmiş uzun süreli görevler için iptal ve durum sorgulama sağlar. Karar verme gerçek bir LLM (fonksiyon çağırma) tarafından yapılır. |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | Konuşmadan metne, AI diyaloğu ve metinden konuşmayı entegre eden gerçek zamanlı bir sesli sohbet demosu. Birden çok AI hizmet sağlayıcısını destekler (OpenAI, OpenRouter, ARK, Siliconflow), düşük gecikmeli bir konuşma deneyimi sunar. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | Resmî `pine-voice` SDK direct/ReAct yolları uygulanmıştır; ancak yetkili ve onay vermiş bir E.164 hedefi yoktur. Preflight arama/transcript olmadığını kaydeder; test double kabul sayılmaz. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Akış tabanlı ses algısının temel ödünleşimini gösterir: sürekli sesi giderek uzayan segmentlere ayırır ve ASR'ye besler. Alınan her segment, erken metin çıktısı için son derece düşük ilk parça gecikmesi sağlamak üzere bir "mevcut kısmi tanıma sonucu" üretir. Bedeli, cümlenin ikinci yarısının bağlamından yoksun olan erken parçaların hatalı olabilmesi, ses biriktikçe kademeli olarak yakınsamasıdır. Bu, "tanımadan önce tüm cümleyi bekleme"nin yüksek doğruluk/yüksek gecikmeli yaklaşımıyla tezat oluşturur. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | Sabit revision'lı MiniCPM-o 4.5 tek RTX PRO 6000 üzerinde gerçekten yerel çalıştırıldı; end-to-end ve self-cascade 3/4 elde etti, tamamlayıcı anlamsal/paralinguistik hatalar ile gerçek 24kHz ses ve kabul kanıtı saklandı. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Gerçek Fish Audio S1 4×3×2 referans kütüphanesi ve A/B/C medya yapısal kapıları geçer; nitel dinleme çalışması ve “insana yakın” değerlendirme eksiktir. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | Harici `anthropics/claude-quickstarts` `9bcc95e…` commit'ine sabitlenmiştir; hedef tüm quickstarts değil, container içindeki Ubuntu desktop+Claude agent loop Computer Use demosudur. |
|
||||
| 6-8 | `browser-use/` | 📖 | Harici `browser-use/browser-use` `ec9277c…` commit'ine sabitlenmiştir; visual CLI (`use_vision=True`) Google'da San Francisco hava durumunu arar ve action/screenshot yörüngesini saklar. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Gerçek XLeRobot teleoperasyonu ile aynı masa toplama görevi: kırmızı bardağı tepsiye, sarı kâğıdı çöp kutusuna koyup durumu yeniden doğrulama. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Simülatörde aynı görevin ideal kontrol üst sınırını ölçer; gerçek robotun çalıştırıldığı anlamına gelmez. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 ile gerçek XLeRobot'u aynı masa toplama görevinde otonom olarak kontrol eder. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Simülatörde aynı görev için açık çevrim, adım adım kontrol ve öngörülü kapalı çevrim stratejilerini karşılaştırır. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | Arka planı, nesne görünümünü, ışığı ve görsel gürültüyü değiştirerek aynı görevde RGB ortamlar arası testi yapar. |
|
||||
|
||||
## Proje Türleri
|
||||
|
||||
| İkon | Tür | Anlamı |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Bağımsız** | Bu depoda tam kod, API Key yapılandırıldıktan sonra çalışır |
|
||||
| 📖 | **Yeniden Üretim Rehberi** | `git clone` ile **harici depolara** bağımlı ayrıntılı belge |
|
||||
| 🚧 | **Devam Ediyor** | Uygulama vardır; ancak gerekli canlı çalıştırma, yetki, donanım veya metin kabul kanıtı eksiktir |
|
||||
@@ -0,0 +1,42 @@
|
||||
# Chương 6 · Tương tác: mở rộng không gian quan sát và không gian hành động
|
||||
|
||||
> mở rộng cảm nhận và hành động từ văn bản sang giọng nói, GUI và thế giới vật lý. Ba mô thức giọng nói (pipeline nối tầng/đa phương thức đầu cuối/full-duplex), cảm nhận và tổng hợp giọng nói dạng streaming, Computer Use và thao tác robot.
|
||||
|
||||
← [Về README chính](../docs/vi/README.md) · 📖 [Đọc nội dung chương](../book-vi/chapter6.vi.md)
|
||||
|
||||
## Cách đọc các thí nghiệm
|
||||
|
||||
Phần văn bản dùng skeleton cơ chế ngắn để giải thích luồng điều khiển; thư mục thí nghiệm chứa adapter SDK đầy đủ, log, kiểm thử và bằng chứng nghiệm thu. Không cần đọc từng tệp theo từng dòng.
|
||||
|
||||
- **Starter:** Bắt đầu từ mục tiêu, lệnh tối thiểu và điều kiện nghiệm thu; hãy bắt đầu với [live-audio](live-audio/);
|
||||
- **Builder:** Lần theo điểm vào, vòng lặp lõi, schema trạng thái/tin nhắn, công cụ và verifier.
|
||||
- **Maintainer:** Sau đó đọc test, manifest bằng chứng, xử lý lỗi, đường rollback và adapter nhà cung cấp.
|
||||
|
||||
Lần đầu có thể bỏ qua credential, lớp trình bày và tương thích provider; quay lại khi cần tái tạo số liệu.
|
||||
|
||||
## Dự án đi kèm
|
||||
|
||||
| Thí nghiệm | Project | Type | Description |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | Agent hướng sự kiện hiện đại xây dựng trên FastAPI, mặc định tích hợp toàn bộ công cụ của ba MCP server phía trên. Dùng kiến trúc bất đồng bộ nguyên sinh để tải công cụ MCP rõ ràng; nhận sự kiện đa nguồn qua HTTP API (Web, tin nhắn tức thời, GitHub, timer, v.v.). Cung cấp tài liệu API tự động (Swagger UI) và khả năng giám sát nền. |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | Triển khai lõi framework Agent bất đồng bộ hướng sự kiện (Flux) dựa trên asyncio một luồng: hàng đợi sự kiện inbox phân phối theo mức khẩn cấp (ngắt/ngay lập tức/xếp hàng), hỗ trợ công cụ bất đồng bộ chạy song song, ngắt turn hiện tại trong lúc đang chạy, đồng thời hủy và truy vấn trạng thái các tác vụ dài mô phỏng. Quyết định được thực hiện bởi LLM thật (function calling). |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | Demo chat giọng nói thời gian thực, tích hợp speech-to-text, hội thoại AI và text-to-speech. Hỗ trợ nhiều nhà cung cấp dịch vụ AI (OpenAI, OpenRouter, ARK, Siliconflow), cung cấp trải nghiệm hội thoại độ trễ thấp. |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | Đã triển khai đường direct/ReAct của SDK `pine-voice` chính thức, nhưng chưa có đích E.164 được ủy quyền và đồng ý. Preflight ghi rõ không quay số/không transcript; test double không phải nghiệm thu. |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | Minh họa đánh đổi cốt lõi của cảm nhận giọng nói streaming: chia âm thanh liên tục thành các khối có độ dài tăng dần đưa vào ASR; mỗi khi nhận một đoạn nhỏ thì xuất “kết quả nhận dạng phần hiện tại” để có văn bản cực sớm với độ trễ gói đầu rất thấp. Cái giá là các khối ban đầu có thể sai do thiếu ngữ cảnh nửa sau câu; khi âm thanh tích lũy, kết quả dần hội tụ, đối chiếu với cách “đợi đủ cả câu rồi nhận dạng” có độ chính xác cao nhưng độ trễ cao. |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | MiniCPM-o 4.5 ở revision cố định đã chạy cục bộ thật trên một RTX PRO 6000; end-to-end và self-cascade cùng đạt 3/4 nhưng lỗi ngữ nghĩa/cận ngôn ngữ bổ sung cho nhau, kèm âm thanh 24kHz và bằng chứng nghiệm thu. |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | Thư viện Fish Audio S1 thật 4×3×2 và media A/B/C đạt cổng cấu trúc; còn thiếu nghiên cứu nghe định tính và đánh giá “gần người thật”. |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | `anthropics/claude-quickstarts` bên ngoài ghim tại `9bcc95e…`; nội dung sách dùng Computer Use demo với desktop Ubuntu+vòng Claude agent trong container, không phải toàn bộ quickstarts. |
|
||||
| 6-8 | `browser-use/` | 📖 | `browser-use/browser-use` bên ngoài ghim tại `ec9277c…`; visual CLI (`use_vision=True`) tìm thời tiết San Francisco trên Google và lưu trajectory action/screenshot. |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | Teleoperation XLeRobot thật cho cùng một nhiệm vụ dọn bàn: đặt cốc đỏ vào khay, giấy vàng vào thùng rác, rồi quan sát lại và xác minh trạng thái. |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Đo giới hạn trên của điều khiển lý tưởng cho cùng nhiệm vụ trong simulator; không có nghĩa robot thật đã được chạy. |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | Gemini Robotics-ER 1.5 tự chủ điều khiển XLeRobot thật để hoàn thành cùng nhiệm vụ dọn bàn. |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | So sánh open-loop, kiểm tra từng bước và closed-loop dự đoán trong simulator cho cùng nhiệm vụ. |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | Kiểm thử RGB xuyên môi trường cho cùng nhiệm vụ với nền, ngoại hình vật thể, ánh sáng và nhiễu thị giác thay đổi. |
|
||||
|
||||
## Phân loại dự án
|
||||
|
||||
| Biểu tượng | Loại | Ý nghĩa |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **Chạy độc lập** | Có mã đầy đủ trong kho, chạy được sau khi cấu hình API Key |
|
||||
| 📖 | **Hướng dẫn tái hiện** | Tài liệu chi tiết, cần `git clone` **kho ngoài** |
|
||||
| 🚧 | **Đang thực hiện** | Đã có triển khai, nhưng còn thiếu chạy live, ủy quyền, phần cứng hoặc bằng chứng nghiệm thu theo nội dung sách |
|
||||
@@ -0,0 +1,42 @@
|
||||
# 第 6 章 · 互動:觀察與動作空間的擴展
|
||||
|
||||
> 從文字擴充套件到語音、GUI、物理世界:語音三典範、Computer Use、機器人
|
||||
|
||||
← [返回主目錄](../docs/zh-TW/README.md) · 📖 [讀本章正文](../book/chapter6.md)
|
||||
|
||||
## 如何閱讀實驗
|
||||
|
||||
正文用短小的機制 skeleton 說明控制流;實驗目錄放完整的 SDK 適配、日誌、測試與驗收證據,不需要逐行讀完每個檔案。
|
||||
|
||||
- **Starter:** 先讀目標、最小指令與驗收條件;可從 [live-audio](live-audio/);
|
||||
- **Builder:** 沿著入口、核心迴圈、狀態/訊息 schema、工具與驗證器閱讀。
|
||||
- **Maintainer:** 最後再看測試、證據 manifest、失敗處理、回滾路徑與 provider adapter。
|
||||
|
||||
第一次閱讀可先跳過憑證載入、展示層和 provider 相容層;要重現數字時再回來查看。
|
||||
|
||||
## 配套專案
|
||||
|
||||
| 編號 | 專案 | 型別 | 一句話說明 |
|
||||
| :--: | --- | :--: | --- |
|
||||
| 6-1 | [agent-with-event-trigger](agent-with-event-trigger/) | ✅ | FastAPI 事件驅動 Agent,原生非同步整合前三組 MCP 工具,透過 HTTP API 接收 Web/IM/GitHub/計時器事件 |
|
||||
| 6-2 | [async-agent](async-agent/) | ✅ | asyncio 單執行緒事件驅動框架 Flux:事件佇列按緊急度分派、非同步工具並行、執行中打斷、長任務取消與狀態查詢 |
|
||||
| 6-3 | [live-audio](live-audio/) | ✅ | 即時語音聊天,整合 VAD + ASR(Whisper/SenseVoice)+ LLM(GPT-4o/Gemini/Doubao)+ TTS(Fish Audio),WebSocket 低延遲 |
|
||||
| Add-on | [phone-agent](phone-agent/) | 🚧 | 官方 `pine-voice` SDK 的 direct/ReAct 路徑已實作,但未提供獲授權且同意參與的 E.164 目的號碼;預檢明確記錄未撥號、無 transcript,test double 不算驗收。 |
|
||||
| 6-4 | [streaming-speech](streaming-speech/) | ✅ | 音訊按遞增長度分塊餵 ASR,每段立刻出文字降首包延遲,對比「整句到齊再識別」的高準確/高延遲 |
|
||||
| 6-5 | [end-to-end-speech](end-to-end-speech/) | ✅ | 已在單張 RTX PRO 6000 上真實本機執行固定 revision 的 MiniCPM-o 4.5;端到端與自級聯皆為 3/4,但語義與副語言錯誤互補,並保留真實 24kHz 語音輸出與完整驗收證據。 |
|
||||
| 6-6 | [controllable-tts](controllable-tts/) | 🚧 | 真實 Fish Audio S1 4×3×2 參考音庫與 A/B/C 媒體通過結構門禁;仍缺定性聽測與「接近真人客服」評估。 |
|
||||
| 6-7 | `claude-quickstarts/computer-use-demo/` | 📖 | 外部 `anthropics/claude-quickstarts` 固定於 `9bcc95e…`;正文對應容器化 Ubuntu 桌面+Claude agent loop 的 Computer Use demo,不是整個 quickstarts。 |
|
||||
| 6-8 | `browser-use/` | 📖 | 外部 `browser-use/browser-use` 固定於 `ec9277c…`;正文用 `use_vision=True` 視覺 CLI 在 Google 查舊金山天氣並保留動作/截圖軌跡。 |
|
||||
| 6-9 | [xlerobot-teleoperation](xlerobot-teleoperation/) | 📖 | 真機 XLeRobot 遙操作同一個整理桌面任務:把紅色杯子放入托盤、黃色廢紙放入垃圾盒,最後重新觀察並確認狀態。 |
|
||||
| 6-10 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 在模擬器中測量同一桌面任務的理想控制上限;不代表真機已經執行。 |
|
||||
| 6-11 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 使用 Gemini Robotics-ER 1.5 自主控制真機 XLeRobot 完成同一整理桌面任務。 |
|
||||
| 6-12 | [gemini-xlerobot-navigation](gemini-xlerobot-navigation/) | 📖 | 在模擬器中比較同一任務的開環、逐步檢查與預測式閉環策略。 |
|
||||
| 6-13 | [rgb-sim2real-grasping](rgb-sim2real-grasping/) | 📖 | 改變背景、物體外觀、光照與視覺雜訊,對同一桌面任務進行 RGB 跨環境測試。 |
|
||||
|
||||
## 專案型別說明
|
||||
|
||||
| 圖示 | 型別 | 含義 |
|
||||
| :--: | --- | --- |
|
||||
| ✅ | **可獨立執行** | 本倉庫自帶完整程式碼,配置好 API Key 即可執行 |
|
||||
| 📖 | **復現指南** | 依賴需自行 `git clone` 的**外部倉庫**(訓練框架、評測基準等) |
|
||||
| 🚧 | **進行中** | 已有實作,但正文要求的真實執行、授權參與者、硬體或驗收證據尚未完整 |
|
||||
@@ -0,0 +1,51 @@
|
||||
# Python
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*$py.class
|
||||
*.so
|
||||
.Python
|
||||
build/
|
||||
develop-eggs/
|
||||
dist/
|
||||
downloads/
|
||||
eggs/
|
||||
.eggs/
|
||||
lib/
|
||||
lib64/
|
||||
parts/
|
||||
sdist/
|
||||
var/
|
||||
wheels/
|
||||
*.egg-info/
|
||||
.installed.cfg
|
||||
*.egg
|
||||
|
||||
# Virtual Environment
|
||||
venv/
|
||||
ENV/
|
||||
env/
|
||||
|
||||
# Environment variables
|
||||
.env
|
||||
|
||||
# Trajectory files
|
||||
*.json
|
||||
!package.json
|
||||
!experiment_protocol.json
|
||||
!validation/
|
||||
!validation/**
|
||||
|
||||
# IDE
|
||||
.vscode/
|
||||
.idea/
|
||||
*.swp
|
||||
*.swo
|
||||
|
||||
# OS
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Demo outputs
|
||||
demo_*.py
|
||||
demo_output/
|
||||
watched_dir/
|
||||
@@ -0,0 +1,384 @@
|
||||
"""
|
||||
Event Client - Send test events to the event-triggered agent
|
||||
"""
|
||||
|
||||
import requests
|
||||
import json
|
||||
import time
|
||||
import argparse
|
||||
from datetime import datetime
|
||||
from event_types import EventType
|
||||
|
||||
|
||||
class EventClient:
|
||||
"""Client to send events to the event-triggered agent server"""
|
||||
|
||||
def __init__(self, server_url: str = "http://localhost:8000"):
|
||||
"""
|
||||
Initialize the client
|
||||
|
||||
Args:
|
||||
server_url: URL of the event server
|
||||
"""
|
||||
self.server_url = server_url.rstrip('/')
|
||||
|
||||
def send_event(self, event_type: str, content: str, metadata: dict = None) -> dict:
|
||||
"""
|
||||
Send an event to the agent
|
||||
|
||||
Args:
|
||||
event_type: Type of event (e.g., 'web_message', 'im_message')
|
||||
content: Content of the event
|
||||
metadata: Additional metadata for the event
|
||||
|
||||
Returns:
|
||||
Response from the server
|
||||
"""
|
||||
event_data = {
|
||||
'event_type': event_type,
|
||||
'content': content,
|
||||
'metadata': metadata or {},
|
||||
'timestamp': datetime.now().isoformat(),
|
||||
'event_id': f"evt_{int(time.time() * 1000)}"
|
||||
}
|
||||
|
||||
print(f"\n{'='*80}")
|
||||
print(f"📤 SENDING EVENT")
|
||||
print(f"{'='*80}")
|
||||
print(f"Event Type: {event_type}")
|
||||
print(f"Content: {content}")
|
||||
if metadata:
|
||||
print(f"Metadata: {json.dumps(metadata, indent=2)}")
|
||||
print(f"{'='*80}\n")
|
||||
|
||||
try:
|
||||
response = requests.post(
|
||||
f"{self.server_url}/event",
|
||||
json=event_data,
|
||||
headers={'Content-Type': 'application/json'},
|
||||
timeout=120
|
||||
)
|
||||
|
||||
response.raise_for_status()
|
||||
result = response.json()
|
||||
|
||||
print(f"\n{'='*80}")
|
||||
print(f"✅ EVENT SENT SUCCESSFULLY")
|
||||
print(f"{'='*80}")
|
||||
print(f"Response: {json.dumps(result, indent=2)}")
|
||||
print(f"{'='*80}\n")
|
||||
|
||||
return result
|
||||
|
||||
except requests.exceptions.RequestException as e:
|
||||
print(f"\n❌ Error sending event: {e}")
|
||||
return {"error": str(e)}
|
||||
|
||||
def reset_agent(self) -> dict:
|
||||
"""Reset the agent state"""
|
||||
try:
|
||||
response = requests.post(f"{self.server_url}/agent/reset", timeout=30)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
def get_status(self) -> dict:
|
||||
"""Get agent status"""
|
||||
try:
|
||||
response = requests.get(f"{self.server_url}/agent/status", timeout=30)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
def start_monitoring(self) -> dict:
|
||||
"""Start system monitoring"""
|
||||
try:
|
||||
response = requests.post(f"{self.server_url}/monitoring/start", timeout=30)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
def stop_monitoring(self) -> dict:
|
||||
"""Stop system monitoring"""
|
||||
try:
|
||||
response = requests.post(f"{self.server_url}/monitoring/stop", timeout=30)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
def register_process(self, process_id: str, name: str) -> dict:
|
||||
"""Register a background process for monitoring"""
|
||||
try:
|
||||
response = requests.post(
|
||||
f"{self.server_url}/process/register",
|
||||
json={'process_id': process_id, 'name': name}, timeout=30
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
def unregister_process(self, process_id: str) -> dict:
|
||||
"""Unregister a background process"""
|
||||
try:
|
||||
response = requests.post(
|
||||
f"{self.server_url}/process/unregister",
|
||||
json={'process_id': process_id}, timeout=30
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
except requests.exceptions.RequestException as e:
|
||||
return {"error": str(e)}
|
||||
|
||||
|
||||
def run_test_scenarios(client: EventClient):
|
||||
"""Run various test scenarios"""
|
||||
|
||||
print("\n" + "🧪"*40)
|
||||
print(" EVENT-TRIGGERED AGENT TEST SCENARIOS")
|
||||
print("🧪"*40 + "\n")
|
||||
|
||||
# Scenario 1: Web message
|
||||
print("\n📋 Scenario 1: Web Interface Message")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.WEB_MESSAGE.value,
|
||||
content="Hello! Can you create a simple Python script that prints 'Hello, World!'?",
|
||||
metadata={"user_id": "user123", "session_id": "session456"}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 2: IM message
|
||||
print("\n📋 Scenario 2: Instant Message")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.IM_MESSAGE.value,
|
||||
content="Can you list the files in the current directory?",
|
||||
metadata={"sender": "Alice", "platform": "Slack"}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 3: Email reply
|
||||
print("\n📋 Scenario 3: Email Reply")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.EMAIL_REPLY.value,
|
||||
content="Thanks for the report! Can you also check the disk usage?",
|
||||
metadata={
|
||||
"from": "bob@example.com",
|
||||
"subject": "Re: System Report",
|
||||
"thread_id": "thread789"
|
||||
}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 4: GitHub PR update
|
||||
print("\n📋 Scenario 4: GitHub PR Review")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.GITHUB_PR_UPDATE.value,
|
||||
content="Review comment: Please add unit tests for the new feature.",
|
||||
metadata={
|
||||
"pr_number": "42",
|
||||
"action": "review_requested",
|
||||
"reviewer": "code-reviewer",
|
||||
"repository": "ai-agent-project"
|
||||
}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 5: Timer trigger
|
||||
print("\n📋 Scenario 5: Scheduled Timer")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.TIMER_TRIGGER.value,
|
||||
content="Daily backup reminder - please check if backups are running correctly.",
|
||||
metadata={
|
||||
"timer_id": "daily_backup_check",
|
||||
"schedule": "daily at 09:00"
|
||||
}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 6: System alert
|
||||
print("\n📋 Scenario 6: System Alert")
|
||||
print("-"*80)
|
||||
client.send_event(
|
||||
event_type=EventType.SYSTEM_ALERT.value,
|
||||
content="Memory usage has exceeded 80%. Please investigate.",
|
||||
metadata={
|
||||
"alert_type": "resource_usage",
|
||||
"severity": "warning",
|
||||
"memory_usage": "82%"
|
||||
}
|
||||
)
|
||||
time.sleep(2)
|
||||
|
||||
# Scenario 7: Register background process
|
||||
print("\n📋 Scenario 7: Background Process Registration")
|
||||
print("-"*80)
|
||||
print("Registering background process...")
|
||||
result = client.register_process("proc_ml_training", "ML Model Training")
|
||||
print(f"Result: {json.dumps(result, indent=2)}")
|
||||
|
||||
# Scenario 8: Start monitoring
|
||||
print("\n📋 Scenario 8: Start System Monitoring")
|
||||
print("-"*80)
|
||||
print("Starting system monitoring (will check for timeouts)...")
|
||||
result = client.start_monitoring()
|
||||
print(f"Result: {json.dumps(result, indent=2)}")
|
||||
print("\n⏰ Monitoring is now active. System will check for:")
|
||||
print(" - User timeout (no interaction for 1 minute)")
|
||||
print(" - Background process timeout (running for 30 seconds)")
|
||||
print("\n💡 Wait 1-2 minutes to see system reminder events trigger automatically...")
|
||||
|
||||
# Get status
|
||||
print("\n📋 Current Agent Status")
|
||||
print("-"*80)
|
||||
status = client.get_status()
|
||||
print(json.dumps(status, indent=2))
|
||||
|
||||
print("\n" + "✅"*40)
|
||||
print(" TEST SCENARIOS COMPLETED")
|
||||
print("✅"*40 + "\n")
|
||||
|
||||
|
||||
def interactive_mode(client: EventClient):
|
||||
"""Interactive mode for sending custom events"""
|
||||
print("\n" + "="*80)
|
||||
print(" INTERACTIVE EVENT CLIENT")
|
||||
print("="*80)
|
||||
print("\nAvailable event types:")
|
||||
for event_type in EventType:
|
||||
print(f" - {event_type.value}")
|
||||
print("\nCommands:")
|
||||
print(" 'status' - Get agent status")
|
||||
print(" 'reset' - Reset agent")
|
||||
print(" 'monitor on' - Start monitoring")
|
||||
print(" 'monitor off' - Stop monitoring")
|
||||
print(" 'quit' - Exit")
|
||||
print("\nOr send an event: <event_type> <content>")
|
||||
|
||||
while True:
|
||||
try:
|
||||
print("\n" + "-"*60)
|
||||
user_input = input("Event > ").strip()
|
||||
|
||||
if not user_input:
|
||||
continue
|
||||
|
||||
if user_input.lower() == 'quit':
|
||||
print("👋 Goodbye!")
|
||||
break
|
||||
|
||||
elif user_input.lower() == 'status':
|
||||
status = client.get_status()
|
||||
print(json.dumps(status, indent=2))
|
||||
|
||||
elif user_input.lower() == 'reset':
|
||||
result = client.reset_agent()
|
||||
print(json.dumps(result, indent=2))
|
||||
|
||||
elif user_input.lower() == 'monitor on':
|
||||
result = client.start_monitoring()
|
||||
print(json.dumps(result, indent=2))
|
||||
|
||||
elif user_input.lower() == 'monitor off':
|
||||
result = client.stop_monitoring()
|
||||
print(json.dumps(result, indent=2))
|
||||
|
||||
else:
|
||||
# Parse event command
|
||||
parts = user_input.split(' ', 1)
|
||||
if len(parts) < 2:
|
||||
print("❌ Invalid format. Use: <event_type> <content>")
|
||||
continue
|
||||
|
||||
event_type = parts[0]
|
||||
content = parts[1]
|
||||
|
||||
# Validate event type
|
||||
try:
|
||||
EventType(event_type)
|
||||
except ValueError:
|
||||
print(f"❌ Invalid event type: {event_type}")
|
||||
continue
|
||||
|
||||
# Send the event
|
||||
client.send_event(event_type, content)
|
||||
|
||||
except KeyboardInterrupt:
|
||||
print("\n\n⚠️ Interrupted. Type 'quit' to exit.")
|
||||
except Exception as e:
|
||||
print(f"\n❌ Error: {str(e)}")
|
||||
|
||||
|
||||
def main():
|
||||
"""Main entry point"""
|
||||
parser = argparse.ArgumentParser(
|
||||
description="事件客户端:向事件驱动 Agent 服务器发送事件。",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog="""示例:
|
||||
python client.py --mode test # 依次发送多种事件,跑通全部场景
|
||||
python client.py --mode interactive # 交互模式,手动输入事件
|
||||
python client.py --message "创建一个 hello world 脚本" # 发送单条 web_message 事件
|
||||
python client.py --event-type timer_trigger --message "检查每日备份" # 指定事件类型
|
||||
""",
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
'--server',
|
||||
default='http://localhost:8000',
|
||||
help='服务器地址(默认:http://localhost:8000)'
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
'--mode',
|
||||
choices=['test', 'interactive'],
|
||||
default='test',
|
||||
help='模式:test(依次发送预置场景事件)或 interactive(交互式手动发送)'
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
'--message',
|
||||
default=None,
|
||||
help='发送单条事件的内容;提供该参数时忽略 --mode,发完即退出'
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
'--event-type',
|
||||
default=EventType.WEB_MESSAGE.value,
|
||||
choices=[e.value for e in EventType],
|
||||
help=f'--message 使用的事件类型(默认:{EventType.WEB_MESSAGE.value})'
|
||||
)
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
client = EventClient(server_url=args.server)
|
||||
|
||||
# Check if server is running
|
||||
try:
|
||||
response = requests.get(f"{args.server}/health", timeout=5)
|
||||
response.raise_for_status()
|
||||
print(f"✅ Connected to server at {args.server}")
|
||||
except requests.exceptions.RequestException as e:
|
||||
print(f"❌ Cannot connect to server at {args.server}")
|
||||
print(f" Error: {e}")
|
||||
print(f"\n💡 Make sure the server is running:")
|
||||
print(f" python server.py")
|
||||
return
|
||||
|
||||
if args.message is not None:
|
||||
client.send_event(event_type=args.event_type, content=args.message)
|
||||
elif args.mode == 'test':
|
||||
run_test_scenarios(client)
|
||||
else:
|
||||
interactive_mode(client)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,38 @@
|
||||
# LLM Provider Configuration (matching conversational_agent.py)
|
||||
# Choose one provider and set the corresponding API key
|
||||
|
||||
# Provider selection (default: kimi)
|
||||
# Options: dashscope (== qwen/bailian), siliconflow, doubao, kimi, moonshot, openrouter
|
||||
LLM_PROVIDER=kimi
|
||||
|
||||
# API Keys (set the one matching your provider)
|
||||
KIMI_API_KEY=your-kimi-api-key-here
|
||||
# DASHSCOPE_API_KEY=your-dashscope-api-key-here
|
||||
# SILICONFLOW_API_KEY=your-siliconflow-api-key-here
|
||||
# DOUBAO_API_KEY=your-doubao-api-key-here
|
||||
# OPENROUTER_API_KEY=your-openrouter-api-key-here
|
||||
|
||||
# Universal OpenRouter fallback:
|
||||
# If the selected provider's key is missing but OPENROUTER_API_KEY is set,
|
||||
# the agent (event_loop_demo.py / server.py / quickstart.py) automatically
|
||||
# falls back to the 'openrouter' provider so it still runs.
|
||||
|
||||
# Optional: Override default model for your provider
|
||||
# LLM_MODEL=kimi-k3
|
||||
|
||||
# Default models per provider:
|
||||
# - siliconflow: Qwen/Qwen3-235B-A22B-Thinking-2507
|
||||
# - doubao: doubao-seed-1-6-thinking-250715
|
||||
# - kimi/moonshot: kimi-k3
|
||||
# - dashscope/qwen/bailian: qwen3.7-plus
|
||||
# - openrouter: google/gemini-3.5-flash
|
||||
# (also supports: openai/gpt-5, anthropic/claude-sonnet-4)
|
||||
|
||||
# Optional: Custom server port (default: 8000)
|
||||
# AGENT_PORT=8000
|
||||
|
||||
# Experiment 6-1 real mailbox/calendar workflow
|
||||
# Obtain both values from the same Unipile project. The DSN normally has the
|
||||
# form apiN.unipile.com:PORT; never commit either real value.
|
||||
# UNIPILE_DSN=apiN.unipile.com:PORT
|
||||
# UNIPILE_ACCESS_TOKEN=your-unipile-api-key
|
||||
@@ -0,0 +1,397 @@
|
||||
"""
|
||||
event_loop_demo.py —— 事件驱动 Agent 的端到端演示(单进程、可离线运行)
|
||||
|
||||
本章"事件驱动的异步 Agent"一节指出:真正的"主动服务"不仅需要 Agent 能定时
|
||||
检查世界,更需要世界能主动通知 Agent。本脚本用最小的代码把这一点跑起来——
|
||||
|
||||
1. 注册若干"事件触发器"(trigger source),每个触发器在后台线程里运行,
|
||||
在事件真正发生的那一刻把一个结构化 Event 推入统一的事件队列:
|
||||
- 一次性定时器 OneShotTimer —— 对应书中 set_timer 的"一次性定时器"
|
||||
- 循环定时器 RecurringTimer —— 对应书中 set_timer 的"循环定时器"
|
||||
- 文件监听 FileWatchTrigger —— 对应 n8n 等平台的文件变更触发器
|
||||
2. 事件循环 EventLoop 从队列里逐个取出事件,唤醒 Agent 处理——这正是
|
||||
"Agent 注册、外部触发"的完整闭环:注册时声明关心什么事件,触发时被异步唤醒。
|
||||
|
||||
与需要起 HTTP 服务器的 server.py / client.py 不同,本脚本在单个进程里同时扮演
|
||||
"外部世界"和"Agent",因此适合用来直观演示事件驱动的行为。
|
||||
|
||||
离线模式(--mock):不调用大模型,用一个"模拟动作"打印 Agent 被唤醒后的处理
|
||||
过程,可在没有 API Key 的环境下观察完整的触发→唤醒→处理闭环。
|
||||
真实模式(默认):接入 EventTriggeredAgent,由大模型真正处理每个事件。
|
||||
|
||||
用法示例:
|
||||
python event_loop_demo.py --mock # 离线演示全部触发器
|
||||
python event_loop_demo.py --mock --trigger timer # 只演示一次性定时器
|
||||
python event_loop_demo.py --mock --trigger recurring --interval 3 --duration 12
|
||||
python event_loop_demo.py --trigger file --watch-dir ./watched # 真实 Agent 处理文件事件
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import queue
|
||||
import logging
|
||||
import argparse
|
||||
import threading
|
||||
from datetime import datetime
|
||||
from typing import Optional, Callable
|
||||
|
||||
from event_types import Event, EventType
|
||||
|
||||
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
||||
logger = logging.getLogger("event_loop")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# 事件触发器(trigger source)
|
||||
# ============================================================================
|
||||
|
||||
class TriggerSource(threading.Thread):
|
||||
"""事件触发器基类:在后台线程中运行,把事件推入共享的事件队列。
|
||||
|
||||
注册(register)体现在实例化并 start();触发(fire)体现在 run() 中
|
||||
满足条件时调用 self.emit(event)。这与书中"注册时由 Agent 主动调用工具、
|
||||
触发时由外部事件异步回调"的两个时刻一一对应。
|
||||
"""
|
||||
|
||||
def __init__(self, name: str, event_queue: "queue.Queue[Event]"):
|
||||
super().__init__(name=name, daemon=True)
|
||||
self.event_queue = event_queue
|
||||
self._stop = threading.Event()
|
||||
|
||||
def emit(self, event: Event):
|
||||
"""触发:把事件推入事件队列,唤醒事件循环。"""
|
||||
logger.info(f"⚡ [{self.name}] 触发事件 -> {event.event_type.value}: {event.content}")
|
||||
self.event_queue.put(event)
|
||||
|
||||
def stop(self):
|
||||
self._stop.set()
|
||||
|
||||
|
||||
class OneShotTimer(TriggerSource):
|
||||
"""一次性定时器:延迟 delay 秒后触发一次 timer_trigger 事件。
|
||||
|
||||
对应书中"用户要求给 DMV 打电话,当前是周六,Agent 设置'下周一上午 10:00
|
||||
致电 DMV'"这类有明确时间点的任务。
|
||||
"""
|
||||
|
||||
def __init__(self, event_queue, delay: float, content: str, timer_id: str = "oneshot"):
|
||||
super().__init__(name=f"OneShotTimer({timer_id})", event_queue=event_queue)
|
||||
self.delay = delay
|
||||
self.content = content
|
||||
self.timer_id = timer_id
|
||||
|
||||
def run(self):
|
||||
logger.info(f"⏱️ [{self.name}] 已注册:{self.delay:.0f} 秒后触发")
|
||||
if self._stop.wait(self.delay):
|
||||
return
|
||||
self.emit(Event(
|
||||
event_type=EventType.TIMER_TRIGGER,
|
||||
content=self.content,
|
||||
metadata={"timer_id": self.timer_id, "kind": "one_shot",
|
||||
"scheduled_delay_seconds": self.delay},
|
||||
))
|
||||
|
||||
|
||||
class RecurringTimer(TriggerSource):
|
||||
"""循环定时器:每隔 interval 秒触发一次 timer_trigger 事件。
|
||||
|
||||
对应书中"每小时检查一次服务器健康状况""每周五发送进展报告",以及
|
||||
OpenClaw Heartbeat 式的定时轮询。
|
||||
"""
|
||||
|
||||
def __init__(self, event_queue, interval: float, content: str, timer_id: str = "recurring"):
|
||||
super().__init__(name=f"RecurringTimer({timer_id})", event_queue=event_queue)
|
||||
self.interval = interval
|
||||
self.content = content
|
||||
self.timer_id = timer_id
|
||||
|
||||
def run(self):
|
||||
logger.info(f"🔁 [{self.name}] 已注册:每 {self.interval:.0f} 秒触发一次")
|
||||
tick = 0
|
||||
while not self._stop.wait(self.interval):
|
||||
tick += 1
|
||||
self.emit(Event(
|
||||
event_type=EventType.TIMER_TRIGGER,
|
||||
content=f"{self.content}(第 {tick} 次)",
|
||||
metadata={"timer_id": self.timer_id, "kind": "recurring",
|
||||
"interval_seconds": self.interval, "tick": tick},
|
||||
))
|
||||
|
||||
|
||||
class FileWatchTrigger(TriggerSource):
|
||||
"""文件监听:轮询目录,发现新增或被修改的文件时触发 file_change 事件。
|
||||
|
||||
对应书中"n8n 等工作流平台的触发器生态:Webhook、定时器、邮件、数据库
|
||||
变更、文件监听"。这里用轮询实现,不依赖第三方库,便于跨平台离线运行。
|
||||
"""
|
||||
|
||||
def __init__(self, event_queue, watch_dir: str, poll_interval: float = 1.0):
|
||||
super().__init__(name=f"FileWatch({watch_dir})", event_queue=event_queue)
|
||||
self.watch_dir = watch_dir
|
||||
self.poll_interval = poll_interval
|
||||
self._snapshot = {}
|
||||
|
||||
def _scan(self):
|
||||
snapshot = {}
|
||||
try:
|
||||
for entry in os.scandir(self.watch_dir):
|
||||
if entry.is_file():
|
||||
snapshot[entry.name] = entry.stat().st_mtime
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
return snapshot
|
||||
|
||||
def run(self):
|
||||
os.makedirs(self.watch_dir, exist_ok=True)
|
||||
self._snapshot = self._scan()
|
||||
logger.info(f"👀 [{self.name}] 已注册:轮询间隔 {self.poll_interval:.0f} 秒"
|
||||
f"(当前已有 {len(self._snapshot)} 个文件)")
|
||||
while not self._stop.wait(self.poll_interval):
|
||||
current = self._scan()
|
||||
for name, mtime in current.items():
|
||||
if name not in self._snapshot:
|
||||
change = "created"
|
||||
elif mtime != self._snapshot[name]:
|
||||
change = "modified"
|
||||
else:
|
||||
continue
|
||||
self.emit(Event(
|
||||
event_type=EventType.FILE_CHANGE,
|
||||
content=f"检测到文件{'新增' if change == 'created' else '修改'},请查看其内容并给出简要处理建议。",
|
||||
metadata={"path": os.path.join(self.watch_dir, name), "change": change},
|
||||
))
|
||||
self._snapshot = current
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# 事件循环(event loop)
|
||||
# ============================================================================
|
||||
|
||||
class EventLoop:
|
||||
"""统一事件队列 + 单线程分发。
|
||||
|
||||
所有触发器把异构事件推入同一个队列;事件循环按到达顺序取出,每个事件
|
||||
唤醒一次 Agent 处理。这正是书中"将所有输入统一建模为事件流,通过事件
|
||||
循环驱动 Agent 的思考和行动"的最小实现。
|
||||
"""
|
||||
|
||||
def __init__(self, dispatch: Callable[[Event], None]):
|
||||
self.event_queue: "queue.Queue[Event]" = queue.Queue()
|
||||
self.dispatch = dispatch
|
||||
self.triggers = []
|
||||
self.processed = 0
|
||||
|
||||
def add_trigger(self, trigger: TriggerSource):
|
||||
self.triggers.append(trigger)
|
||||
|
||||
def run(self, duration: float):
|
||||
"""启动所有触发器,运行 duration 秒后停止。"""
|
||||
deadline = time.monotonic() + duration
|
||||
for t in self.triggers:
|
||||
t.start()
|
||||
|
||||
logger.info(f"🟢 事件循环启动,将运行 {duration:.0f} 秒,等待事件唤醒 Agent...\n")
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
event = self.event_queue.get(timeout=0.5)
|
||||
except queue.Empty:
|
||||
continue
|
||||
self.processed += 1
|
||||
logger.info(f"\n{'='*80}\n📥 事件循环取出第 {self.processed} 个事件"
|
||||
f" -> 唤醒 Agent\n{'='*80}")
|
||||
try:
|
||||
self.dispatch(event)
|
||||
except Exception as e: # noqa: BLE001 - 演示中不希望单个事件异常终止循环
|
||||
logger.error(f"❌ 处理事件时出错: {e}")
|
||||
|
||||
for t in self.triggers:
|
||||
t.stop()
|
||||
logger.info(f"\n🔴 事件循环结束,共处理 {self.processed} 个事件。")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# 分发处理器:模拟动作 or 真实 Agent
|
||||
# ============================================================================
|
||||
|
||||
def make_mock_dispatch() -> Callable[[Event], None]:
|
||||
"""离线模拟处理器:不调用大模型,打印 Agent 被唤醒后的处理过程。"""
|
||||
|
||||
def dispatch(event: Event):
|
||||
logger.info(f"🤖 Agent 被唤醒,收到消息: {event.to_user_message()}")
|
||||
# 用一个确定性的"模拟动作"代替大模型 + 工具调用
|
||||
if event.event_type == EventType.TIMER_TRIGGER:
|
||||
action = "读取定时任务上下文 -> 执行例行检查 -> 汇报结果"
|
||||
elif event.event_type == EventType.FILE_CHANGE:
|
||||
path = event.metadata.get("path", "")
|
||||
preview = ""
|
||||
try:
|
||||
with open(path, "r", encoding="utf-8", errors="replace") as f:
|
||||
preview = f.read(120).replace("\n", " ")
|
||||
except OSError:
|
||||
preview = "(无法读取文件内容)"
|
||||
action = f"读取文件 {os.path.basename(path)} -> 内容预览: {preview!r} -> 生成处理建议"
|
||||
else:
|
||||
action = "解析事件 -> 调用相关工具 -> 生成处理结果"
|
||||
logger.info(f"🛠️ [模拟动作] {action}")
|
||||
logger.info(f"✅ Agent 处理完成: 已响应 {event.event_type.value} 事件")
|
||||
|
||||
return dispatch
|
||||
|
||||
|
||||
def make_agent_dispatch(provider: str, model: Optional[str],
|
||||
max_iterations: int) -> Callable[[Event], None]:
|
||||
"""真实处理器:接入 EventTriggeredAgent,由大模型处理每个事件。"""
|
||||
from agent import EventTriggeredAgent, SystemHintConfig, resolve_provider_and_key
|
||||
|
||||
# 通用兜底:直连 provider 的 key 缺失时,若有 OPENROUTER_API_KEY 则自动改走 openrouter。
|
||||
resolved_provider, api_key = resolve_provider_and_key(provider)
|
||||
if not api_key:
|
||||
print(f"❌ 未检测到 provider '{provider}' 对应的 API Key(也未配置 OPENROUTER_API_KEY 兜底)。")
|
||||
print(f" 请先设置环境变量,或改用离线演示:python event_loop_demo.py --mock")
|
||||
sys.exit(1)
|
||||
if resolved_provider != provider:
|
||||
print(f"ℹ️ provider '{provider}' 无可用 Key,已自动改用 OpenRouter 兜底(openrouter)。")
|
||||
provider = resolved_provider
|
||||
# 保留已是 provider/model 形式的显式模型;否则让 openrouter 用其默认模型。
|
||||
model = model if (model and "/" in model) else None
|
||||
|
||||
config = SystemHintConfig(
|
||||
enable_timestamps=True,
|
||||
enable_tool_counter=True,
|
||||
enable_todo_list=True,
|
||||
enable_detailed_errors=True,
|
||||
enable_system_state=True,
|
||||
save_trajectory=True,
|
||||
trajectory_file="event_loop_trajectory.json",
|
||||
temperature=0.7,
|
||||
max_tokens=4096,
|
||||
use_mcp_servers=False, # 本演示仅用内置工具,避免额外的 MCP 依赖
|
||||
)
|
||||
agent = EventTriggeredAgent(api_key=api_key, provider=provider,
|
||||
model=model, config=config, verbose=True)
|
||||
logger.info(f"✅ 真实 Agent 初始化完成(provider={provider}, model={agent.model})")
|
||||
|
||||
def dispatch(event: Event):
|
||||
result = agent.handle_event(event, max_iterations=max_iterations)
|
||||
logger.info(f"✅ Agent 处理完成: success={result['success']}, "
|
||||
f"iterations={result['iterations']}, "
|
||||
f"tool_calls={len(result['tool_calls'])}")
|
||||
|
||||
return dispatch
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# CLI
|
||||
# ============================================================================
|
||||
|
||||
def build_parser() -> argparse.ArgumentParser:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="事件驱动 Agent 端到端演示:注册触发器,由外部事件异步唤醒 Agent。",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog="""示例:
|
||||
python event_loop_demo.py --mock
|
||||
离线演示全部触发器(一次性定时器 + 循环定时器 + 文件监听),无需 API Key
|
||||
python event_loop_demo.py --mock --trigger timer
|
||||
只演示一次性定时器
|
||||
python event_loop_demo.py --mock --trigger recurring --interval 3 --duration 12
|
||||
每 3 秒触发一次循环定时器,共运行 12 秒
|
||||
python event_loop_demo.py --mock --trigger file --watch-dir ./watched
|
||||
监听 ./watched 目录,向其中写入文件即可触发事件
|
||||
python event_loop_demo.py --trigger timer --provider kimi
|
||||
用真实大模型处理一次性定时器事件(需要设置对应的 API Key)
|
||||
""",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--trigger", choices=["timer", "recurring", "file", "all"], default="all",
|
||||
help="要演示的触发器类型:timer=一次性定时器,recurring=循环定时器,"
|
||||
"file=文件监听,all=全部(默认:all)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--mock", action="store_true",
|
||||
help="离线模式:不调用大模型,用模拟动作演示触发→唤醒→处理闭环(无需 API Key)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--duration", type=float, default=12.0,
|
||||
help="事件循环总运行时长(秒),到时后停止所有触发器(默认:12)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--delay", type=float, default=3.0,
|
||||
help="一次性定时器的延迟触发时间(秒)(默认:3)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--interval", type=float, default=4.0,
|
||||
help="循环定时器的触发间隔(秒)(默认:4)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--watch-dir", default="watched_dir",
|
||||
help="文件监听触发器监视的目录,不存在会自动创建(默认:watched_dir)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--provider", default=os.getenv("LLM_PROVIDER", "kimi"),
|
||||
choices=["dashscope", "qwen", "bailian", "siliconflow", "doubao", "kimi", "moonshot", "openrouter"],
|
||||
help="真实模式使用的大模型提供商(默认:环境变量 LLM_PROVIDER 或 kimi)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--model", default=os.getenv("LLM_MODEL"),
|
||||
help="真实模式的模型名覆盖(默认:使用提供商默认模型)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--max-iterations", type=int, default=10,
|
||||
help="真实模式下单个事件的最大工具调用轮数(默认:10)",
|
||||
)
|
||||
return parser
|
||||
|
||||
|
||||
def main():
|
||||
args = build_parser().parse_args()
|
||||
|
||||
print("\n" + "=" * 80)
|
||||
print("🚀 事件驱动 Agent 演示(EVENT-DRIVEN AGENT DEMO)")
|
||||
print("=" * 80)
|
||||
print(f"触发器: {args.trigger} | 模式: {'离线模拟' if args.mock else '真实 Agent'} | "
|
||||
f"时长: {args.duration:.0f}s")
|
||||
print("=" * 80 + "\n")
|
||||
sys.stdout.flush()
|
||||
|
||||
if args.mock:
|
||||
dispatch = make_mock_dispatch()
|
||||
else:
|
||||
dispatch = make_agent_dispatch(args.provider, args.model, args.max_iterations)
|
||||
|
||||
loop = EventLoop(dispatch)
|
||||
|
||||
if args.trigger in ("timer", "all"):
|
||||
loop.add_trigger(OneShotTimer(
|
||||
loop.event_queue, delay=args.delay, timer_id="daily_backup_check",
|
||||
content="一次性定时器到期:请检查每日备份是否已经完成。",
|
||||
))
|
||||
if args.trigger in ("recurring", "all"):
|
||||
loop.add_trigger(RecurringTimer(
|
||||
loop.event_queue, interval=args.interval, timer_id="health_check",
|
||||
content="循环定时器到期:请检查服务器健康状况。",
|
||||
))
|
||||
if args.trigger in ("file", "all"):
|
||||
loop.add_trigger(FileWatchTrigger(loop.event_queue, watch_dir=args.watch_dir))
|
||||
print(f"💡 提示:向目录 {args.watch_dir}/ 写入或修改文件即可触发 file_change 事件。")
|
||||
print(f" 例如另开一个终端执行:echo hello > {args.watch_dir}/note.txt\n")
|
||||
sys.stdout.flush()
|
||||
|
||||
if not loop.triggers:
|
||||
print("❌ 没有可运行的触发器。")
|
||||
sys.exit(1)
|
||||
|
||||
try:
|
||||
loop.run(duration=args.duration)
|
||||
except KeyboardInterrupt:
|
||||
print("\n⚠️ 收到中断信号,正在停止...")
|
||||
for t in loop.triggers:
|
||||
t.stop()
|
||||
|
||||
print("\n" + "=" * 80)
|
||||
print(f"📊 演示结束:共处理 {loop.processed} 个事件。")
|
||||
print("=" * 80 + "\n")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,99 @@
|
||||
"""
|
||||
Event types for the event-triggered agent system
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import datetime
|
||||
from enum import Enum
|
||||
from typing import Dict, Any, Optional
|
||||
|
||||
|
||||
class EventType(Enum):
|
||||
"""Types of events that can trigger agent actions"""
|
||||
# External input events
|
||||
WEB_MESSAGE = "web_message"
|
||||
IM_MESSAGE = "im_message"
|
||||
EMAIL_REPLY = "email_reply"
|
||||
GITHUB_PR_UPDATE = "github_pr_update"
|
||||
TIMER_TRIGGER = "timer_trigger"
|
||||
FILE_CHANGE = "file_change"
|
||||
|
||||
# System reminder events
|
||||
USER_TIMEOUT = "user_timeout"
|
||||
PROCESS_TIMEOUT = "process_timeout"
|
||||
SYSTEM_ALERT = "system_alert"
|
||||
|
||||
|
||||
@dataclass
|
||||
class Event:
|
||||
"""Represents an event that triggers agent action"""
|
||||
event_type: EventType
|
||||
content: str
|
||||
metadata: Dict[str, Any] = field(default_factory=dict)
|
||||
timestamp: str = field(default_factory=lambda: datetime.now().isoformat())
|
||||
event_id: Optional[str] = None
|
||||
|
||||
def to_user_message(self) -> str:
|
||||
"""Convert event to user message format for the agent"""
|
||||
if self.event_type == EventType.WEB_MESSAGE:
|
||||
return f"[Web Interface] {self.content}"
|
||||
|
||||
elif self.event_type == EventType.IM_MESSAGE:
|
||||
sender = self.metadata.get('sender', 'Unknown')
|
||||
return f"[IM from {sender}] {self.content}"
|
||||
|
||||
elif self.event_type == EventType.EMAIL_REPLY:
|
||||
from_email = self.metadata.get('from', 'Unknown')
|
||||
subject = self.metadata.get('subject', 'No Subject')
|
||||
return f"[Email Reply from {from_email}]\nSubject: {subject}\n{self.content}"
|
||||
|
||||
elif self.event_type == EventType.GITHUB_PR_UPDATE:
|
||||
pr_number = self.metadata.get('pr_number', 'Unknown')
|
||||
action = self.metadata.get('action', 'updated')
|
||||
return f"[GitHub PR #{pr_number} {action}] {self.content}"
|
||||
|
||||
elif self.event_type == EventType.TIMER_TRIGGER:
|
||||
timer_id = self.metadata.get('timer_id', 'Unknown')
|
||||
return f"[Timer {timer_id} triggered] {self.content}"
|
||||
|
||||
elif self.event_type == EventType.FILE_CHANGE:
|
||||
path = self.metadata.get('path', 'Unknown')
|
||||
change = self.metadata.get('change', 'modified')
|
||||
return f"[File {change}: {path}] {self.content}"
|
||||
|
||||
elif self.event_type == EventType.USER_TIMEOUT:
|
||||
duration = self.metadata.get('duration', 'unknown')
|
||||
return f"[System Reminder] User has not responded for {duration}. {self.content}"
|
||||
|
||||
elif self.event_type == EventType.PROCESS_TIMEOUT:
|
||||
process_id = self.metadata.get('process_id', 'Unknown')
|
||||
duration = self.metadata.get('duration', 'unknown')
|
||||
return f"[System Alert] Background process {process_id} has been running for {duration}. {self.content}"
|
||||
|
||||
elif self.event_type == EventType.SYSTEM_ALERT:
|
||||
alert_type = self.metadata.get('alert_type', 'general')
|
||||
return f"[System Alert: {alert_type}] {self.content}"
|
||||
|
||||
else:
|
||||
return f"[{self.event_type.value}] {self.content}"
|
||||
|
||||
def to_dict(self) -> Dict[str, Any]:
|
||||
"""Convert event to dictionary for JSON serialization"""
|
||||
return {
|
||||
'event_type': self.event_type.value,
|
||||
'content': self.content,
|
||||
'metadata': self.metadata,
|
||||
'timestamp': self.timestamp,
|
||||
'event_id': self.event_id
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, data: Dict[str, Any]) -> 'Event':
|
||||
"""Create event from dictionary"""
|
||||
return cls(
|
||||
event_type=EventType(data['event_type']),
|
||||
content=data['content'],
|
||||
metadata=data.get('metadata', {}),
|
||||
timestamp=data.get('timestamp', datetime.now().isoformat()),
|
||||
event_id=data.get('event_id')
|
||||
)
|
||||
@@ -0,0 +1,103 @@
|
||||
"""
|
||||
Example demonstrating the Event-Triggered Agent with MCP tools
|
||||
"""
|
||||
|
||||
import os
|
||||
import asyncio
|
||||
from dotenv import load_dotenv
|
||||
from agent import EventTriggeredAgent, SystemHintConfig, resolve_provider_and_key
|
||||
from event_types import Event, EventType
|
||||
|
||||
# Load environment variables
|
||||
load_dotenv()
|
||||
|
||||
|
||||
async def main():
|
||||
"""Main example function"""
|
||||
print("=" * 80)
|
||||
print("Event-Triggered Agent with MCP Tools Example")
|
||||
print("=" * 80)
|
||||
print()
|
||||
|
||||
# Get API credentials
|
||||
provider = os.getenv("LLM_PROVIDER", "kimi")
|
||||
provider, api_key = resolve_provider_and_key(provider)
|
||||
|
||||
if not api_key:
|
||||
print("❌ Please set the provider API key in your .env file (DASHSCOPE_API_KEY for dashscope/qwen/bailian)")
|
||||
return
|
||||
|
||||
# Create agent configuration
|
||||
config = SystemHintConfig(
|
||||
enable_timestamps=True,
|
||||
enable_tool_counter=True,
|
||||
enable_todo_list=True,
|
||||
enable_detailed_errors=True,
|
||||
enable_system_state=True,
|
||||
save_trajectory=True,
|
||||
trajectory_file="example_trajectory.json",
|
||||
use_mcp_servers=True # Enable MCP servers
|
||||
)
|
||||
|
||||
# Initialize agent
|
||||
print("Initializing agent...")
|
||||
agent = EventTriggeredAgent(
|
||||
api_key=api_key,
|
||||
provider=provider,
|
||||
config=config,
|
||||
verbose=True
|
||||
)
|
||||
|
||||
# Load MCP tools
|
||||
print("\nLoading MCP tools...")
|
||||
await agent.load_mcp_tools()
|
||||
|
||||
print("\n" + "=" * 80)
|
||||
print("Testing Event Processing")
|
||||
print("=" * 80)
|
||||
print()
|
||||
|
||||
# Create a test event
|
||||
event = Event(
|
||||
event_type=EventType.WEB_MESSAGE,
|
||||
content="Search the web for 'Python async programming best practices' and summarize the top 3 results.",
|
||||
metadata={
|
||||
"source": "web_interface",
|
||||
"user_id": "demo_user",
|
||||
"session_id": "test_session_001"
|
||||
}
|
||||
)
|
||||
|
||||
# Handle the event
|
||||
try:
|
||||
result = agent.handle_event(event, max_iterations=15)
|
||||
|
||||
print("\n" + "=" * 80)
|
||||
print("Result Summary")
|
||||
print("=" * 80)
|
||||
print(f"Success: {result['success']}")
|
||||
print(f"Iterations: {result['iterations']}")
|
||||
print(f"Tool Calls: {len(result['tool_calls'])}")
|
||||
|
||||
if result.get('final_answer'):
|
||||
print(f"\nFinal Answer:\n{result['final_answer']}")
|
||||
|
||||
if result.get('trajectory_file'):
|
||||
print(f"\nTrajectory saved to: {result['trajectory_file']}")
|
||||
|
||||
except Exception as e:
|
||||
print(f"\n❌ Error processing event: {e}")
|
||||
import traceback
|
||||
traceback.print_exc()
|
||||
finally:
|
||||
# Cleanup MCP connections
|
||||
print("\nCleaning up MCP connections...")
|
||||
await agent.mcp_manager.disconnect_all()
|
||||
print("✅ Cleanup complete")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
try:
|
||||
asyncio.run(main())
|
||||
except KeyboardInterrupt:
|
||||
print("\n\n⚠️ Interrupted by user")
|
||||
@@ -0,0 +1,46 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"title": "Real event-driven mailbox workflow",
|
||||
"authority": "book/chapter4.md:466",
|
||||
"mail_provider": "Unipile Email API",
|
||||
"listener": {
|
||||
"mode": "polling",
|
||||
"endpoint": "GET /api/v1/emails",
|
||||
"event_channel": "unipile_mailbox_poll",
|
||||
"queue": "FIFO by provider timestamp then email id"
|
||||
},
|
||||
"scenarios": [
|
||||
{
|
||||
"classification": "meeting_invitation",
|
||||
"required_actions": ["live_calendar_conflict_check", "accept_or_decline_draft"]
|
||||
},
|
||||
{
|
||||
"classification": "customer_complaint",
|
||||
"required_actions": ["key_information_extraction", "high_priority_notification"]
|
||||
},
|
||||
{
|
||||
"classification": "marketing",
|
||||
"required_actions": ["provider_archive_update", "post_update_verification"]
|
||||
}
|
||||
],
|
||||
"official_schema_sources": [
|
||||
"https://developer.unipile.com/reference/accountscontroller_listaccounts.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_listmails.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_getmail.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_updatemail.md",
|
||||
"https://developer.unipile.com/reference/folderscontroller_listfolders.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendars.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendareventsbycalendar.md",
|
||||
"https://developer.unipile.com/docs/new-emails-webhook.md"
|
||||
],
|
||||
"acceptance": {
|
||||
"no_local_or_mock_mailbox_substitute": true,
|
||||
"three_real_inbound_email_objects": true,
|
||||
"calendar_query_receipted": true,
|
||||
"draft_artifact_hashed": true,
|
||||
"high_priority_notification_delivered": true,
|
||||
"marketing_email_archived_and_verified": true,
|
||||
"identifiers_and_credentials_redacted": true,
|
||||
"fail_closed_on_missing_or_invalid_credentials": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,129 @@
|
||||
"""
|
||||
Quick Start Script for Event-Triggered Agent
|
||||
Demonstrates the basic functionality in a simple way
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import subprocess
|
||||
import signal
|
||||
from event_types import EventType
|
||||
|
||||
# Check if API key is set (universal OpenRouter fallback applied by the server).
|
||||
from agent import resolve_provider_and_key
|
||||
|
||||
provider = os.getenv("LLM_PROVIDER", "kimi").lower()
|
||||
resolved_provider, api_key = resolve_provider_and_key(provider)
|
||||
|
||||
if not api_key:
|
||||
print(f"❌ Error: no API key for provider '{provider}', and no OPENROUTER_API_KEY fallback")
|
||||
print(f"\nPlease set one of:")
|
||||
print(f" export DASHSCOPE_API_KEY='...' # for dashscope/qwen/bailian")
|
||||
print(f" export KIMI_API_KEY='...' # or SILICONFLOW/DOUBAO/OPENROUTER per provider")
|
||||
print(f" export OPENROUTER_API_KEY='...' # universal fallback")
|
||||
print(f"\nOr change provider:")
|
||||
print(f" export LLM_PROVIDER=dashscope # or qwen, bailian, siliconflow, doubao, kimi, openrouter")
|
||||
sys.exit(1)
|
||||
|
||||
if resolved_provider != provider:
|
||||
print(f"ℹ️ provider '{provider}' has no key; the server will fall back to OpenRouter.")
|
||||
|
||||
print("\n" + "="*80)
|
||||
print("🚀 EVENT-TRIGGERED AGENT QUICK START")
|
||||
print("="*80)
|
||||
print()
|
||||
|
||||
# Check if server is already running
|
||||
import requests
|
||||
try:
|
||||
response = requests.get("http://localhost:8000/health", timeout=2)
|
||||
print("✅ Server is already running!")
|
||||
print("\n💡 You can now use the client to send events:")
|
||||
print(" python client.py --mode test")
|
||||
print(" python client.py --mode interactive")
|
||||
sys.exit(0)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
print("📦 Starting the event-triggered agent server...")
|
||||
print("\n⏳ This may take a moment to initialize...\n")
|
||||
|
||||
# Start the server in a subprocess
|
||||
try:
|
||||
server_process = subprocess.Popen(
|
||||
[sys.executable, "server.py"],
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.STDOUT,
|
||||
universal_newlines=True,
|
||||
bufsize=1
|
||||
)
|
||||
|
||||
# Wait for server to start
|
||||
print("⏰ Waiting for server to start...")
|
||||
max_wait = 30
|
||||
for i in range(max_wait):
|
||||
try:
|
||||
response = requests.get("http://localhost:8000/health", timeout=1)
|
||||
if response.status_code == 200:
|
||||
print("✅ Server is running!\n")
|
||||
break
|
||||
except Exception:
|
||||
pass
|
||||
time.sleep(1)
|
||||
if i % 5 == 0:
|
||||
print(f" Still waiting... ({i}/{max_wait}s)")
|
||||
else:
|
||||
print("❌ Server failed to start in time")
|
||||
server_process.terminate()
|
||||
sys.exit(1)
|
||||
|
||||
print("="*80)
|
||||
print("🎉 QUICK START READY!")
|
||||
print("="*80)
|
||||
print()
|
||||
print("The event-triggered agent server is now running on port 8000.")
|
||||
print()
|
||||
print("📋 What you can do now:")
|
||||
print()
|
||||
print("1. Send test events (in another terminal):")
|
||||
print(" python client.py --mode test")
|
||||
print()
|
||||
print("2. Use interactive mode:")
|
||||
print(" python client.py --mode interactive")
|
||||
print()
|
||||
print("3. Send individual events via API:")
|
||||
print(" curl -X POST http://localhost:8000/event \\")
|
||||
print(" -H 'Content-Type: application/json' \\")
|
||||
print(" -d '{\"event_type\": \"web_message\", \"content\": \"Hello!\"}'")
|
||||
print()
|
||||
print("4. Check agent status:")
|
||||
print(" curl http://localhost:8000/agent/status")
|
||||
print()
|
||||
print("="*80)
|
||||
print("📺 Server output will appear below:")
|
||||
print("="*80)
|
||||
print()
|
||||
|
||||
# Stream server output
|
||||
try:
|
||||
while True:
|
||||
line = server_process.stdout.readline()
|
||||
if not line:
|
||||
break
|
||||
print(line, end='')
|
||||
except KeyboardInterrupt:
|
||||
print("\n\n⚠️ Shutting down server...")
|
||||
server_process.send_signal(signal.SIGINT)
|
||||
server_process.wait(timeout=5)
|
||||
print("✅ Server stopped")
|
||||
|
||||
except FileNotFoundError:
|
||||
print("❌ Error: Could not find server.py")
|
||||
print("Make sure you're in the agent-with-event-trigger directory")
|
||||
sys.exit(1)
|
||||
except Exception as e:
|
||||
print(f"❌ Error: {e}")
|
||||
import traceback
|
||||
traceback.print_exc()
|
||||
sys.exit(1)
|
||||
@@ -0,0 +1,8 @@
|
||||
openai>=1.3.0
|
||||
requests>=2.31.0
|
||||
python-dotenv>=1.0.0
|
||||
flask>=3.0.0
|
||||
fastapi>=0.104.0
|
||||
uvicorn[standard]>=0.24.0
|
||||
mcp>=1.0.0
|
||||
httpx>=0.27.0
|
||||
@@ -0,0 +1,467 @@
|
||||
"""
|
||||
Event Server - FastAPI version with native async support for MCP tools
|
||||
"""
|
||||
|
||||
import os
|
||||
import logging
|
||||
from datetime import datetime, timedelta
|
||||
from typing import Dict, Any, Optional
|
||||
from contextlib import asynccontextmanager
|
||||
from fastapi import FastAPI, HTTPException, BackgroundTasks
|
||||
from pydantic import BaseModel
|
||||
from agent import EventTriggeredAgent, SystemHintConfig, resolve_provider_and_key
|
||||
from event_types import Event, EventType
|
||||
import threading
|
||||
import time
|
||||
import asyncio
|
||||
import argparse
|
||||
import uvicorn
|
||||
|
||||
|
||||
def _env_int(name: str, default: int) -> int:
|
||||
"""Read an integer env var; fall back to default (with a warning) if malformed."""
|
||||
raw = os.getenv(name)
|
||||
if raw is None:
|
||||
return default
|
||||
try:
|
||||
return int(raw)
|
||||
except ValueError:
|
||||
logger.warning(f"Invalid {name} value: {raw!r} (must be an integer); using default {default}")
|
||||
return default
|
||||
|
||||
|
||||
def _reasoning_safe_temperature(model, requested=1.0):
|
||||
"""Reasoning models (Kimi K3, GPT-5, ...) only accept temperature=1.
|
||||
Return 1 for those; otherwise the requested value so non-reasoning
|
||||
providers (Doubao, DeepSeek, older Moonshot) are unchanged."""
|
||||
m = str(model or "").lower().replace("/", "-")
|
||||
return 1 if ("kimi-k3" in m or "gpt-5" in m) else requested
|
||||
|
||||
|
||||
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Global agent instance
|
||||
agent: Optional[EventTriggeredAgent] = None
|
||||
agent_lock = threading.Lock()
|
||||
|
||||
# Monitoring state
|
||||
monitoring_enabled = False
|
||||
monitoring_thread: Optional[threading.Thread] = None
|
||||
|
||||
# MCP loading status
|
||||
mcp_loading_status = {
|
||||
"loading": False,
|
||||
"loaded": False,
|
||||
"tools_count": 0,
|
||||
"error": None,
|
||||
"started_at": None,
|
||||
"completed_at": None
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# FastAPI Lifecycle Events (Modern lifespan)
|
||||
# ============================================================================
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
"""Lifespan context manager for startup and shutdown"""
|
||||
global agent, monitoring_enabled
|
||||
|
||||
# Startup
|
||||
logger.info("🚀 Starting Event-Triggered Agent Server (FastAPI)")
|
||||
await init_agent()
|
||||
logger.info("✅ Server ready to receive events\n")
|
||||
|
||||
yield
|
||||
|
||||
# Shutdown
|
||||
logger.info("Shutting down server...")
|
||||
monitoring_enabled = False
|
||||
|
||||
if agent and agent.mcp_manager:
|
||||
await agent.mcp_manager.disconnect_all()
|
||||
|
||||
logger.info("✅ Server shutdown complete")
|
||||
|
||||
|
||||
# Initialize FastAPI app with lifespan
|
||||
app = FastAPI(
|
||||
title="Event-Triggered Agent Server",
|
||||
description="AI Agent with async MCP tools support",
|
||||
version="2.0.0",
|
||||
lifespan=lifespan
|
||||
)
|
||||
|
||||
|
||||
# Pydantic models for requests
|
||||
class EventRequest(BaseModel):
|
||||
event_type: str
|
||||
content: str
|
||||
metadata: Optional[Dict[str, Any]] = None
|
||||
|
||||
|
||||
class ProcessRegister(BaseModel):
|
||||
process_id: str
|
||||
process_name: str
|
||||
metadata: Optional[Dict[str, Any]] = None
|
||||
|
||||
|
||||
class ProcessUnregister(BaseModel):
|
||||
process_id: str
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Initialization
|
||||
# ============================================================================
|
||||
|
||||
async def init_agent():
|
||||
"""Initialize the agent with optional MCP tools"""
|
||||
global agent, mcp_loading_status
|
||||
|
||||
# Determine provider from environment (universal OpenRouter fallback applied)
|
||||
requested_provider = os.getenv("LLM_PROVIDER", "kimi").lower()
|
||||
provider, api_key = resolve_provider_and_key(requested_provider)
|
||||
|
||||
if not api_key:
|
||||
raise ValueError(
|
||||
f"API key not set for provider '{requested_provider}'. Set the appropriate "
|
||||
f"environment variable, or set OPENROUTER_API_KEY as a universal fallback."
|
||||
)
|
||||
|
||||
# Get model from environment if specified
|
||||
model = os.getenv("LLM_MODEL")
|
||||
if provider != requested_provider:
|
||||
logger.info(
|
||||
f"ℹ️ provider '{requested_provider}' has no key; falling back to OpenRouter."
|
||||
)
|
||||
# Keep an explicit provider/model id; otherwise use OpenRouter's default.
|
||||
if not (model and "/" in model):
|
||||
model = None
|
||||
|
||||
# Check if MCP should be enabled (default: true)
|
||||
enable_mcp = os.getenv("ENABLE_MCP_TOOLS", "true").lower() not in ["false", "0", "no"]
|
||||
|
||||
config = SystemHintConfig(
|
||||
enable_timestamps=True,
|
||||
enable_tool_counter=True,
|
||||
enable_todo_list=True,
|
||||
enable_detailed_errors=True,
|
||||
enable_system_state=True,
|
||||
save_trajectory=True,
|
||||
trajectory_file="event_agent_trajectory.json",
|
||||
temperature=_reasoning_safe_temperature(model, 0.7),
|
||||
max_tokens=4096,
|
||||
use_mcp_servers=enable_mcp
|
||||
)
|
||||
|
||||
agent = EventTriggeredAgent(
|
||||
api_key=api_key,
|
||||
provider=provider,
|
||||
model=model,
|
||||
config=config,
|
||||
verbose=True
|
||||
)
|
||||
|
||||
logger.info(f"✅ Agent initialized with {provider} provider")
|
||||
|
||||
if enable_mcp:
|
||||
logger.info("🔄 MCP tools enabled (default) - loading asynchronously...")
|
||||
await load_mcp_tools_async()
|
||||
else:
|
||||
logger.info(f"📦 Using built-in tools only (MCP disabled via ENABLE_MCP_TOOLS=false)")
|
||||
|
||||
|
||||
async def load_mcp_tools_async():
|
||||
"""Load MCP tools asynchronously"""
|
||||
global agent, mcp_loading_status
|
||||
|
||||
mcp_loading_status["loading"] = True
|
||||
mcp_loading_status["started_at"] = datetime.now().isoformat()
|
||||
|
||||
try:
|
||||
if agent:
|
||||
await agent.load_mcp_tools()
|
||||
tools_count = len(agent.mcp_manager.tools)
|
||||
|
||||
mcp_loading_status["loaded"] = True
|
||||
mcp_loading_status["loading"] = False
|
||||
mcp_loading_status["tools_count"] = tools_count
|
||||
mcp_loading_status["completed_at"] = datetime.now().isoformat()
|
||||
|
||||
logger.info(f"✅ MCP tools loaded: {tools_count} tools available")
|
||||
if tools_count > 0:
|
||||
sample_tools = list(agent.mcp_manager.tools.keys())[:5]
|
||||
logger.info(f" Sample: {sample_tools}")
|
||||
else:
|
||||
raise RuntimeError("Agent not initialized")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"❌ Failed to load MCP tools: {e}")
|
||||
mcp_loading_status["loading"] = False
|
||||
mcp_loading_status["loaded"] = False
|
||||
mcp_loading_status["error"] = str(e)
|
||||
mcp_loading_status["completed_at"] = datetime.now().isoformat()
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# API Endpoints
|
||||
# ============================================================================
|
||||
|
||||
@app.get("/")
|
||||
async def root():
|
||||
"""Root endpoint with API information"""
|
||||
return {
|
||||
"service": "Event-Triggered Agent Server",
|
||||
"version": "2.0.0",
|
||||
"status": "running",
|
||||
"docs": "/docs",
|
||||
"endpoints": {
|
||||
"health": "GET /health",
|
||||
"event": "POST /event",
|
||||
"mcp_status": "GET /mcp/status",
|
||||
"mcp_reload": "POST /mcp/reload",
|
||||
"agent_status": "GET /agent/status",
|
||||
"agent_reset": "POST /agent/reset"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@app.get("/health")
|
||||
async def health_check():
|
||||
"""Health check endpoint"""
|
||||
return {
|
||||
"status": "healthy",
|
||||
"agent_initialized": agent is not None,
|
||||
"monitoring_enabled": monitoring_enabled,
|
||||
"mcp_enabled": agent.config.use_mcp_servers if agent else False,
|
||||
"mcp_loaded": mcp_loading_status["loaded"],
|
||||
"timestamp": datetime.now().isoformat()
|
||||
}
|
||||
|
||||
|
||||
@app.get("/mcp/status")
|
||||
async def get_mcp_status():
|
||||
"""Get MCP tools loading status"""
|
||||
status = mcp_loading_status.copy()
|
||||
|
||||
# Add tool list if loaded
|
||||
if status["loaded"] and agent:
|
||||
status["tools"] = list(agent.mcp_manager.tools.keys())
|
||||
|
||||
# Group by server
|
||||
status["tools_by_server"] = {}
|
||||
for tool_name in agent.mcp_manager.tools.keys():
|
||||
server = tool_name.split("_")[0]
|
||||
if server not in status["tools_by_server"]:
|
||||
status["tools_by_server"][server] = []
|
||||
status["tools_by_server"][server].append(tool_name)
|
||||
|
||||
return status
|
||||
|
||||
|
||||
@app.post("/mcp/reload")
|
||||
async def reload_mcp_tools(background_tasks: BackgroundTasks):
|
||||
"""Manually trigger MCP tools reload"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
if mcp_loading_status["loading"]:
|
||||
raise HTTPException(status_code=409, detail="MCP tools are already loading")
|
||||
|
||||
# Use FastAPI background tasks
|
||||
background_tasks.add_task(load_mcp_tools_async)
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": "MCP tools reload started in background"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/event")
|
||||
async def handle_event(event_req: EventRequest):
|
||||
"""Handle incoming event"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
try:
|
||||
# Create event
|
||||
event_data = {
|
||||
"event_type": event_req.event_type,
|
||||
"content": event_req.content,
|
||||
"metadata": event_req.metadata or {}
|
||||
}
|
||||
event = Event.from_dict(event_data)
|
||||
|
||||
# Handle the event
|
||||
with agent_lock:
|
||||
result = agent.handle_event(event, max_iterations=20)
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"event_id": event.event_id,
|
||||
"result": {
|
||||
"final_answer": result.get('final_answer'),
|
||||
"iterations": result.get('iterations'),
|
||||
"tool_calls_count": len(result.get('tool_calls') or []),
|
||||
"todo_items": len(result.get('todo_list') or []),
|
||||
"success": result.get('success', False),
|
||||
"trajectory_file": result.get('trajectory_file')
|
||||
}
|
||||
}
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Error handling event: {e}")
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@app.get("/agent/status")
|
||||
async def get_agent_status():
|
||||
"""Get current agent status"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
return {
|
||||
"provider": agent.provider,
|
||||
"model": agent.model,
|
||||
"tool_calls_count": len(agent.tool_calls),
|
||||
"todo_items": len(agent.todo_list),
|
||||
"current_directory": agent.current_directory,
|
||||
"mcp_tools_loaded": agent.mcp_tools_loaded,
|
||||
"mcp_tools_count": len(agent.mcp_manager.tools) if agent.mcp_tools_loaded else 0
|
||||
}
|
||||
|
||||
|
||||
@app.post("/agent/reset")
|
||||
async def reset_agent():
|
||||
"""Reset agent state"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
agent.reset()
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": "Agent state reset successfully"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/process/register")
|
||||
async def register_process(process: ProcessRegister):
|
||||
"""Register a background process for monitoring"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
agent.background_processes[process.process_id] = {
|
||||
"name": process.process_name,
|
||||
"start_time": datetime.now().isoformat(),
|
||||
"metadata": process.metadata or {},
|
||||
"reminded": False
|
||||
}
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": f"Process '{process.process_name}' registered"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/process/unregister")
|
||||
async def unregister_process(process: ProcessUnregister):
|
||||
"""Unregister a background process"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
if process.process_id in agent.background_processes:
|
||||
del agent.background_processes[process.process_id]
|
||||
return {
|
||||
"success": True,
|
||||
"message": f"Process {process.process_id} unregistered"
|
||||
}
|
||||
else:
|
||||
return {
|
||||
"success": False,
|
||||
"message": f"Process {process.process_id} not found"
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Main Entry Point
|
||||
# ============================================================================
|
||||
|
||||
def build_parser() -> argparse.ArgumentParser:
|
||||
"""构建命令行参数解析器(命令行参数优先级高于环境变量)。"""
|
||||
parser = argparse.ArgumentParser(
|
||||
description="事件驱动 Agent 的 HTTP 服务器(FastAPI):"
|
||||
"对外暴露 /event 等接口,把 Webhook 式的外部回调转成事件唤醒 Agent。",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog="""示例:
|
||||
python server.py # 使用默认配置(端口 8000,启用 MCP 工具)
|
||||
python server.py --port 9000 # 自定义端口
|
||||
python server.py --provider doubao # 指定大模型提供商
|
||||
python server.py --no-mcp # 只用内置工具,不加载 MCP 工具
|
||||
之后用客户端发送事件:python client.py --mode test
|
||||
""",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--host", default=os.getenv("AGENT_HOST", "0.0.0.0"),
|
||||
help="监听地址(默认:0.0.0.0)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--port", type=int, default=_env_int("AGENT_PORT", 8000),
|
||||
help="监听端口(默认:环境变量 AGENT_PORT 或 8000)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--provider", default=None,
|
||||
choices=["dashscope", "qwen", "bailian", "siliconflow", "doubao", "kimi", "moonshot", "openrouter"],
|
||||
help="大模型提供商(默认:环境变量 LLM_PROVIDER 或 kimi)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--model", default=None,
|
||||
help="模型名覆盖(默认:使用提供商默认模型)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--no-mcp", action="store_true",
|
||||
help="禁用 MCP 工具,只使用内置工具(等价于 ENABLE_MCP_TOOLS=false)",
|
||||
)
|
||||
return parser
|
||||
|
||||
|
||||
def main():
|
||||
"""Main entry point"""
|
||||
args = build_parser().parse_args()
|
||||
|
||||
# 命令行参数覆盖环境变量:init_agent() 在 lifespan 中读取这些环境变量
|
||||
if args.provider:
|
||||
os.environ["LLM_PROVIDER"] = args.provider
|
||||
if args.model:
|
||||
os.environ["LLM_MODEL"] = args.model
|
||||
if args.no_mcp:
|
||||
os.environ["ENABLE_MCP_TOOLS"] = "false"
|
||||
|
||||
print("\n" + "="*80)
|
||||
print("🤖 EVENT-TRIGGERED AGENT SERVER (FastAPI)")
|
||||
print("="*80)
|
||||
print()
|
||||
|
||||
print(f"✅ Starting server on {args.host}:{args.port}")
|
||||
print(f"📡 API Documentation: http://localhost:{args.port}/docs")
|
||||
print(f"📊 ReDoc: http://localhost:{args.port}/redoc")
|
||||
print()
|
||||
print("="*80 + "\n")
|
||||
|
||||
# Run with uvicorn
|
||||
uvicorn.run(
|
||||
app,
|
||||
host=args.host,
|
||||
port=args.port,
|
||||
log_level="info"
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,416 @@
|
||||
"""
|
||||
Event Server - FastAPI version with native async support for MCP tools
|
||||
"""
|
||||
|
||||
import os
|
||||
import logging
|
||||
from datetime import datetime, timedelta
|
||||
from typing import Dict, Any, Optional
|
||||
from contextlib import asynccontextmanager
|
||||
from fastapi import FastAPI, HTTPException, BackgroundTasks
|
||||
from pydantic import BaseModel
|
||||
from agent import EventTriggeredAgent, SystemHintConfig, resolve_provider_and_key
|
||||
from event_types import Event, EventType
|
||||
import threading
|
||||
import time
|
||||
import asyncio
|
||||
import uvicorn
|
||||
|
||||
|
||||
def _env_int(name: str, default: int) -> int:
|
||||
"""Read an integer env var; fall back to default (with a warning) if malformed."""
|
||||
raw = os.getenv(name)
|
||||
if raw is None:
|
||||
return default
|
||||
try:
|
||||
return int(raw)
|
||||
except ValueError:
|
||||
logger.warning(f"Invalid {name} value: {raw!r} (must be an integer); using default {default}")
|
||||
return default
|
||||
|
||||
|
||||
def _reasoning_safe_temperature(model, requested=1.0):
|
||||
"""Reasoning models (Kimi K3, GPT-5, ...) only accept temperature=1.
|
||||
Return 1 for those; otherwise the requested value so non-reasoning
|
||||
providers (Doubao, DeepSeek, older Moonshot) are unchanged."""
|
||||
m = str(model or "").lower().replace("/", "-")
|
||||
return 1 if ("kimi-k3" in m or "gpt-5" in m) else requested
|
||||
|
||||
|
||||
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Global agent instance
|
||||
agent: Optional[EventTriggeredAgent] = None
|
||||
agent_lock = threading.Lock()
|
||||
|
||||
# Monitoring state
|
||||
monitoring_enabled = False
|
||||
monitoring_thread: Optional[threading.Thread] = None
|
||||
|
||||
# MCP loading status
|
||||
mcp_loading_status = {
|
||||
"loading": False,
|
||||
"loaded": False,
|
||||
"tools_count": 0,
|
||||
"error": None,
|
||||
"started_at": None,
|
||||
"completed_at": None
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# FastAPI Lifecycle Events (Modern lifespan)
|
||||
# ============================================================================
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
"""Lifespan context manager for startup and shutdown"""
|
||||
global agent, monitoring_enabled
|
||||
|
||||
# Startup
|
||||
logger.info("🚀 Starting Event-Triggered Agent Server (FastAPI)")
|
||||
await init_agent()
|
||||
logger.info("✅ Server ready to receive events\n")
|
||||
|
||||
yield
|
||||
|
||||
# Shutdown
|
||||
logger.info("Shutting down server...")
|
||||
monitoring_enabled = False
|
||||
|
||||
if agent and agent.mcp_manager:
|
||||
await agent.mcp_manager.disconnect_all()
|
||||
|
||||
logger.info("✅ Server shutdown complete")
|
||||
|
||||
|
||||
# Initialize FastAPI app with lifespan
|
||||
app = FastAPI(
|
||||
title="Event-Triggered Agent Server",
|
||||
description="AI Agent with async MCP tools support",
|
||||
version="2.0.0",
|
||||
lifespan=lifespan
|
||||
)
|
||||
|
||||
|
||||
# Pydantic models for requests
|
||||
class EventRequest(BaseModel):
|
||||
event_type: str
|
||||
content: str
|
||||
metadata: Optional[Dict[str, Any]] = None
|
||||
|
||||
|
||||
class ProcessRegister(BaseModel):
|
||||
process_id: str
|
||||
process_name: str
|
||||
metadata: Optional[Dict[str, Any]] = None
|
||||
|
||||
|
||||
class ProcessUnregister(BaseModel):
|
||||
process_id: str
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Initialization
|
||||
# ============================================================================
|
||||
|
||||
async def init_agent():
|
||||
"""Initialize the agent with optional MCP tools"""
|
||||
global agent, mcp_loading_status
|
||||
|
||||
# Determine provider and key, applying the universal OpenRouter fallback.
|
||||
requested_provider = os.getenv("LLM_PROVIDER", "kimi").lower()
|
||||
provider, api_key = resolve_provider_and_key(requested_provider)
|
||||
|
||||
if not api_key:
|
||||
raise ValueError(
|
||||
f"API key not set for provider '{requested_provider}'. Set the appropriate "
|
||||
"environment variable or OPENROUTER_API_KEY."
|
||||
)
|
||||
|
||||
# Get model from environment if specified
|
||||
model = os.getenv("LLM_MODEL")
|
||||
if provider == "openrouter" and provider != requested_provider and model and "/" not in model:
|
||||
model = None
|
||||
|
||||
# Check if MCP should be enabled (default: true)
|
||||
enable_mcp = os.getenv("ENABLE_MCP_TOOLS", "true").lower() not in ["false", "0", "no"]
|
||||
|
||||
config = SystemHintConfig(
|
||||
enable_timestamps=True,
|
||||
enable_tool_counter=True,
|
||||
enable_todo_list=True,
|
||||
enable_detailed_errors=True,
|
||||
enable_system_state=True,
|
||||
save_trajectory=True,
|
||||
trajectory_file="event_agent_trajectory.json",
|
||||
temperature=_reasoning_safe_temperature(model, 0.7),
|
||||
max_tokens=4096,
|
||||
use_mcp_servers=enable_mcp
|
||||
)
|
||||
|
||||
agent = EventTriggeredAgent(
|
||||
api_key=api_key,
|
||||
provider=provider,
|
||||
model=model,
|
||||
config=config,
|
||||
verbose=True
|
||||
)
|
||||
|
||||
logger.info(f"✅ Agent initialized with {provider} provider")
|
||||
|
||||
if enable_mcp:
|
||||
logger.info("🔄 MCP tools enabled (default) - loading asynchronously...")
|
||||
await load_mcp_tools_async()
|
||||
else:
|
||||
logger.info(f"📦 Using built-in tools only (MCP disabled via ENABLE_MCP_TOOLS=false)")
|
||||
|
||||
|
||||
async def load_mcp_tools_async():
|
||||
"""Load MCP tools asynchronously"""
|
||||
global agent, mcp_loading_status
|
||||
|
||||
mcp_loading_status["loading"] = True
|
||||
mcp_loading_status["started_at"] = datetime.now().isoformat()
|
||||
|
||||
try:
|
||||
if agent:
|
||||
await agent.load_mcp_tools()
|
||||
tools_count = len(agent.mcp_manager.tools)
|
||||
|
||||
mcp_loading_status["loaded"] = True
|
||||
mcp_loading_status["loading"] = False
|
||||
mcp_loading_status["tools_count"] = tools_count
|
||||
mcp_loading_status["completed_at"] = datetime.now().isoformat()
|
||||
|
||||
logger.info(f"✅ MCP tools loaded: {tools_count} tools available")
|
||||
if tools_count > 0:
|
||||
sample_tools = list(agent.mcp_manager.tools.keys())[:5]
|
||||
logger.info(f" Sample: {sample_tools}")
|
||||
else:
|
||||
raise RuntimeError("Agent not initialized")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"❌ Failed to load MCP tools: {e}")
|
||||
mcp_loading_status["loading"] = False
|
||||
mcp_loading_status["loaded"] = False
|
||||
mcp_loading_status["error"] = str(e)
|
||||
mcp_loading_status["completed_at"] = datetime.now().isoformat()
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# API Endpoints
|
||||
# ============================================================================
|
||||
|
||||
@app.get("/")
|
||||
async def root():
|
||||
"""Root endpoint with API information"""
|
||||
return {
|
||||
"service": "Event-Triggered Agent Server",
|
||||
"version": "2.0.0",
|
||||
"status": "running",
|
||||
"docs": "/docs",
|
||||
"endpoints": {
|
||||
"health": "GET /health",
|
||||
"event": "POST /event",
|
||||
"mcp_status": "GET /mcp/status",
|
||||
"mcp_reload": "POST /mcp/reload",
|
||||
"agent_status": "GET /agent/status",
|
||||
"agent_reset": "POST /agent/reset"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@app.get("/health")
|
||||
async def health_check():
|
||||
"""Health check endpoint"""
|
||||
return {
|
||||
"status": "healthy",
|
||||
"agent_initialized": agent is not None,
|
||||
"monitoring_enabled": monitoring_enabled,
|
||||
"mcp_enabled": agent.config.use_mcp_servers if agent else False,
|
||||
"mcp_loaded": mcp_loading_status["loaded"],
|
||||
"timestamp": datetime.now().isoformat()
|
||||
}
|
||||
|
||||
|
||||
@app.get("/mcp/status")
|
||||
async def get_mcp_status():
|
||||
"""Get MCP tools loading status"""
|
||||
status = mcp_loading_status.copy()
|
||||
|
||||
# Add tool list if loaded
|
||||
if status["loaded"] and agent:
|
||||
status["tools"] = list(agent.mcp_manager.tools.keys())
|
||||
|
||||
# Group by server
|
||||
status["tools_by_server"] = {}
|
||||
for tool_name in agent.mcp_manager.tools.keys():
|
||||
server = tool_name.split("_")[0]
|
||||
if server not in status["tools_by_server"]:
|
||||
status["tools_by_server"][server] = []
|
||||
status["tools_by_server"][server].append(tool_name)
|
||||
|
||||
return status
|
||||
|
||||
|
||||
@app.post("/mcp/reload")
|
||||
async def reload_mcp_tools(background_tasks: BackgroundTasks):
|
||||
"""Manually trigger MCP tools reload"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
if mcp_loading_status["loading"]:
|
||||
raise HTTPException(status_code=409, detail="MCP tools are already loading")
|
||||
|
||||
# Use FastAPI background tasks
|
||||
background_tasks.add_task(load_mcp_tools_async)
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": "MCP tools reload started in background"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/event")
|
||||
async def handle_event(event_req: EventRequest):
|
||||
"""Handle incoming event"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
try:
|
||||
# Create event
|
||||
event_data = {
|
||||
"event_type": event_req.event_type,
|
||||
"content": event_req.content,
|
||||
"metadata": event_req.metadata or {}
|
||||
}
|
||||
event = Event.from_dict(event_data)
|
||||
|
||||
# Handle the event
|
||||
with agent_lock:
|
||||
result = agent.handle_event(event, max_iterations=20)
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"event_id": event.event_id,
|
||||
"result": {
|
||||
"final_answer": result.get('final_answer'),
|
||||
"iterations": result.get('iterations'),
|
||||
"tool_calls_count": len(result.get('tool_calls') or []),
|
||||
"todo_items": len(result.get('todo_list') or []),
|
||||
"success": result.get('success', False),
|
||||
"trajectory_file": result.get('trajectory_file')
|
||||
}
|
||||
}
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Error handling event: {e}")
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@app.get("/agent/status")
|
||||
async def get_agent_status():
|
||||
"""Get current agent status"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
return {
|
||||
"provider": agent.provider,
|
||||
"model": agent.model,
|
||||
"tool_calls_count": len(agent.tool_calls),
|
||||
"todo_items": len(agent.todo_list),
|
||||
"current_directory": agent.current_directory,
|
||||
"mcp_tools_loaded": agent.mcp_tools_loaded,
|
||||
"mcp_tools_count": len(agent.mcp_manager.tools) if agent.mcp_tools_loaded else 0
|
||||
}
|
||||
|
||||
|
||||
@app.post("/agent/reset")
|
||||
async def reset_agent():
|
||||
"""Reset agent state"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
agent.reset()
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": "Agent state reset successfully"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/process/register")
|
||||
async def register_process(process: ProcessRegister):
|
||||
"""Register a background process for monitoring"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
agent.background_processes[process.process_id] = {
|
||||
"name": process.process_name,
|
||||
"start_time": datetime.now().isoformat(),
|
||||
"metadata": process.metadata or {},
|
||||
"reminded": False
|
||||
}
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"message": f"Process '{process.process_name}' registered"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/process/unregister")
|
||||
async def unregister_process(process: ProcessUnregister):
|
||||
"""Unregister a background process"""
|
||||
if agent is None:
|
||||
raise HTTPException(status_code=500, detail="Agent not initialized")
|
||||
|
||||
with agent_lock:
|
||||
if process.process_id in agent.background_processes:
|
||||
del agent.background_processes[process.process_id]
|
||||
return {
|
||||
"success": True,
|
||||
"message": f"Process {process.process_id} unregistered"
|
||||
}
|
||||
else:
|
||||
return {
|
||||
"success": False,
|
||||
"message": f"Process {process.process_id} not found"
|
||||
}
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Main Entry Point
|
||||
# ============================================================================
|
||||
|
||||
def main():
|
||||
"""Main entry point"""
|
||||
print("\n" + "="*80)
|
||||
print("🤖 EVENT-TRIGGERED AGENT SERVER (FastAPI)")
|
||||
print("="*80)
|
||||
print()
|
||||
|
||||
# Get port from environment
|
||||
port = _env_int('AGENT_PORT', 8000)
|
||||
|
||||
print(f"✅ Starting server on port {port}")
|
||||
print(f"📡 API Documentation: http://localhost:{port}/docs")
|
||||
print(f"📊 ReDoc: http://localhost:{port}/redoc")
|
||||
print()
|
||||
print("="*80 + "\n")
|
||||
|
||||
# Run with uvicorn
|
||||
uvicorn.run(
|
||||
app,
|
||||
host="0.0.0.0",
|
||||
port=port,
|
||||
log_level="info"
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,138 @@
|
||||
"""Offline contract tests; these are not substitutes for the live campaign receipt."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import json
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
MODULE_PATH = HERE / "unipile_mailbox_experiment.py"
|
||||
SPEC = importlib.util.spec_from_file_location("experiment_6_1_unipile", MODULE_PATH)
|
||||
experiment = importlib.util.module_from_spec(SPEC)
|
||||
assert SPEC.loader is not None
|
||||
sys.modules[SPEC.name] = experiment
|
||||
SPEC.loader.exec_module(experiment)
|
||||
|
||||
|
||||
def _email(kind: str, date: str, suffix: str) -> dict:
|
||||
base = {"id": f"email-{suffix}", "account_id": "account-real",
|
||||
"date": date, "role": "inbox", "folders": ["Inbox"]}
|
||||
if kind == "meeting_invitation":
|
||||
return {**base, "subject": "Meeting invitation: design review",
|
||||
"body_plain": "START_UTC: 2026-08-03T10:00:00.000Z\n"
|
||||
"END_UTC: 2026-08-03T11:00:00.000Z"}
|
||||
if kind == "customer_complaint":
|
||||
return {**base, "subject": "Customer complaint: delayed order",
|
||||
"body_plain": "Customer complaint for order #E44-TEST. Please escalate."}
|
||||
return {**base, "subject": "Marketing newsletter",
|
||||
"body_plain": "Marketing newsletter promotion. Click to unsubscribe."}
|
||||
|
||||
|
||||
def _provider(request: httpx.Request) -> httpx.Response:
|
||||
assert request.headers.get("X-API-KEY") == "unit-secret"
|
||||
if request.method == "GET" and request.url.path == "/api/v1/calendars":
|
||||
return httpx.Response(200, json={"data": [{"id": "calendar-real",
|
||||
"is_primary": True}]})
|
||||
if request.method == "GET" and request.url.path.endswith("/events"):
|
||||
return httpx.Response(200, json={"data": []})
|
||||
if request.method == "GET" and request.url.path == "/api/v1/folders":
|
||||
return httpx.Response(200, json={"items": [{"id": "folder-archive",
|
||||
"name": "Archive",
|
||||
"role": "archive"}]})
|
||||
if request.method == "PUT" and request.url.path == "/api/v1/emails/email-marketing":
|
||||
assert json.loads(request.content) == {"folders": ["archive"]}
|
||||
return httpx.Response(200, json={"object": "EmailUpdated"})
|
||||
if request.method == "GET" and request.url.path == "/api/v1/emails/email-marketing":
|
||||
return httpx.Response(200, json={"id": "email-marketing",
|
||||
"role": "archive", "folders": ["Archive"]})
|
||||
return httpx.Response(404, json={"type": "unexpected_test_request"})
|
||||
|
||||
|
||||
def test_classification_is_unique_and_meeting_interval_is_exact():
|
||||
email = _email("meeting_invitation", "2026-08-01T00:00:01.000Z", "meeting")
|
||||
assert experiment.classify_email(email) == "meeting_invitation"
|
||||
start, end = experiment.meeting_interval(email)
|
||||
assert start.isoformat() == "2026-08-03T10:00:00+00:00"
|
||||
assert (end - start).total_seconds() == 3600
|
||||
ambiguous = {**email, "subject": "Meeting invitation and marketing newsletter",
|
||||
"body_plain": email["body_plain"] + "\nMarketing unsubscribe"}
|
||||
with pytest.raises(ValueError, match="not unique"):
|
||||
experiment.classify_email(ambiguous)
|
||||
|
||||
|
||||
def test_real_api_error_is_receipted_and_raises():
|
||||
def unauthorized(_: httpx.Request) -> httpx.Response:
|
||||
return httpx.Response(401, json={"status": 401,
|
||||
"type": "errors/missing_credentials",
|
||||
"title": "Missing credentials"})
|
||||
|
||||
client = experiment.UnipileClient(
|
||||
"api.example.invalid:12345", "unit-secret",
|
||||
transport=httpx.MockTransport(unauthorized),
|
||||
)
|
||||
with pytest.raises(experiment.UnipileAPIError):
|
||||
client.list_accounts()
|
||||
assert client.calls == [{
|
||||
**client.calls[0], "status": 401, "success": False,
|
||||
"credential_scheme": "X-API-KEY",
|
||||
"error_type": "errors/missing_credentials",
|
||||
}]
|
||||
assert "unit-secret" not in experiment.canonical_json(client.calls)
|
||||
client.close()
|
||||
|
||||
def test_three_email_workflow_and_acceptance(tmp_path):
|
||||
client = experiment.UnipileClient(
|
||||
"api.example.invalid:12345", "unit-secret",
|
||||
transport=httpx.MockTransport(_provider),
|
||||
)
|
||||
runner = experiment.MailboxExperiment(
|
||||
client, tmp_path, "account-real", "account-real"
|
||||
)
|
||||
# Deliberately unordered input proves that the FIFO queue uses provider time.
|
||||
emails = [
|
||||
_email("marketing", "2026-08-01T00:00:03.000Z", "marketing"),
|
||||
_email("meeting_invitation", "2026-08-01T00:00:01.000Z", "meeting"),
|
||||
_email("customer_complaint", "2026-08-01T00:00:02.000Z", "complaint"),
|
||||
]
|
||||
runner.process_queue(emails)
|
||||
assert [row["classification"] for row in runner.workflows] == [
|
||||
"meeting_invitation", "customer_complaint", "marketing"
|
||||
]
|
||||
result = experiment.derive_acceptance(
|
||||
runner.events, runner.workflows, client.calls,
|
||||
[{"object": "EmailSent"}] * 3,
|
||||
credential_secret="unit-secret", dsn_secret="api.example.invalid:12345",
|
||||
)
|
||||
assert result["status"] == "passed"
|
||||
assert all(result["gates"].values())
|
||||
assert (tmp_path / "artifacts" / "meeting_reply_draft.txt").stat().st_size > 100
|
||||
assert (tmp_path / "artifacts" / "high_priority_notifications.jsonl").is_file()
|
||||
assert "unit-secret" not in experiment.canonical_json(client.calls)
|
||||
client.close()
|
||||
|
||||
|
||||
def test_acceptance_fails_when_provider_archive_verification_is_missing(tmp_path):
|
||||
client = experiment.UnipileClient(
|
||||
"api.example.invalid:12345", "unit-secret",
|
||||
transport=httpx.MockTransport(_provider),
|
||||
)
|
||||
runner = experiment.MailboxExperiment(client, tmp_path, "account-real", "account-real")
|
||||
runner.process_queue([
|
||||
_email("meeting_invitation", "2026-08-01T00:00:01.000Z", "meeting"),
|
||||
_email("customer_complaint", "2026-08-01T00:00:02.000Z", "complaint"),
|
||||
_email("marketing", "2026-08-01T00:00:03.000Z", "marketing"),
|
||||
])
|
||||
runner.workflows[-1]["archive"]["verified"] = False
|
||||
result = experiment.derive_acceptance(
|
||||
runner.events, runner.workflows, client.calls, [{}, {}, {}],
|
||||
credential_secret="unit-secret", dsn_secret="api.example.invalid:12345",
|
||||
)
|
||||
assert result["status"] == "failed"
|
||||
assert not result["gates"]["marketing_archived_and_verified_through_provider"]
|
||||
client.close()
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Test import bootstrap for the agent-with-event-trigger experiment."""
|
||||
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
|
||||
EXPERIMENT_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(EXPERIMENT_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(EXPERIMENT_ROOT))
|
||||
@@ -0,0 +1,137 @@
|
||||
"""
|
||||
Simple demo script to test the event-triggered agent locally
|
||||
without needing the server/client architecture
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
EXPERIMENT_ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(EXPERIMENT_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(EXPERIMENT_ROOT))
|
||||
|
||||
from agent import EventTriggeredAgent, SystemHintConfig, resolve_provider_and_key
|
||||
from event_types import Event, EventType
|
||||
|
||||
|
||||
def _reasoning_safe_temperature(model, requested=1.0):
|
||||
"""Reasoning models (Kimi K3, GPT-5, ...) only accept temperature=1.
|
||||
Return 1 for those; otherwise the requested value so non-reasoning
|
||||
providers (Doubao, DeepSeek, older Moonshot) are unchanged."""
|
||||
m = str(model or "").lower().replace("/", "-")
|
||||
return 1 if ("kimi-k3" in m or "gpt-5" in m) else requested
|
||||
|
||||
|
||||
def main():
|
||||
"""Run a simple demo of the event-triggered agent"""
|
||||
|
||||
print("\n" + "="*80)
|
||||
print("🧪 EVENT-TRIGGERED AGENT DEMO")
|
||||
print("="*80)
|
||||
print()
|
||||
|
||||
# Get provider and API key (including DashScope/Bailian aliases and fallback)
|
||||
provider = os.getenv("LLM_PROVIDER", "kimi").lower()
|
||||
provider, api_key = resolve_provider_and_key(provider)
|
||||
|
||||
if not api_key:
|
||||
print(f"❌ Error: Please set API key for provider '{provider}'")
|
||||
print(" export DASHSCOPE_API_KEY='your-api-key-here' # for dashscope/qwen/bailian")
|
||||
return
|
||||
|
||||
# Get optional model override
|
||||
model = os.getenv("LLM_MODEL")
|
||||
|
||||
# Create agent with full system hints (matching conversational_agent.py config)
|
||||
config = SystemHintConfig(
|
||||
enable_timestamps=True,
|
||||
enable_tool_counter=True,
|
||||
enable_todo_list=True,
|
||||
enable_detailed_errors=True,
|
||||
enable_system_state=True,
|
||||
save_trajectory=True,
|
||||
trajectory_file="demo_trajectory.json",
|
||||
temperature=_reasoning_safe_temperature(model, 0.7), # Matching conversational_agent.py
|
||||
max_tokens=4096 # Matching conversational_agent.py
|
||||
)
|
||||
|
||||
agent = EventTriggeredAgent(
|
||||
api_key=api_key,
|
||||
provider=provider,
|
||||
model=model,
|
||||
config=config,
|
||||
verbose=True
|
||||
)
|
||||
|
||||
print("✅ Agent initialized\n")
|
||||
|
||||
# Demo 1: Web message
|
||||
print("\n" + "-"*80)
|
||||
print("📋 Demo 1: Web Interface Message")
|
||||
print("-"*80)
|
||||
|
||||
event1 = Event(
|
||||
event_type=EventType.WEB_MESSAGE,
|
||||
content="Create a simple Python script that prints 'Hello, Event-Triggered Agent!' and save it as demo_hello.py",
|
||||
metadata={"user_id": "demo_user"}
|
||||
)
|
||||
|
||||
result1 = agent.handle_event(event1, max_iterations=10)
|
||||
print(f"\n✅ Event handled. Success: {result1['success']}")
|
||||
print(f" Iterations: {result1['iterations']}")
|
||||
print(f" Tool calls: {len(result1['tool_calls'])}")
|
||||
|
||||
# Demo 2: IM message
|
||||
print("\n" + "-"*80)
|
||||
print("📋 Demo 2: Instant Message")
|
||||
print("-"*80)
|
||||
|
||||
event2 = Event(
|
||||
event_type=EventType.IM_MESSAGE,
|
||||
content="Can you run the script you just created?",
|
||||
metadata={"sender": "Alice", "platform": "Slack"}
|
||||
)
|
||||
|
||||
result2 = agent.handle_event(event2, max_iterations=10)
|
||||
print(f"\n✅ Event handled. Success: {result2['success']}")
|
||||
print(f" Iterations: {result2['iterations']}")
|
||||
print(f" Tool calls: {len(result2['tool_calls'])}")
|
||||
|
||||
# Demo 3: System alert
|
||||
print("\n" + "-"*80)
|
||||
print("📋 Demo 3: System Alert")
|
||||
print("-"*80)
|
||||
|
||||
event3 = Event(
|
||||
event_type=EventType.SYSTEM_ALERT,
|
||||
content="Please check the current directory and list all Python files.",
|
||||
metadata={"alert_type": "routine_check"}
|
||||
)
|
||||
|
||||
result3 = agent.handle_event(event3, max_iterations=10)
|
||||
print(f"\n✅ Event handled. Success: {result3['success']}")
|
||||
print(f" Iterations: {result3['iterations']}")
|
||||
print(f" Tool calls: {len(result3['tool_calls'])}")
|
||||
|
||||
# Summary
|
||||
print("\n" + "="*80)
|
||||
print("📊 DEMO SUMMARY")
|
||||
print("="*80)
|
||||
print(f"Total events processed: 3")
|
||||
print(f"Total tool calls: {len(agent.tool_calls)}")
|
||||
print(f"Total conversation messages: {len(agent.conversation_history)}")
|
||||
print(f"Trajectory saved to: {config.trajectory_file}")
|
||||
print()
|
||||
print("✅ Demo completed successfully!")
|
||||
print()
|
||||
print("💡 Next steps:")
|
||||
print(" - Check demo_hello.py to see the created file")
|
||||
print(" - View demo_trajectory.json to see the full conversation")
|
||||
print(" - Run 'python server.py' and 'python client.py' for the full system")
|
||||
print("="*80 + "\n")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,73 @@
|
||||
"""Regression tests for the code interpreter's execution namespace."""
|
||||
|
||||
import importlib.util
|
||||
import sys
|
||||
import types
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
|
||||
def _load_agent_module():
|
||||
openai = types.ModuleType("openai")
|
||||
openai.OpenAI = type("OpenAI", (), {})
|
||||
|
||||
mcp = types.ModuleType("mcp")
|
||||
mcp.__path__ = []
|
||||
mcp.ClientSession = type("ClientSession", (), {})
|
||||
mcp.StdioServerParameters = type("StdioServerParameters", (), {})
|
||||
|
||||
mcp_client = types.ModuleType("mcp.client")
|
||||
mcp_client.__path__ = []
|
||||
mcp_stdio = types.ModuleType("mcp.client.stdio")
|
||||
mcp_stdio.stdio_client = lambda *args, **kwargs: None
|
||||
mcp_types = types.ModuleType("mcp.types")
|
||||
mcp_types.TextContent = type("TextContent", (), {})
|
||||
|
||||
stubs = {
|
||||
"openai": openai,
|
||||
"mcp": mcp,
|
||||
"mcp.client": mcp_client,
|
||||
"mcp.client.stdio": mcp_stdio,
|
||||
"mcp.types": mcp_types,
|
||||
}
|
||||
module_path = Path(__file__).resolve().parents[1] / "agent.py"
|
||||
module_name = "_event_trigger_agent_under_test"
|
||||
spec = importlib.util.spec_from_file_location(module_name, module_path)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
|
||||
with patch.dict(sys.modules, stubs):
|
||||
sys.modules[module_name] = module
|
||||
sys.path.insert(0, str(module_path.parent))
|
||||
try:
|
||||
spec.loader.exec_module(module)
|
||||
finally:
|
||||
sys.path.pop(0)
|
||||
|
||||
return module
|
||||
|
||||
|
||||
class CodeInterpreterNamespaceTests(unittest.TestCase):
|
||||
def test_code_interpreter_shares_names_with_defined_functions(self):
|
||||
agent_module = _load_agent_module()
|
||||
agent = agent_module.EventTriggeredAgent.__new__(agent_module.EventTriggeredAgent)
|
||||
|
||||
result = agent._tool_code_interpreter(
|
||||
"value = 5\n"
|
||||
"def double():\n"
|
||||
" return value * 2\n"
|
||||
"print(double())"
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
result,
|
||||
{
|
||||
"success": True,
|
||||
"stdout": "10\n",
|
||||
"stderr": "",
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,36 @@
|
||||
"""Regression test: malformed AGENT_PORT must not crash server startup.
|
||||
|
||||
Both server variants parsed AGENT_PORT with bare int() (one inside
|
||||
build_parser's default, one in main), so AGENT_PORT=abc crashed with an
|
||||
unhandled ValueError at startup. They now fall back to 8000 with a warning.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
|
||||
import server
|
||||
import server_fastapi
|
||||
|
||||
|
||||
def test_env_int_falls_back_on_malformed(monkeypatch):
|
||||
monkeypatch.setenv("AGENT_PORT", "abc")
|
||||
assert server._env_int("AGENT_PORT", 8000) == 8000
|
||||
assert server_fastapi._env_int("AGENT_PORT", 8000) == 8000
|
||||
|
||||
|
||||
def test_env_int_parses_valid_value(monkeypatch):
|
||||
monkeypatch.setenv("AGENT_PORT", "9000")
|
||||
assert server._env_int("AGENT_PORT", 8000) == 9000
|
||||
assert server_fastapi._env_int("AGENT_PORT", 8000) == 9000
|
||||
|
||||
|
||||
def test_env_int_default_when_unset(monkeypatch):
|
||||
monkeypatch.delenv("AGENT_PORT", raising=False)
|
||||
assert server._env_int("AGENT_PORT", 8000) == 8000
|
||||
|
||||
|
||||
def test_build_parser_survives_malformed_env(monkeypatch):
|
||||
monkeypatch.setenv("AGENT_PORT", "not-a-port")
|
||||
args = server.build_parser().parse_args([])
|
||||
assert args.port == 8000
|
||||
@@ -0,0 +1,30 @@
|
||||
"""
|
||||
Test suite locking out TypeError in event server response formatting
|
||||
when handle_event returns a result dictionary with tool_calls or todo_list set to None.
|
||||
"""
|
||||
|
||||
def test_server_result_formatting_handles_null_lists():
|
||||
"""
|
||||
Ensure event response dictionary formats tool_calls_count and todo_items without TypeError
|
||||
when tool_calls or todo_list is None.
|
||||
"""
|
||||
result = {
|
||||
'final_answer': 'Done',
|
||||
'iterations': 1,
|
||||
'tool_calls': None,
|
||||
'todo_list': None,
|
||||
'success': True,
|
||||
'trajectory_file': None
|
||||
}
|
||||
|
||||
formatted = {
|
||||
"final_answer": result.get('final_answer'),
|
||||
"iterations": result.get('iterations'),
|
||||
"tool_calls_count": len(result.get('tool_calls') or []),
|
||||
"todo_items": len(result.get('todo_list') or []),
|
||||
"success": result.get('success', False),
|
||||
"trajectory_file": result.get('trajectory_file')
|
||||
}
|
||||
|
||||
assert formatted["tool_calls_count"] == 0
|
||||
assert formatted["todo_items"] == 0
|
||||
@@ -0,0 +1,760 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run Chapter 4 Experiment 6-1 against a real Unipile mailbox.
|
||||
|
||||
The listener uses documented mailbox polling, which the manuscript explicitly
|
||||
allows as an alternative to push notifications. It never substitutes local
|
||||
mail files for provider objects. Missing/invalid credentials produce a durable
|
||||
blocked receipt instead of a successful demonstration.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from collections import deque
|
||||
from datetime import datetime, timedelta, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
PROTOCOL_PATH = HERE / "experiment_protocol.json"
|
||||
VALIDATION_ROOT = HERE / "validation" / "experiment_6_1"
|
||||
UTC = timezone.utc
|
||||
SIMULATION_PATTERN = re.compile(
|
||||
r"\b(mock(?:ed)?|placeholder|synthetic|simulat(?:ed|ion))\b", re.IGNORECASE
|
||||
)
|
||||
CREDENTIAL_PATTERN = re.compile(r"\b(?:sk|gh[opusr])-[A-Za-z0-9_-]{12,}\b")
|
||||
|
||||
|
||||
def canonical_json(value: Any) -> str:
|
||||
return json.dumps(value, ensure_ascii=False, sort_keys=True,
|
||||
separators=(",", ":"), default=str)
|
||||
|
||||
|
||||
def sha256(value: bytes | str) -> str:
|
||||
if isinstance(value, str):
|
||||
value = value.encode()
|
||||
return hashlib.sha256(value).hexdigest()
|
||||
|
||||
|
||||
def write_json(path: Path, value: Any) -> None:
|
||||
text = json.dumps(value, ensure_ascii=False, indent=2, default=str) + "\n"
|
||||
if CREDENTIAL_PATTERN.search(text):
|
||||
raise ValueError(f"credential-shaped value in {path}")
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(text, encoding="utf-8")
|
||||
|
||||
|
||||
def iso_millis(value: datetime) -> str:
|
||||
return value.astimezone(UTC).isoformat(timespec="milliseconds").replace("+00:00", "Z")
|
||||
|
||||
|
||||
def parse_datetime(value: Any) -> datetime | None:
|
||||
if isinstance(value, dict):
|
||||
value = value.get("date_time") or value.get("dateTime") or value.get("datetime") \
|
||||
or value.get("date")
|
||||
if not value:
|
||||
return None
|
||||
text = str(value).strip().replace("Z", "+00:00")
|
||||
try:
|
||||
parsed = datetime.fromisoformat(text)
|
||||
except ValueError:
|
||||
return None
|
||||
if parsed.tzinfo is None:
|
||||
parsed = parsed.replace(tzinfo=UTC)
|
||||
return parsed.astimezone(UTC)
|
||||
|
||||
|
||||
def redacted(value: Any, key: str = "") -> Any:
|
||||
"""Retain audit shape while hashing identities and message bodies."""
|
||||
lower = key.lower()
|
||||
if isinstance(value, str) and re.fullmatch(r"[^@\s]+@[^@\s]+\.[^@\s]+", value):
|
||||
return {"sha256": sha256(value)[:20], "present": True, "kind": "email_address"}
|
||||
if any(token in lower for token in ("token", "authorization", "api_key", "credential")):
|
||||
return "<redacted>"
|
||||
if lower in {"body", "body_plain", "text", "content"} and isinstance(value, str):
|
||||
return {"sha256": sha256(value), "characters": len(value)}
|
||||
if lower == "subject" and isinstance(value, str):
|
||||
# Experiment subjects contain no personal data and prove scenario fidelity.
|
||||
return value
|
||||
if lower.endswith("_id") or lower in {"id", "identifier", "email", "from", "to"}:
|
||||
if isinstance(value, (str, int)):
|
||||
return {"sha256": sha256(str(value))[:20], "present": bool(value)}
|
||||
if isinstance(value, dict):
|
||||
return {str(child_key): redacted(child, str(child_key))
|
||||
for child_key, child in value.items()}
|
||||
if isinstance(value, list):
|
||||
return [redacted(child, key) for child in value]
|
||||
return value
|
||||
|
||||
|
||||
class UnipileAPIError(RuntimeError):
|
||||
def __init__(self, method: str, path: str, status: int, payload: Any):
|
||||
self.method = method
|
||||
self.path = path
|
||||
self.status = status
|
||||
self.payload = payload
|
||||
title = payload.get("title") if isinstance(payload, dict) else None
|
||||
error_type = payload.get("type") if isinstance(payload, dict) else None
|
||||
super().__init__(f"{method} {path} returned {status}: {error_type or title or 'API error'}")
|
||||
|
||||
|
||||
class UnipileClient:
|
||||
"""Small receipt-producing adapter for the official Email/Calendar API."""
|
||||
|
||||
def __init__(self, dsn: str, access_token: str, *, timeout: float = 45,
|
||||
transport: httpx.BaseTransport | None = None):
|
||||
if not dsn or not access_token:
|
||||
raise ValueError("UNIPILE_DSN and UNIPILE_ACCESS_TOKEN are required")
|
||||
base = dsn.strip().rstrip("/")
|
||||
if not base.startswith(("http://", "https://")):
|
||||
base = "https://" + base
|
||||
self.base_url = base
|
||||
self.access_token = access_token.strip()
|
||||
self.calls: list[dict[str, Any]] = []
|
||||
self.http = httpx.Client(
|
||||
base_url=base,
|
||||
timeout=timeout,
|
||||
follow_redirects=True,
|
||||
headers={"X-API-KEY": self.access_token,
|
||||
"User-Agent": "ai-agent-book-experiment/4.4"},
|
||||
transport=transport,
|
||||
)
|
||||
|
||||
def close(self) -> None:
|
||||
self.http.close()
|
||||
|
||||
def request(self, method: str, path: str, *, params: dict[str, Any] | None = None,
|
||||
json_body: dict[str, Any] | None = None,
|
||||
multipart: dict[str, Any] | None = None,
|
||||
expected: set[int] | None = None) -> Any:
|
||||
expected = expected or {200}
|
||||
started = time.perf_counter()
|
||||
request_kwargs: dict[str, Any] = {"params": params}
|
||||
if json_body is not None:
|
||||
request_kwargs["json"] = json_body
|
||||
if multipart is not None:
|
||||
request_kwargs["files"] = {key: (None, value) for key, value in multipart.items()}
|
||||
try:
|
||||
response = self.http.request(method, path, **request_kwargs)
|
||||
try:
|
||||
payload: Any = response.json()
|
||||
except ValueError:
|
||||
payload = {"non_json_sha256": sha256(response.content),
|
||||
"bytes": len(response.content)}
|
||||
receipt = {
|
||||
"method": method.upper(),
|
||||
"path": path,
|
||||
"request": redacted({"params": params or {},
|
||||
"json": json_body,
|
||||
"multipart": multipart}),
|
||||
"credential_scheme": "X-API-KEY",
|
||||
"status": response.status_code,
|
||||
"success": response.status_code in expected,
|
||||
"latency_seconds": round(time.perf_counter() - started, 3),
|
||||
"response_sha256": sha256(response.content),
|
||||
"response_bytes": len(response.content),
|
||||
"response_shape": sorted(payload) if isinstance(payload, dict)
|
||||
else type(payload).__name__,
|
||||
"error_type": payload.get("type") if isinstance(payload, dict) else None,
|
||||
"error_title": payload.get("title") if isinstance(payload, dict) else None,
|
||||
}
|
||||
self.calls.append(receipt)
|
||||
if response.status_code not in expected:
|
||||
raise UnipileAPIError(method.upper(), path, response.status_code, payload)
|
||||
return payload
|
||||
except UnipileAPIError:
|
||||
raise
|
||||
except Exception as exc:
|
||||
self.calls.append({
|
||||
"method": method.upper(), "path": path,
|
||||
"request": redacted({"params": params or {}, "json": json_body,
|
||||
"multipart": multipart}),
|
||||
"credential_scheme": "X-API-KEY", "status": None, "success": False,
|
||||
"latency_seconds": round(time.perf_counter() - started, 3),
|
||||
"error_type": type(exc).__name__, "error_title": str(exc)[:300],
|
||||
})
|
||||
raise
|
||||
|
||||
def list_accounts(self) -> list[dict[str, Any]]:
|
||||
payload = self.request("GET", "/api/v1/accounts", params={"limit": 100})
|
||||
items = payload.get("items", []) if isinstance(payload, dict) else []
|
||||
if not isinstance(items, list):
|
||||
raise ValueError("Unipile accounts response did not contain an items list")
|
||||
return items
|
||||
|
||||
def list_folders(self, account_id: str) -> list[dict[str, Any]]:
|
||||
payload = self.request("GET", "/api/v1/folders",
|
||||
params={"account_id": account_id})
|
||||
return payload.get("items", []) if isinstance(payload, dict) else []
|
||||
|
||||
def list_emails(self, account_id: str, *, after: datetime,
|
||||
folder: str | None = None, limit: int = 100) -> list[dict[str, Any]]:
|
||||
params: dict[str, Any] = {
|
||||
"account_id": account_id, "after": iso_millis(after),
|
||||
"limit": min(max(limit, 1), 250), "meta_only": False,
|
||||
}
|
||||
if folder:
|
||||
params["folder"] = folder
|
||||
payload = self.request("GET", "/api/v1/emails", params=params)
|
||||
items = payload.get("items", []) if isinstance(payload, dict) else []
|
||||
if not isinstance(items, list):
|
||||
raise ValueError("Unipile email response did not contain an items list")
|
||||
return items
|
||||
|
||||
def get_email(self, email_id: str) -> dict[str, Any]:
|
||||
payload = self.request("GET", f"/api/v1/emails/{email_id}")
|
||||
if not isinstance(payload, dict):
|
||||
raise ValueError("Unipile email response was not an object")
|
||||
return payload
|
||||
|
||||
def update_email_folders(self, email_id: str, folders: list[str]) -> dict[str, Any]:
|
||||
payload = self.request("PUT", f"/api/v1/emails/{email_id}",
|
||||
json_body={"folders": folders})
|
||||
if not isinstance(payload, dict):
|
||||
raise ValueError("Unipile update response was not an object")
|
||||
return payload
|
||||
|
||||
def send_email(self, account_id: str, recipient: str, subject: str,
|
||||
body: str) -> dict[str, Any]:
|
||||
payload = self.request(
|
||||
"POST", "/api/v1/emails", expected={201}, multipart={
|
||||
"account_id": account_id,
|
||||
"to": json.dumps([{"display_name": "Experiment 6-1 mailbox",
|
||||
"identifier": recipient}]),
|
||||
"subject": subject,
|
||||
"body": body,
|
||||
},
|
||||
)
|
||||
if not isinstance(payload, dict):
|
||||
raise ValueError("Unipile send response was not an object")
|
||||
return payload
|
||||
|
||||
def list_calendars(self, account_id: str) -> list[dict[str, Any]]:
|
||||
payload = self.request("GET", "/api/v1/calendars",
|
||||
params={"account_id": account_id, "limit": 100})
|
||||
data = payload.get("data", []) if isinstance(payload, dict) else []
|
||||
if not isinstance(data, list):
|
||||
raise ValueError("Unipile calendars response did not contain a data list")
|
||||
return data
|
||||
|
||||
def list_calendar_events(self, account_id: str, calendar_id: str,
|
||||
start: datetime, end: datetime) -> list[dict[str, Any]]:
|
||||
payload = self.request(
|
||||
"GET", f"/api/v1/calendars/{calendar_id}/events", params={
|
||||
"account_id": account_id,
|
||||
"start": iso_millis(start - timedelta(days=1)),
|
||||
"end": iso_millis(end + timedelta(days=1)),
|
||||
"expand_recurring": True, "limit": 250,
|
||||
},
|
||||
)
|
||||
data = payload.get("data", []) if isinstance(payload, dict) else []
|
||||
if not isinstance(data, list):
|
||||
raise ValueError("Unipile events response did not contain a data list")
|
||||
return data
|
||||
|
||||
|
||||
def account_email(account: dict[str, Any]) -> str | None:
|
||||
"""Find an email-looking account identity without exposing it in receipts."""
|
||||
preferred = ("email", "identifier", "username", "user", "name")
|
||||
for key in preferred:
|
||||
value = account.get(key)
|
||||
if isinstance(value, str) and re.fullmatch(r"[^@\s]+@[^@\s]+\.[^@\s]+", value):
|
||||
return value
|
||||
for value in account.values():
|
||||
if isinstance(value, dict):
|
||||
found = account_email(value)
|
||||
if found:
|
||||
return found
|
||||
return None
|
||||
|
||||
|
||||
def email_text(email: dict[str, Any]) -> str:
|
||||
for key in ("body_plain", "body", "text"):
|
||||
value = email.get(key)
|
||||
if isinstance(value, str) and value.strip():
|
||||
return value
|
||||
return ""
|
||||
|
||||
|
||||
def classify_email(email: dict[str, Any]) -> str:
|
||||
text = f"{email.get('subject', '')}\n{email_text(email)}".lower()
|
||||
matches = []
|
||||
if "meeting invitation" in text and "start_utc:" in text and "end_utc:" in text:
|
||||
matches.append("meeting_invitation")
|
||||
if "customer complaint" in text and re.search(r"order\s*#[a-z0-9-]+", text):
|
||||
matches.append("customer_complaint")
|
||||
if "marketing" in text and ("unsubscribe" in text or "newsletter" in text):
|
||||
matches.append("marketing")
|
||||
if len(matches) != 1:
|
||||
raise ValueError(f"email classification was not unique: {matches}")
|
||||
return matches[0]
|
||||
|
||||
|
||||
def meeting_interval(email: dict[str, Any]) -> tuple[datetime, datetime]:
|
||||
text = email_text(email)
|
||||
start_match = re.search(r"START_UTC:\s*([^\s]+)", text, re.IGNORECASE)
|
||||
end_match = re.search(r"END_UTC:\s*([^\s]+)", text, re.IGNORECASE)
|
||||
start = parse_datetime(start_match.group(1)) if start_match else None
|
||||
end = parse_datetime(end_match.group(1)) if end_match else None
|
||||
if not start or not end or end <= start:
|
||||
raise ValueError("meeting email lacked a valid START_UTC/END_UTC interval")
|
||||
return start, end
|
||||
|
||||
|
||||
def event_overlaps(event: dict[str, Any], start: datetime, end: datetime) -> bool:
|
||||
event_start = parse_datetime(event.get("start") or event.get("start_at"))
|
||||
event_end = parse_datetime(event.get("end") or event.get("end_at"))
|
||||
cancelled = bool(event.get("is_cancelled")) or str(event.get("status", "")).lower() == "cancelled"
|
||||
return bool(event_start and event_end and not cancelled
|
||||
and event_start < end and event_end > start)
|
||||
|
||||
|
||||
def canonical_event(email: dict[str, Any], sequence: int) -> dict[str, Any]:
|
||||
email_id = str(email.get("id", ""))
|
||||
account_id = str(email.get("account_id", ""))
|
||||
if not email_id or not account_id:
|
||||
raise ValueError("provider email object lacked id/account_id")
|
||||
return {
|
||||
"sequence": sequence,
|
||||
"source": {"type": "email", "provider": "unipile",
|
||||
"email_id_sha256": sha256(email_id)[:20],
|
||||
"account_id_sha256": sha256(account_id)[:20]},
|
||||
"channel": "unipile_mailbox_poll",
|
||||
"content": {"subject": email.get("subject", ""),
|
||||
"body_sha256": sha256(email_text(email)),
|
||||
"body_characters": len(email_text(email))},
|
||||
"context": {"provider_date": email.get("date"), "role": email.get("role"),
|
||||
"folders_count": len(email.get("folders") or [])},
|
||||
"provider_receipt_sha256": sha256(canonical_json(redacted(email))),
|
||||
}
|
||||
|
||||
|
||||
def provider_date(email: dict[str, Any]) -> tuple[datetime, str]:
|
||||
return (parse_datetime(email.get("date")) or datetime.min.replace(tzinfo=UTC),
|
||||
str(email.get("id", "")))
|
||||
|
||||
|
||||
class MailboxExperiment:
|
||||
def __init__(self, client: UnipileClient, campaign_dir: Path,
|
||||
account_id: str, calendar_account_id: str):
|
||||
self.client = client
|
||||
self.campaign_dir = campaign_dir
|
||||
self.account_id = account_id
|
||||
self.calendar_account_id = calendar_account_id
|
||||
self.events: list[dict[str, Any]] = []
|
||||
self.workflows: list[dict[str, Any]] = []
|
||||
|
||||
def _meeting(self, email: dict[str, Any]) -> dict[str, Any]:
|
||||
start, end = meeting_interval(email)
|
||||
calendars = self.client.list_calendars(self.calendar_account_id)
|
||||
if not calendars:
|
||||
raise RuntimeError("calendar conflict check returned no calendars")
|
||||
calendar = next((row for row in calendars
|
||||
if row.get("is_primary") or row.get("is_default")), calendars[0])
|
||||
calendar_id = str(calendar.get("id", ""))
|
||||
if not calendar_id:
|
||||
raise ValueError("selected calendar lacked an id")
|
||||
events = self.client.list_calendar_events(
|
||||
self.calendar_account_id, calendar_id, start, end
|
||||
)
|
||||
conflicts = [event for event in events if event_overlaps(event, start, end)]
|
||||
disposition = "decline" if conflicts else "accept"
|
||||
draft = (
|
||||
f"Subject: Re: {email.get('subject', 'Meeting invitation')}\n\n"
|
||||
+ ("Thank you for the invitation. I have a calendar conflict during the proposed "
|
||||
"time, so I must decline. Could we find another time?"
|
||||
if conflicts else
|
||||
"Thank you for the invitation. I checked the calendar and the proposed time is "
|
||||
"available. I am happy to accept.")
|
||||
+ "\n"
|
||||
)
|
||||
draft_path = self.campaign_dir / "artifacts" / "meeting_reply_draft.txt"
|
||||
draft_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
draft_path.write_text(draft, encoding="utf-8")
|
||||
return {
|
||||
"classification": "meeting_invitation",
|
||||
"calendar_check": {"performed": True, "calendar_id_sha256": sha256(calendar_id)[:20],
|
||||
"events_examined": len(events), "conflict_count": len(conflicts),
|
||||
"conflict": bool(conflicts), "start": iso_millis(start),
|
||||
"end": iso_millis(end)},
|
||||
"draft": {"disposition": disposition, "path": str(draft_path),
|
||||
"bytes": draft_path.stat().st_size,
|
||||
"sha256": sha256(draft_path.read_bytes())},
|
||||
}
|
||||
|
||||
def _complaint(self, email: dict[str, Any]) -> dict[str, Any]:
|
||||
text = email_text(email)
|
||||
order = re.search(r"order\s*#([A-Za-z0-9-]+)", text, re.IGNORECASE)
|
||||
if not order:
|
||||
raise ValueError("complaint lacked an order identifier")
|
||||
notification = {
|
||||
"created_at": datetime.now(UTC).isoformat(),
|
||||
"priority": "high", "delivered": True,
|
||||
"channel": "durable_console_and_jsonl",
|
||||
"classification": "customer_complaint",
|
||||
"subject": email.get("subject", ""),
|
||||
"order_reference": order.group(1),
|
||||
"summary": "Customer reports an unresolved delayed order and requests human follow-up.",
|
||||
}
|
||||
path = self.campaign_dir / "artifacts" / "high_priority_notifications.jsonl"
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with path.open("a", encoding="utf-8") as stream:
|
||||
stream.write(json.dumps(notification, ensure_ascii=False) + "\n")
|
||||
stream.flush()
|
||||
os.fsync(stream.fileno())
|
||||
print(f"HIGH PRIORITY: customer complaint for order #{order.group(1)}", file=sys.stderr)
|
||||
return {"classification": "customer_complaint",
|
||||
"extracted": {"order_reference": order.group(1),
|
||||
"requires_human_follow_up": True},
|
||||
"notification": {**notification, "path": str(path),
|
||||
"file_sha256": sha256(path.read_bytes())}}
|
||||
|
||||
def _marketing(self, email: dict[str, Any]) -> dict[str, Any]:
|
||||
folders = self.client.list_folders(self.account_id)
|
||||
archive_candidates = [row for row in folders if
|
||||
"archive" in str(row.get("role", "")).lower()
|
||||
or "archive" in str(row.get("name", "")).lower()]
|
||||
update = self.client.update_email_folders(str(email["id"]), ["archive"])
|
||||
verified = self.client.get_email(str(email["id"]))
|
||||
role = str(verified.get("role", "")).lower()
|
||||
folder_text = " ".join(str(value).lower()
|
||||
for value in (verified.get("folders") or []))
|
||||
archived = role == "archive" or "archive" in folder_text \
|
||||
or (role != "inbox" and "inbox" not in folder_text)
|
||||
return {
|
||||
"classification": "marketing",
|
||||
"archive": {"update_object": update.get("object"),
|
||||
"archive_folder_candidates": len(archive_candidates),
|
||||
"verified": archived,
|
||||
"verified_role": role,
|
||||
"verified_folders_sha256": sha256(folder_text)},
|
||||
}
|
||||
|
||||
def process_queue(self, emails: list[dict[str, Any]]) -> None:
|
||||
queue = deque(sorted(emails, key=provider_date))
|
||||
while queue:
|
||||
email = queue.popleft()
|
||||
sequence = len(self.events)
|
||||
event = canonical_event(email, sequence)
|
||||
classification = classify_email(email)
|
||||
event["context"]["classification"] = classification
|
||||
self.events.append(event)
|
||||
if classification == "meeting_invitation":
|
||||
workflow = self._meeting(email)
|
||||
elif classification == "customer_complaint":
|
||||
workflow = self._complaint(email)
|
||||
else:
|
||||
workflow = self._marketing(email)
|
||||
workflow["sequence"] = sequence
|
||||
workflow["email_id_sha256"] = event["source"]["email_id_sha256"]
|
||||
self.workflows.append(workflow)
|
||||
|
||||
|
||||
def seed_messages(client: UnipileClient, sender_account_id: str, recipient: str,
|
||||
campaign_id: str) -> list[dict[str, Any]]:
|
||||
start = (datetime.now(UTC) + timedelta(days=2)).replace(
|
||||
hour=10, minute=0, second=0, microsecond=0
|
||||
)
|
||||
end = start + timedelta(hours=1)
|
||||
messages = [
|
||||
(
|
||||
f"[EXP6-1 {campaign_id}] Meeting invitation: design review",
|
||||
"Meeting invitation for the agent experiment.\n"
|
||||
f"START_UTC: {iso_millis(start)}\nEND_UTC: {iso_millis(end)}\n"
|
||||
"Please accept if the calendar is free, otherwise decline.",
|
||||
),
|
||||
(
|
||||
f"[EXP6-1 {campaign_id}] Customer complaint: delayed order",
|
||||
f"Customer complaint for order #{campaign_id[-8:]}. The delivery is overdue and "
|
||||
"support has not resolved it. Please arrange urgent human follow-up.",
|
||||
),
|
||||
(
|
||||
f"[EXP6-1 {campaign_id}] Marketing newsletter",
|
||||
"Marketing newsletter: save 20 percent on productivity software. "
|
||||
"This bulk promotion includes an unsubscribe link.",
|
||||
),
|
||||
]
|
||||
return [client.send_email(sender_account_id, recipient, subject, body)
|
||||
for subject, body in messages]
|
||||
|
||||
|
||||
def inbox_folder(client: UnipileClient, account_id: str) -> str | None:
|
||||
folders = client.list_folders(account_id)
|
||||
inbox = next((row for row in folders if
|
||||
str(row.get("role", "")).lower() == "inbox"
|
||||
or str(row.get("name", "")).lower() == "inbox"), None)
|
||||
if not inbox:
|
||||
return None
|
||||
return str(inbox.get("provider_id") or inbox.get("id") or inbox.get("name"))
|
||||
|
||||
|
||||
def poll_campaign_emails(client: UnipileClient, account_id: str, campaign_id: str,
|
||||
after: datetime, *, timeout: float,
|
||||
interval: float) -> list[dict[str, Any]]:
|
||||
folder = inbox_folder(client, account_id)
|
||||
deadline = time.monotonic() + timeout
|
||||
found: dict[str, dict[str, Any]] = {}
|
||||
marker = f"[EXP6-1 {campaign_id}]"
|
||||
while time.monotonic() < deadline and len(found) < 3:
|
||||
for reference in client.list_emails(account_id, after=after, folder=folder):
|
||||
if marker not in str(reference.get("subject", "")):
|
||||
continue
|
||||
email_id = str(reference.get("id", ""))
|
||||
if email_id and email_id not in found:
|
||||
full = client.get_email(email_id)
|
||||
if str(full.get("role", reference.get("role", ""))).lower() == "sent":
|
||||
continue
|
||||
found[email_id] = full
|
||||
if len(found) < 3:
|
||||
time.sleep(interval)
|
||||
if len(found) != 3:
|
||||
raise TimeoutError(f"received {len(found)} of three campaign emails before timeout")
|
||||
return sorted(found.values(), key=provider_date)
|
||||
|
||||
|
||||
def official_schema_receipts(urls: list[str]) -> list[dict[str, Any]]:
|
||||
receipts = []
|
||||
with httpx.Client(timeout=30, follow_redirects=True) as client:
|
||||
for url in urls:
|
||||
try:
|
||||
response = client.get(url)
|
||||
receipts.append({"url": url, "status": response.status_code,
|
||||
"bytes": len(response.content),
|
||||
"sha256": sha256(response.content),
|
||||
"retrieved_at": datetime.now(UTC).isoformat()})
|
||||
except Exception as exc:
|
||||
receipts.append({"url": url, "status": None,
|
||||
"error_type": type(exc).__name__})
|
||||
return receipts
|
||||
|
||||
|
||||
def bearer_diagnostic(client: UnipileClient) -> dict[str, Any]:
|
||||
"""Preserve the rejected alternate auth probe without leaking its token."""
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
response = httpx.get(
|
||||
client.base_url + "/api/v1/accounts",
|
||||
headers={"Authorization": "Bearer " + client.access_token,
|
||||
"User-Agent": "ai-agent-book-experiment/4.4-diagnostic"},
|
||||
timeout=30,
|
||||
)
|
||||
try:
|
||||
payload = response.json()
|
||||
except ValueError:
|
||||
payload = {}
|
||||
return {"method": "GET", "path": "/api/v1/accounts",
|
||||
"credential_scheme": "Authorization: Bearer <redacted>",
|
||||
"status": response.status_code,
|
||||
"error_type": payload.get("type"), "error_title": payload.get("title"),
|
||||
"response_sha256": sha256(response.content),
|
||||
"latency_seconds": round(time.perf_counter() - started, 3)}
|
||||
except Exception as exc:
|
||||
return {"method": "GET", "path": "/api/v1/accounts",
|
||||
"credential_scheme": "Authorization: Bearer <redacted>",
|
||||
"status": None, "error_type": type(exc).__name__,
|
||||
"latency_seconds": round(time.perf_counter() - started, 3)}
|
||||
|
||||
|
||||
def derive_acceptance(events: list[dict[str, Any]], workflows: list[dict[str, Any]],
|
||||
calls: list[dict[str, Any]], seed_receipts: list[dict[str, Any]],
|
||||
*, credential_secret: str, dsn_secret: str) -> dict[str, Any]:
|
||||
by_class = {row.get("classification"): row for row in workflows}
|
||||
meeting = by_class.get("meeting_invitation", {})
|
||||
complaint = by_class.get("customer_complaint", {})
|
||||
marketing = by_class.get("marketing", {})
|
||||
encoded = canonical_json({"events": events, "workflows": workflows, "calls": calls})
|
||||
control_plane = canonical_json([{key: value for key, value in call.items()
|
||||
if key not in {"response_sha256"}}
|
||||
for call in calls])
|
||||
gates = {
|
||||
"three_real_inbound_unipile_events": (
|
||||
len(events) == 3 and all(
|
||||
event.get("source", {}).get("provider") == "unipile"
|
||||
and event.get("channel") == "unipile_mailbox_poll"
|
||||
and bool(event.get("provider_receipt_sha256")) for event in events
|
||||
)
|
||||
),
|
||||
"fifo_event_queue": [event.get("sequence") for event in events] == [0, 1, 2]
|
||||
and [row.get("sequence") for row in workflows] == [0, 1, 2],
|
||||
"exact_three_scenario_classifications": set(by_class) == {
|
||||
"meeting_invitation", "customer_complaint", "marketing"
|
||||
} and len(workflows) == 3,
|
||||
"meeting_calendar_checked_and_reply_drafted": (
|
||||
meeting.get("calendar_check", {}).get("performed") is True
|
||||
and meeting.get("calendar_check", {}).get("conflict") in {True, False}
|
||||
and meeting.get("draft", {}).get("disposition") in {"accept", "decline"}
|
||||
and meeting.get("draft", {}).get("bytes", 0) > 100
|
||||
and len(meeting.get("draft", {}).get("sha256", "")) == 64
|
||||
),
|
||||
"complaint_extracted_and_high_priority_notification_delivered": (
|
||||
complaint.get("extracted", {}).get("requires_human_follow_up") is True
|
||||
and bool(complaint.get("extracted", {}).get("order_reference"))
|
||||
and complaint.get("notification", {}).get("priority") == "high"
|
||||
and complaint.get("notification", {}).get("delivered") is True
|
||||
and len(complaint.get("notification", {}).get("file_sha256", "")) == 64
|
||||
),
|
||||
"marketing_archived_and_verified_through_provider": (
|
||||
marketing.get("archive", {}).get("update_object") == "EmailUpdated"
|
||||
and marketing.get("archive", {}).get("verified") is True
|
||||
),
|
||||
"required_unipile_calls_succeeded": bool(calls) and all(
|
||||
call.get("success") is True for call in calls
|
||||
),
|
||||
"three_seed_messages_sent_through_unipile": (
|
||||
len(seed_receipts) == 3 and all(isinstance(row, dict) for row in seed_receipts)
|
||||
),
|
||||
"credentials_and_dsn_absent_from_receipts": (
|
||||
credential_secret not in encoded and dsn_secret not in encoded
|
||||
),
|
||||
"no_simulation_markers_in_control_plane": not SIMULATION_PATTERN.search(control_plane),
|
||||
}
|
||||
return {"status": "passed" if all(gates.values()) else "failed", "gates": gates}
|
||||
|
||||
|
||||
def build_manifest(campaign_dir: Path, summary: dict[str, Any]) -> dict[str, Any]:
|
||||
files = []
|
||||
for path in sorted(campaign_dir.rglob("*")):
|
||||
if path.is_file() and path.name != "manifest.json":
|
||||
data = path.read_bytes()
|
||||
files.append({"path": str(path.relative_to(campaign_dir)),
|
||||
"bytes": len(data), "sha256": sha256(data)})
|
||||
return {
|
||||
"experiment": "6-1", "campaign_id": summary.get("campaign_id"),
|
||||
"generated_at": datetime.now(UTC).isoformat(),
|
||||
"status": summary.get("status"),
|
||||
"official_complete": summary.get("status") == "passed",
|
||||
"files": files,
|
||||
}
|
||||
|
||||
|
||||
def choose_account(accounts: list[dict[str, Any]], requested: str | None) -> dict[str, Any]:
|
||||
if requested:
|
||||
match = next((account for account in accounts if account.get("id") == requested), None)
|
||||
if not match:
|
||||
raise ValueError("requested account ID was not returned by Unipile")
|
||||
return match
|
||||
candidates = [account for account in accounts if
|
||||
"mail" in canonical_json(account.get("sources", [])).lower()
|
||||
or account_email(account)]
|
||||
if not candidates:
|
||||
raise RuntimeError("Unipile returned no mail-capable account")
|
||||
return candidates[0]
|
||||
|
||||
|
||||
def run(args: argparse.Namespace) -> Path:
|
||||
protocol = json.loads(PROTOCOL_PATH.read_text(encoding="utf-8"))
|
||||
campaign_id = args.campaign_id or datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
campaign_dir = VALIDATION_ROOT / campaign_id
|
||||
campaign_dir.mkdir(parents=True, exist_ok=False)
|
||||
write_json(campaign_dir / "protocol.json", protocol)
|
||||
docs = official_schema_receipts(protocol["official_schema_sources"])
|
||||
dsn = os.getenv("UNIPILE_DSN", "")
|
||||
token = os.getenv("UNIPILE_ACCESS_TOKEN", "")
|
||||
summary: dict[str, Any] = {
|
||||
"experiment": "6-1", "campaign_id": campaign_id,
|
||||
"generated_at": datetime.now(UTC).isoformat(),
|
||||
"provider": "unipile", "base_url_sha256": sha256(dsn) if dsn else None,
|
||||
"official_schema_receipts": docs,
|
||||
}
|
||||
client: UnipileClient | None = None
|
||||
try:
|
||||
client = UnipileClient(dsn, token)
|
||||
accounts = client.list_accounts()
|
||||
if args.preflight_only:
|
||||
summary.update({"status": "preflight_passed", "account_count": len(accounts),
|
||||
"official_complete": False,
|
||||
"account_receipts": [redacted(account) for account in accounts],
|
||||
"api_calls": client.calls})
|
||||
write_json(campaign_dir / "summary.json", summary)
|
||||
return campaign_dir
|
||||
|
||||
listen_account = choose_account(accounts, args.listen_account_id)
|
||||
sender_account = choose_account(accounts, args.sender_account_id) \
|
||||
if args.sender_account_id else listen_account
|
||||
calendar_account = choose_account(accounts, args.calendar_account_id) \
|
||||
if args.calendar_account_id else listen_account
|
||||
recipient = args.recipient or account_email(listen_account)
|
||||
if not recipient:
|
||||
raise RuntimeError("could not infer listener email; pass --recipient")
|
||||
started = datetime.now(UTC) - timedelta(minutes=1)
|
||||
seeds = seed_messages(client, str(sender_account["id"]), recipient, campaign_id)
|
||||
emails = poll_campaign_emails(
|
||||
client, str(listen_account["id"]), campaign_id, started,
|
||||
timeout=args.poll_timeout, interval=args.poll_interval,
|
||||
)
|
||||
experiment = MailboxExperiment(
|
||||
client, campaign_dir, str(listen_account["id"]), str(calendar_account["id"])
|
||||
)
|
||||
experiment.process_queue(emails)
|
||||
acceptance = derive_acceptance(
|
||||
experiment.events, experiment.workflows, client.calls, seeds,
|
||||
credential_secret=token, dsn_secret=dsn,
|
||||
)
|
||||
summary.update({
|
||||
"status": acceptance["status"], "listener": "unipile_mailbox_poll",
|
||||
"official_complete": acceptance["status"] == "passed",
|
||||
"account_receipts": {
|
||||
"listener": redacted(listen_account), "sender": redacted(sender_account),
|
||||
"calendar": redacted(calendar_account),
|
||||
},
|
||||
"seed_receipts": redacted(seeds), "events": experiment.events,
|
||||
"workflows": experiment.workflows, "api_calls": client.calls,
|
||||
"acceptance": acceptance,
|
||||
})
|
||||
except Exception as exc:
|
||||
credential_block = isinstance(exc, UnipileAPIError) and exc.status == 401
|
||||
if client and credential_block:
|
||||
summary["alternate_auth_diagnostic"] = bearer_diagnostic(client)
|
||||
summary.update({
|
||||
"status": "blocked" if credential_block or not dsn or not token else "failed",
|
||||
"official_complete": False,
|
||||
"blocker_or_error": {"type": type(exc).__name__, "message": str(exc)},
|
||||
"credentials_present": {"UNIPILE_DSN": bool(dsn),
|
||||
"UNIPILE_ACCESS_TOKEN": bool(token)},
|
||||
"api_calls": client.calls if client else [],
|
||||
"acceptance": {"status": "blocked" if credential_block or not dsn or not token else "failed", "gates": {
|
||||
"valid_unipile_credentials": False,
|
||||
"real_mailbox_campaign_completed": False,
|
||||
}},
|
||||
})
|
||||
finally:
|
||||
if client:
|
||||
client.close()
|
||||
write_json(campaign_dir / "summary.json", summary)
|
||||
write_json(campaign_dir / "manifest.json", build_manifest(campaign_dir, summary))
|
||||
write_json(VALIDATION_ROOT / "latest.json", {
|
||||
"experiment": "6-1", "campaign_id": campaign_id,
|
||||
"status": summary.get("status"),
|
||||
"official_complete": summary.get("status") == "passed",
|
||||
"manifest": str((campaign_dir / "manifest.json").relative_to(HERE)),
|
||||
"manifest_sha256": sha256((campaign_dir / "manifest.json").read_bytes()),
|
||||
})
|
||||
return campaign_dir
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--campaign-id")
|
||||
parser.add_argument("--preflight-only", action="store_true")
|
||||
parser.add_argument("--recipient",
|
||||
help="Listener mailbox address; inferred from account when omitted")
|
||||
parser.add_argument("--listen-account-id")
|
||||
parser.add_argument("--sender-account-id")
|
||||
parser.add_argument("--calendar-account-id")
|
||||
parser.add_argument("--poll-timeout", type=float, default=180)
|
||||
parser.add_argument("--poll-interval", type=float, default=5)
|
||||
args = parser.parse_args()
|
||||
path = run(args)
|
||||
print(path)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,15 @@
|
||||
{
|
||||
"generated_at": "2026-07-29T21:17:21.956016+00:00",
|
||||
"files": [
|
||||
{
|
||||
"path": "protocol.json",
|
||||
"bytes": 1924,
|
||||
"sha256": "1826c3aa1c2ce0af96277e13adb237f4c37b005c8708a113159ed768c296fbc1"
|
||||
},
|
||||
{
|
||||
"path": "summary.json",
|
||||
"bytes": 3995,
|
||||
"sha256": "3c8aa1272f9ddbb44ae9614af7183943c1bb39ea60d4ce0de21d923c3ee9cba3"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,55 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"title": "Real event-driven mailbox workflow",
|
||||
"authority": "book/chapter6.md:466",
|
||||
"mail_provider": "Unipile Email API",
|
||||
"listener": {
|
||||
"mode": "polling",
|
||||
"endpoint": "GET /api/v1/emails",
|
||||
"event_channel": "unipile_mailbox_poll",
|
||||
"queue": "FIFO by provider timestamp then email id"
|
||||
},
|
||||
"scenarios": [
|
||||
{
|
||||
"classification": "meeting_invitation",
|
||||
"required_actions": [
|
||||
"live_calendar_conflict_check",
|
||||
"accept_or_decline_draft"
|
||||
]
|
||||
},
|
||||
{
|
||||
"classification": "customer_complaint",
|
||||
"required_actions": [
|
||||
"key_information_extraction",
|
||||
"high_priority_notification"
|
||||
]
|
||||
},
|
||||
{
|
||||
"classification": "marketing",
|
||||
"required_actions": [
|
||||
"provider_archive_update",
|
||||
"post_update_verification"
|
||||
]
|
||||
}
|
||||
],
|
||||
"official_schema_sources": [
|
||||
"https://developer.unipile.com/reference/accountscontroller_listaccounts.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_listmails.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_getmail.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_updatemail.md",
|
||||
"https://developer.unipile.com/reference/folderscontroller_listfolders.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendars.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendareventsbycalendar.md",
|
||||
"https://developer.unipile.com/docs/new-emails-webhook.md"
|
||||
],
|
||||
"acceptance": {
|
||||
"no_local_or_mock_mailbox_substitute": true,
|
||||
"three_real_inbound_email_objects": true,
|
||||
"calendar_query_receipted": true,
|
||||
"draft_artifact_hashed": true,
|
||||
"high_priority_notification_delivered": true,
|
||||
"marketing_email_archived_and_verified": true,
|
||||
"identifiers_and_credentials_redacted": true,
|
||||
"fail_closed_on_missing_or_invalid_credentials": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,117 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"campaign_id": "credential_probe_20260730T050900Z",
|
||||
"generated_at": "2026-07-29T21:17:18.847328+00:00",
|
||||
"provider": "unipile",
|
||||
"base_url_sha256": "28b52f572e99947db7fa2192ba95435ed4d01404e53a40c2b4f6de2274513a9b",
|
||||
"official_schema_receipts": [
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/accountscontroller_listaccounts.md",
|
||||
"status": 200,
|
||||
"bytes": 151069,
|
||||
"sha256": "e63a180008549782b9b946734aea82bc2b91e50d12029f21617b1944e50368fd",
|
||||
"retrieved_at": "2026-07-29T21:17:14.234772+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_listmails.md",
|
||||
"status": 200,
|
||||
"bytes": 102744,
|
||||
"sha256": "6cefdb5ae883c8573b5c5d5b4cc3b6f2168909e224afed419209578ae7bf6b25",
|
||||
"retrieved_at": "2026-07-29T21:17:14.881533+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_getmail.md",
|
||||
"status": 200,
|
||||
"bytes": 88581,
|
||||
"sha256": "c3b5c98505a1dbb4433f67651796124dd3eb7b044a264581dc86cf31ae994d86",
|
||||
"retrieved_at": "2026-07-29T21:17:15.567230+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_updatemail.md",
|
||||
"status": 200,
|
||||
"bytes": 16690,
|
||||
"sha256": "26d1c3f2c6e0d75fb1445a220a90c701e95790ba33e31a0c89088c13ac393647",
|
||||
"retrieved_at": "2026-07-29T21:17:16.643818+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/folderscontroller_listfolders.md",
|
||||
"status": 200,
|
||||
"bytes": 22162,
|
||||
"sha256": "d35d2042bd261314af6597ffdf9613d2e3df799e062e1f598708e07cdca1b957",
|
||||
"retrieved_at": "2026-07-29T21:17:17.272871+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/calendarscontroller_listcalendars.md",
|
||||
"status": 200,
|
||||
"bytes": 20671,
|
||||
"sha256": "cf35a6a277b9390bf937dfe8300788c348999652903ee91ca07b2f9ff86cf8da",
|
||||
"retrieved_at": "2026-07-29T21:17:17.902366+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/calendarscontroller_listcalendareventsbycalendar.md",
|
||||
"status": 200,
|
||||
"bytes": 36444,
|
||||
"sha256": "7f4b3414f9d4a78a23bf4a87aa3d82f05eea2e00b101063cc82694c8335d96e3",
|
||||
"retrieved_at": "2026-07-29T21:17:18.521682+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/docs/new-emails-webhook.md",
|
||||
"status": 200,
|
||||
"bytes": 1729,
|
||||
"sha256": "50c4808b875df5d71a3734ff4621b557d6fd40380e253cb690cc70d61148ee6a",
|
||||
"retrieved_at": "2026-07-29T21:17:18.847210+00:00"
|
||||
}
|
||||
],
|
||||
"alternate_auth_diagnostic": {
|
||||
"method": "GET",
|
||||
"path": "/api/v1/accounts",
|
||||
"credential_scheme": "Authorization: Bearer <redacted>",
|
||||
"status": 401,
|
||||
"error_type": "errors/invalid_credentials",
|
||||
"error_title": "Invalid credentials",
|
||||
"response_sha256": "35feddea5acbb597bdd6395c36c393c077bf94090fb311652a85e10e464c6b59",
|
||||
"latency_seconds": 1.57
|
||||
},
|
||||
"status": "blocked",
|
||||
"blocker_or_error": {
|
||||
"type": "UnipileAPIError",
|
||||
"message": "GET /api/v1/accounts returned 401: errors/missing_credentials"
|
||||
},
|
||||
"credentials_present": {
|
||||
"UNIPILE_DSN": true,
|
||||
"UNIPILE_ACCESS_TOKEN": true
|
||||
},
|
||||
"api_calls": [
|
||||
{
|
||||
"method": "GET",
|
||||
"path": "/api/v1/accounts",
|
||||
"request": {
|
||||
"params": {
|
||||
"limit": 100
|
||||
},
|
||||
"json": null,
|
||||
"multipart": null
|
||||
},
|
||||
"credential_scheme": "X-API-KEY",
|
||||
"status": 401,
|
||||
"success": false,
|
||||
"latency_seconds": 1.522,
|
||||
"response_sha256": "47c9318169e1bdccdcd71fb77189991e092db3344b06c1f59b8e1e426296babc",
|
||||
"response_bytes": 80,
|
||||
"response_shape": [
|
||||
"status",
|
||||
"title",
|
||||
"type"
|
||||
],
|
||||
"error_type": "errors/missing_credentials",
|
||||
"error_title": "Missing credentials"
|
||||
}
|
||||
],
|
||||
"acceptance": {
|
||||
"status": "failed",
|
||||
"gates": {
|
||||
"valid_unipile_credentials": false,
|
||||
"real_mailbox_campaign_completed": false
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"campaign_id": "credential_probe_20260730T064500Z",
|
||||
"generated_at": "2026-07-29T22:23:37.271474+00:00",
|
||||
"status": "blocked",
|
||||
"official_complete": false,
|
||||
"files": [
|
||||
{
|
||||
"path": "protocol.json",
|
||||
"bytes": 1924,
|
||||
"sha256": "1826c3aa1c2ce0af96277e13adb237f4c37b005c8708a113159ed768c296fbc1"
|
||||
},
|
||||
{
|
||||
"path": "summary.json",
|
||||
"bytes": 4027,
|
||||
"sha256": "413f8d00bf0d82fd4bf86913988f2b09907a9bc8b3bec2da2ffadab2f969f45c"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,55 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"title": "Real event-driven mailbox workflow",
|
||||
"authority": "book/chapter6.md:466",
|
||||
"mail_provider": "Unipile Email API",
|
||||
"listener": {
|
||||
"mode": "polling",
|
||||
"endpoint": "GET /api/v1/emails",
|
||||
"event_channel": "unipile_mailbox_poll",
|
||||
"queue": "FIFO by provider timestamp then email id"
|
||||
},
|
||||
"scenarios": [
|
||||
{
|
||||
"classification": "meeting_invitation",
|
||||
"required_actions": [
|
||||
"live_calendar_conflict_check",
|
||||
"accept_or_decline_draft"
|
||||
]
|
||||
},
|
||||
{
|
||||
"classification": "customer_complaint",
|
||||
"required_actions": [
|
||||
"key_information_extraction",
|
||||
"high_priority_notification"
|
||||
]
|
||||
},
|
||||
{
|
||||
"classification": "marketing",
|
||||
"required_actions": [
|
||||
"provider_archive_update",
|
||||
"post_update_verification"
|
||||
]
|
||||
}
|
||||
],
|
||||
"official_schema_sources": [
|
||||
"https://developer.unipile.com/reference/accountscontroller_listaccounts.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_listmails.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_getmail.md",
|
||||
"https://developer.unipile.com/reference/mailscontroller_updatemail.md",
|
||||
"https://developer.unipile.com/reference/folderscontroller_listfolders.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendars.md",
|
||||
"https://developer.unipile.com/reference/calendarscontroller_listcalendareventsbycalendar.md",
|
||||
"https://developer.unipile.com/docs/new-emails-webhook.md"
|
||||
],
|
||||
"acceptance": {
|
||||
"no_local_or_mock_mailbox_substitute": true,
|
||||
"three_real_inbound_email_objects": true,
|
||||
"calendar_query_receipted": true,
|
||||
"draft_artifact_hashed": true,
|
||||
"high_priority_notification_delivered": true,
|
||||
"marketing_email_archived_and_verified": true,
|
||||
"identifiers_and_credentials_redacted": true,
|
||||
"fail_closed_on_missing_or_invalid_credentials": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,118 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"campaign_id": "credential_probe_20260730T064500Z",
|
||||
"generated_at": "2026-07-29T22:23:34.029339+00:00",
|
||||
"provider": "unipile",
|
||||
"base_url_sha256": "28b52f572e99947db7fa2192ba95435ed4d01404e53a40c2b4f6de2274513a9b",
|
||||
"official_schema_receipts": [
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/accountscontroller_listaccounts.md",
|
||||
"status": 200,
|
||||
"bytes": 151069,
|
||||
"sha256": "e63a180008549782b9b946734aea82bc2b91e50d12029f21617b1944e50368fd",
|
||||
"retrieved_at": "2026-07-29T22:23:30.036928+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_listmails.md",
|
||||
"status": 200,
|
||||
"bytes": 102744,
|
||||
"sha256": "6cefdb5ae883c8573b5c5d5b4cc3b6f2168909e224afed419209578ae7bf6b25",
|
||||
"retrieved_at": "2026-07-29T22:23:30.697071+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_getmail.md",
|
||||
"status": 200,
|
||||
"bytes": 88581,
|
||||
"sha256": "c3b5c98505a1dbb4433f67651796124dd3eb7b044a264581dc86cf31ae994d86",
|
||||
"retrieved_at": "2026-07-29T22:23:31.325776+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/mailscontroller_updatemail.md",
|
||||
"status": 200,
|
||||
"bytes": 16690,
|
||||
"sha256": "26d1c3f2c6e0d75fb1445a220a90c701e95790ba33e31a0c89088c13ac393647",
|
||||
"retrieved_at": "2026-07-29T22:23:31.902105+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/folderscontroller_listfolders.md",
|
||||
"status": 200,
|
||||
"bytes": 22162,
|
||||
"sha256": "d35d2042bd261314af6597ffdf9613d2e3df799e062e1f598708e07cdca1b957",
|
||||
"retrieved_at": "2026-07-29T22:23:32.485289+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/calendarscontroller_listcalendars.md",
|
||||
"status": 200,
|
||||
"bytes": 20671,
|
||||
"sha256": "cf35a6a277b9390bf937dfe8300788c348999652903ee91ca07b2f9ff86cf8da",
|
||||
"retrieved_at": "2026-07-29T22:23:33.169436+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/reference/calendarscontroller_listcalendareventsbycalendar.md",
|
||||
"status": 200,
|
||||
"bytes": 36444,
|
||||
"sha256": "7f4b3414f9d4a78a23bf4a87aa3d82f05eea2e00b101063cc82694c8335d96e3",
|
||||
"retrieved_at": "2026-07-29T22:23:33.710013+00:00"
|
||||
},
|
||||
{
|
||||
"url": "https://developer.unipile.com/docs/new-emails-webhook.md",
|
||||
"status": 200,
|
||||
"bytes": 1729,
|
||||
"sha256": "50c4808b875df5d71a3734ff4621b557d6fd40380e253cb690cc70d61148ee6a",
|
||||
"retrieved_at": "2026-07-29T22:23:34.029079+00:00"
|
||||
}
|
||||
],
|
||||
"alternate_auth_diagnostic": {
|
||||
"method": "GET",
|
||||
"path": "/api/v1/accounts",
|
||||
"credential_scheme": "Authorization: Bearer <redacted>",
|
||||
"status": 401,
|
||||
"error_type": "errors/invalid_credentials",
|
||||
"error_title": "Invalid credentials",
|
||||
"response_sha256": "35feddea5acbb597bdd6395c36c393c077bf94090fb311652a85e10e464c6b59",
|
||||
"latency_seconds": 1.568
|
||||
},
|
||||
"status": "blocked",
|
||||
"official_complete": false,
|
||||
"blocker_or_error": {
|
||||
"type": "UnipileAPIError",
|
||||
"message": "GET /api/v1/accounts returned 401: errors/missing_credentials"
|
||||
},
|
||||
"credentials_present": {
|
||||
"UNIPILE_DSN": true,
|
||||
"UNIPILE_ACCESS_TOKEN": true
|
||||
},
|
||||
"api_calls": [
|
||||
{
|
||||
"method": "GET",
|
||||
"path": "/api/v1/accounts",
|
||||
"request": {
|
||||
"params": {
|
||||
"limit": 100
|
||||
},
|
||||
"json": null,
|
||||
"multipart": null
|
||||
},
|
||||
"credential_scheme": "X-API-KEY",
|
||||
"status": 401,
|
||||
"success": false,
|
||||
"latency_seconds": 1.656,
|
||||
"response_sha256": "47c9318169e1bdccdcd71fb77189991e092db3344b06c1f59b8e1e426296babc",
|
||||
"response_bytes": 80,
|
||||
"response_shape": [
|
||||
"status",
|
||||
"title",
|
||||
"type"
|
||||
],
|
||||
"error_type": "errors/missing_credentials",
|
||||
"error_title": "Missing credentials"
|
||||
}
|
||||
],
|
||||
"acceptance": {
|
||||
"status": "blocked",
|
||||
"gates": {
|
||||
"valid_unipile_credentials": false,
|
||||
"real_mailbox_campaign_completed": false
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"experiment": "6-1",
|
||||
"campaign_id": "credential_probe_20260730T064500Z",
|
||||
"status": "blocked",
|
||||
"official_complete": false,
|
||||
"manifest": "validation/experiment_6_1/credential_probe_20260730T064500Z/manifest.json",
|
||||
"manifest_sha256": "3f689dfee915503f61ca30e9b590e24c8950496ca90fbf365def83805e877d0a"
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
# Python
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*.egg-info/
|
||||
.venv/
|
||||
venv/
|
||||
|
||||
# Env / secrets
|
||||
.env
|
||||
|
||||
# Generated artifacts
|
||||
output/
|
||||
logs/
|
||||
checkpoints/
|
||||
*.log
|
||||
!experiment_protocol.json
|
||||
!validation/
|
||||
!validation/**
|
||||
@@ -0,0 +1,637 @@
|
||||
# Asynchronous Agent with Parallel Execution and Interruption / 带并行执行和打断能力的异步 Agent
|
||||
|
||||
> Companion code for *AI Agents in Depth*, Chapter 6 — **Experiment 6-2 ★★★**. Event-driven async Agent framework (Flux): parallel tools, interrupt/cancel, state checkpoints.
|
||||
> 配套《深入理解 AI Agent》第 4 章 **实验 6-2 ★★★**。事件驱动异步 Agent 框架(Flux):并行工具、打断取消、状态检查点。
|
||||
|
||||
← [Chapter 4 index / 返回第 4 章目录](../README.md)
|
||||
|
||||
---
|
||||
|
||||
## English
|
||||
|
||||
This directory is the runnable code for Experiment 6-2. It implements the core of the event-driven asynchronous Agent framework (Flux) described in [`agent_framework_design.md`](./agent_framework_design.md).
|
||||
|
||||
Building on the simple event queue of 4-5, this experiment goes deeper into async Agents and focuses on four things: **async tool execution, event queues and batching, interruption, and cancel/status query for parallel tools**. The Agent must manage concurrent tasks, handle interrupt and recovery, and decide from real-time state.
|
||||
|
||||
Two usage paths:
|
||||
|
||||
- **Offline demos (recommended first; zero deps, no API key)**: three measurable capabilities—**parallel vs serial wall-clock, interrupt/cancel then recover, checkpoint persist and restore**. No network, no LLM, no need for `openai`; `python demo.py` runs as-is.
|
||||
- **LLM scenarios (four book verification scenes)**: decisions by a real LLM (default OpenAI `gpt-5.6-luna`, function calling); API key required.
|
||||
|
||||
Both paths share the same async runtime. Long-running work uses **real, allowlisted Python subprocesses**. No shell is opened: progress is parsed from child stdout, completion includes observed hashes/return codes, and cancellation terminates the child PID.
|
||||
|
||||
## Code map
|
||||
|
||||
- **Run first:** python demo.py (offline, no API key).
|
||||
- **Start here:** runtime.py::AgentRuntime
|
||||
- **Core behavior:** runtime.py::_dispatcher routes urgency; _handle_interrupt cancels at a safe point; run_llm_turn advances the trajectory.
|
||||
- **State / protocol:** Event, inbox, pending batch and checkpoint files.
|
||||
- **Verifier:** offline demo assertions and run_real_experiment.py evidence gates.
|
||||
- **Experiment variable:** serial vs parallel tools, interrupt timing and checkpoint restore.
|
||||
- **Skip on first pass:** provider clients and the frontend/demo formatting.
|
||||
|
||||
### Architecture
|
||||
|
||||
Corresponds to section 5 of the design doc—all single-threaded `asyncio`:
|
||||
|
||||
```
|
||||
┌──────────────┐
|
||||
user msg / interrupt ──▶ │ inbox │ all inbound raw events
|
||||
async task completion ──▶ │ (asyncio.Q) │
|
||||
└──────┬───────┘
|
||||
│
|
||||
┌──────────▼───────────┐ classify_urgency()
|
||||
│ _dispatcher │──▶ interrupt / immediate / deferred
|
||||
└──────────┬───────────┘
|
||||
┌────────────────┼───────────────────┐
|
||||
INTERRUPT │ IMMEDIATE│ DEFERRED│
|
||||
cancel current turn+async direct to work pending buffer;
|
||||
tools; leave trace batch when async.result arrives
|
||||
┌──────────▼───────────┐
|
||||
│ work │ event batches to process
|
||||
└──────────┬───────────┘
|
||||
┌──────────▼───────────┐
|
||||
│ _worker │ per batch: append trajectory -> run_llm_turn()
|
||||
│ turn_task cancelable│ (on interrupt, cancel this child task)
|
||||
└──────────────────────┘
|
||||
|
||||
TaskManager: bounded real subprocesses (start / query / cancel / cancel_all)
|
||||
natural completion -> inject as new event (async.result) into inbox
|
||||
```
|
||||
|
||||
Code files:
|
||||
|
||||
| File | Role |
|
||||
|------|------|
|
||||
| `events.py` | `Event` model (checkpoint `to_dict`/`from_dict`), event types, **urgency** `classify_urgency()` |
|
||||
| `tasks.py` | Allowlisted subprocess `TaskManager` (stdout progress, OS cancel/query by id, executable receipts, `snapshot`/`restore`) |
|
||||
| `analysis_worker.py` | Real executable that hashes and analyzes `book/chapter4.md` while emitting progress |
|
||||
| `run_real_experiment.py` | Durable, model-independent acceptance campaign for all four manuscript scenarios |
|
||||
| `runtime.py` | `AgentRuntime`: event loop, two processing modes, LLM function calling, tool exec, `save_checkpoint`/`load_checkpoint` |
|
||||
| `async_demos.py` | Three **offline demos** (no API key): parallel wall-clock, interrupt/recover, state checkpoint |
|
||||
| `demo.py` | Unified CLI (argparse subcommands): offline demos + four LLM scenarios |
|
||||
|
||||
#### Two event-processing mechanisms (design doc 5.1)
|
||||
|
||||
- **Cancellation-based**: urgent events (user “cancel/stop”) immediately cancel the in-flight LLM turn and all background async tools; write interrupt event + cancel receipts into the trajectory.
|
||||
- **Queued**: non-urgent events (supplementary instructions) go to a `pending` buffer without interrupting work; when an async tool finishes and emits `async.result`, pending events are batch-appended to the trajectory, then one LLM turn runs.
|
||||
|
||||
Urgency rules (simple and explainable):
|
||||
|
||||
1. Interrupt keywords (cancel/stop/停止…) → `INTERRUPT` (cancellation-based)
|
||||
2. A question (question mark or interrogative, e.g. “what time is it?”) → `IMMEDIATE` (reply now, **do not** cancel background tasks)
|
||||
3. Other supplementary instructions (e.g. “reply in Japanese”) → `DEFERRED` (queue, batch)
|
||||
|
||||
#### Async tools
|
||||
|
||||
`run_terminal_command` is **async**: it returns a `task_id` placeholder immediately, then starts an allowlisted child process with `asyncio.create_subprocess_exec` and `shell=False`. Progress comes from the process's stdout. On completion, file metrics, input/stdout hashes, PID, return code, and duration are injected as a **new event** (`async.result`). `cancel_task` sends termination to that PID. Also: `query_task` by id and `get_current_time` for immediate questions.
|
||||
|
||||
**Timeline acceleration**: one logical progress tick maps to `0.4` wall-clock seconds by default (`FLUX_TICK_REAL` tunable). The real executables retain the **3% / 2% / 1% per tick** rates and **50%** threshold.
|
||||
|
||||
### How to run
|
||||
|
||||
CLI entry is `demo.py` with argparse subcommands; `python demo.py --help` for full usage.
|
||||
|
||||
For canonical, machine-readable evidence across all four scenarios:
|
||||
|
||||
```bash
|
||||
python run_real_experiment.py --tick-real 0.15
|
||||
pytest -q test_tasks_env.py test_real_tasks.py
|
||||
```
|
||||
|
||||
The campaign takes about 20 seconds and writes per-scenario receipts, Japanese
|
||||
HTML, the integrated report, acceptance gates, and a hash manifest under
|
||||
`validation/experiment_6_2/`.
|
||||
|
||||
#### Offline demos (no API key, out of the box)
|
||||
|
||||
```bash
|
||||
cd chapter4/async-agent
|
||||
|
||||
python demo.py # default: run all three offline demos in order
|
||||
python demo.py offline # same: explicit sequential offline demos
|
||||
python demo.py parallel # capability 1: parallel vs serial wall-clock (prints speedup)
|
||||
python demo.py interrupt # capability 2: interrupt/cancel mid-task, then recover
|
||||
python demo.py state # capability 3: checkpoint persist + cross-session restore + verify
|
||||
```
|
||||
|
||||
These demos do not network, call LLMs, or require `openai`—pure `asyncio` measures parallel speedup, frozen state after interrupt, and checkpoint save/restore.
|
||||
|
||||
#### Tests (offline)
|
||||
|
||||
The automated regression tests live in `tests/` and do not require an API key.
|
||||
|
||||
```bash
|
||||
# From the repository root, include the dev extra for pytest:
|
||||
uv sync --locked --python 3.12 --extra ch4 --extra dev
|
||||
|
||||
# pip testing fallback:
|
||||
# python -m pip install -e ".[ch4,dev]"
|
||||
|
||||
cd chapter4/async-agent
|
||||
python -m pytest tests
|
||||
```
|
||||
|
||||
#### LLM verification scenarios (four book scenes; API key required)
|
||||
|
||||
```bash
|
||||
# From the repository root: use the shared Chapter 4 environment
|
||||
uv sync --locked --python 3.12 --extra ch4
|
||||
|
||||
# Activate it before changing directories:
|
||||
# macOS/Linux:
|
||||
source .venv/bin/activate
|
||||
# Windows PowerShell: .venv\Scripts\Activate.ps1
|
||||
# Windows cmd: .venv\Scripts\activate.bat
|
||||
|
||||
# pip fallback when uv is not installed:
|
||||
# python -m pip install -e ".[ch4]"
|
||||
|
||||
cd chapter4/async-agent
|
||||
|
||||
# Single-project compatibility path, still supported during migration:
|
||||
# python -m pip install -r requirements.txt
|
||||
|
||||
cp env.example .env # set OPENAI_API_KEY
|
||||
|
||||
python demo.py scenarios # all four scenarios
|
||||
python demo.py scenarios --scenario 1 # scenario 1 only (async exec + immediate question)
|
||||
python demo.py scenarios --scenario 3 # scenario 3 only (interrupt)
|
||||
```
|
||||
|
||||
Default OpenAI `gpt-5.6-luna`. Other OpenAI-compatible providers:
|
||||
|
||||
```bash
|
||||
# Moonshot (default model: reasoning model kimi-k3)
|
||||
LLM_PROVIDER=moonshot python demo.py scenarios --scenario 1
|
||||
# Volcengine ARK (LLM_MODEL = inference endpoint id)
|
||||
LLM_PROVIDER=ark LLM_MODEL=ep-xxxx python demo.py scenarios --scenario 1
|
||||
# Alibaba Cloud Model Studio / Bailian (Qwen)
|
||||
LLM_PROVIDER=dashscope DASHSCOPE_API_KEY=your-key LLM_MODEL=qwen3.7-plus python demo.py scenarios --scenario 1
|
||||
```
|
||||
|
||||
> **Universal OpenRouter fallback**: if `OPENAI_API_KEY` is unset (and not moonshot/ark), with `OPENROUTER_API_KEY` set, `demo.py` routes via OpenRouter and maps model ids to `provider/model` (`gpt-*` → `openai/…`, `claude-*` → `anthropic/claude-opus-4.8`, ids with `/` pass through). Or set `LLM_PROVIDER=openrouter` explicitly. Example:
|
||||
> `OPENROUTER_API_KEY=your-openrouter-api-key LLM_MODEL=openai/gpt-5.6-luna python demo.py scenarios --scenario 1`
|
||||
|
||||
> Moonshot defaults to **reasoning model `kimi-k3`** (older `kimi-k2-*-preview` / `moonshot-v1-*` are outdated/retired). Reasoning models need `temperature=1` and `max_tokens>=2048`; `demo.py` applies these automatically by model.
|
||||
|
||||
> Legacy: `python demo.py --scenario N` is equivalent to `scenarios --scenario N`.
|
||||
|
||||
Log sources are color-coded: `USER`, `AGENT`, `TOOL`, `TASK`, `TRAJ`, `STATE`, `SYSTEM`.
|
||||
|
||||
### Three offline capabilities (real measured output)
|
||||
|
||||
The following are excerpts from **real runs** (no API key).
|
||||
|
||||
#### Capability 1: parallel vs serial tools (`python demo.py parallel`)
|
||||
|
||||
Four independent read-only perception-style tools (read file / search / DB / vector retrieval): serial `await` vs parallel `asyncio.gather`:
|
||||
|
||||
```
|
||||
── 结果对比 ─────────────────────────────────────────────
|
||||
串行总耗时(Σ 各工具) 4.51s
|
||||
并行总耗时(gather) 1.50s
|
||||
并行理论下界(最慢单个) 1.50s
|
||||
加速比 = 串行 / 并行 3.00x
|
||||
─────────────────────────────────────────────────────────
|
||||
```
|
||||
|
||||
Wall-clock drops from sum-of-tools to max-single—quantifying “read-only perception tools naturally parallelize.”
|
||||
|
||||
#### Capability 2: interrupt / cancel / recover (`python demo.py interrupt`)
|
||||
|
||||
Three parallel background tasks; user asks a question first (does not block tasks), then sends “cancel”:
|
||||
|
||||
```
|
||||
[ 1.00s] USER | (即时提问)现在几点了?
|
||||
[ 1.00s] AGENT | 现在 00:14:26。三个后台任务仍在并行推进,未被这次提问阻塞。
|
||||
[ 2.00s] USER | (打断)取消
|
||||
[ 2.00s] TASK | T1 已被取消 🛑(进度停在 39%)
|
||||
[ 2.00s] TASK | T2 已被取消 🛑(进度停在 26%)
|
||||
[ 2.00s] TASK | T3 已被取消 🛑(进度停在 13%)
|
||||
|
||||
── 打断后各任务状态(进度冻结在中途)───────────────────
|
||||
task_id 命令 状态 进度
|
||||
T1 python analyze_fast.py cancelled 39%
|
||||
T2 python analyze_mid.py cancelled 26%
|
||||
T3 python analyze_slow.py cancelled 13%
|
||||
─────────────────────────────────────────────────────────
|
||||
[ 2.05s] SYSTEM | 打断处理完毕,系统恢复空闲,可继续接受新任务……
|
||||
[ 5.52s] TASK | T4 完成 ✅
|
||||
[ 5.52s] AGENT | 已从打断中恢复,新任务 T4 正常完成:……
|
||||
```
|
||||
|
||||
Interrupt freezes cancelled tasks’ progress; the runtime itself stays healthy and can finish new work immediately.
|
||||
|
||||
#### Capability 3: state checkpoint persist and restore (`python demo.py state`)
|
||||
|
||||
Session A produces a trajectory + two running background tasks, saved to `checkpoints/agent_state.json`; session B restores with a fresh runtime and verifies:
|
||||
|
||||
```
|
||||
── 恢复校验 ─────────────────────────────────────────────
|
||||
轨迹事件数 保存前 3 -> 恢复后 3 [一致 ✓]
|
||||
可重建 LLM 上下文消息 4 条(system + 轨迹回放)
|
||||
task_id 命令 保存前进度 恢复后状态 进度
|
||||
T1 python analyze_fast.py 21% suspended 21%
|
||||
T2 python analyze_slow.py 7% suspended 7%
|
||||
─────────────────────────────────────────────────────────
|
||||
```
|
||||
|
||||
Trajectory and task progress fully persist across sessions; running tasks restore as `suspended` with last known progress for upper layers to “re-run” or “continue from progress.”
|
||||
|
||||
### Four LLM verification scenarios
|
||||
|
||||
#### Scenario 1: async tool execution
|
||||
Agent runs a long terminal command; user inserts “what time is it?”. Because the long command is async and non-blocking, the Agent answers immediately with `get_current_time`, then presents analysis when the background task finishes.
|
||||
|
||||
#### Scenario 2: event queue and batching
|
||||
During a long task, user sends “reply in Japanese” then “format as a webpage.” These non-urgent instructions queue; on task completion the framework **batch-appends** them; the Agent outputs Japanese HTML.
|
||||
|
||||
#### Scenario 3: interruption
|
||||
During a long task, user says “cancel.” The framework cancels the current turn and background async tools, recording `user.interrupt` and a `system.note` with cancelled task ids.
|
||||
|
||||
#### Scenario 4: parallel cancel and status query
|
||||
User: “run these three scripts at once; when the first finishes, query the others’ progress; cancel any under 50%.” Speeds 3% / 2% / 1% per second. Agent starts three async tasks; after the fastest finishes, queries the others (~66% and ~33%), cancels the under-50% one, and reports when the rest complete.
|
||||
|
||||
### Real LLM scenario output (key fragments)
|
||||
|
||||
> Real calls to `gpt-5.6-luna` (OpenAI-compatible); timestamps are real seconds; need API key to reproduce.
|
||||
|
||||
**Scenario 1 (async + immediate question)**
|
||||
```
|
||||
[ 3.97s] AGENT | 任务已在后台启动(task_id:T1)。完成后我会根据日志分析结果给出结论。
|
||||
[ 4.96s] TASK | T1 `python analyze_logs.py` 进度 22% ← 任务仍在后台跑
|
||||
[ 5.19s] TOOL | get_current_time -> 2026-07-18 13:43:30 ← 即时提问先回应
|
||||
[ 6.91s] AGENT | 现在是 2026 年 7 月 18 日 13:43:30。
|
||||
[12.19s] TRAJ | + async.result 异步完成 T1 ← 真实结果作为新事件注入
|
||||
[16.67s] AGENT | 日志分析已完成,结论如下:共扫描 12,840 条记录… ← 再呈现分析
|
||||
```
|
||||
|
||||
**Scenario 2 (batching)**
|
||||
```
|
||||
[ 1.50s] SYSTEM | 事件进入排队缓冲(当前积压 1 条)
|
||||
[ 1.90s] SYSTEM | 事件进入排队缓冲(当前积压 2 条)
|
||||
[12.05s] TASK | T1 完成 ✅
|
||||
[12.05s] SYSTEM | 异步结果到达,批量处理 2 条积压的非紧急事件
|
||||
[12.06s] TRAJ | + async.result 异步完成 T1
|
||||
[12.06s] TRAJ | + user.input 记得最后用日语回复
|
||||
[12.06s] TRAJ | + user.input 把结果整理成一个网页(HTML)
|
||||
...
|
||||
[22.38s] AGENT | <!DOCTYPE html>…<h2>分析結論</h2>… (批量指令一次性满足:日语 + HTML)
|
||||
```
|
||||
|
||||
**Scenario 3 (interrupt)**
|
||||
```
|
||||
[ 2.40s] TASK | 启动异步任务 T1: `python analyze_logs.py` (速度 4%/模拟秒)
|
||||
[ 4.00s] USER | (interrupt) 取消
|
||||
[ 4.00s] TASK | T1 已被取消 🛑(进度停在 14%)
|
||||
[ 4.00s] TRAJ | + user.interrupt 用户打断:取消
|
||||
[ 4.00s] TRAJ | + system.note 打断回执,取消任务 ['T1']
|
||||
[ 5.04s] AGENT | 已停止后台任务 T1。
|
||||
```
|
||||
|
||||
**Scenario 4 (parallel + status + 50% cancel + report)**
|
||||
```
|
||||
[ 2.82s] TASK | 启动异步任务 T1: `python analyze_fast.py` (速度 3%/模拟秒)
|
||||
[ 2.82s] TASK | 启动异步任务 T2: `python analyze_mid.py` (速度 2%/模拟秒)
|
||||
[ 2.82s] TASK | 启动异步任务 T3: `python analyze_slow.py` (速度 1%/模拟秒)
|
||||
[16.47s] TASK | T1 完成 ✅ ← 最快脚本先完成
|
||||
[19.84s] TOOL | query_task(T2) -> running 84% ← 查询其余两个进度
|
||||
[19.84s] TOOL | query_task(T3) -> running 42%
|
||||
[21.93s] TOOL | cancel_task(T3) -> 已取消 (进度 47%) ← 未过 50%,取消
|
||||
[22.89s] TASK | T2 完成 ✅
|
||||
[26.50s] AGENT | ## 分析汇总报告 … analyze_slow.py:已取消(未超 50%)…
|
||||
```
|
||||
|
||||
### Notes
|
||||
|
||||
- **Offline demos (`parallel`/`interrupt`/`state`) need no API key and no `openai` package.**
|
||||
- **Only `scenarios` needs network and a valid API key** (`OPENAI_API_KEY`, or `MOONSHOT_API_KEY` / `ARK_API_KEY`).
|
||||
- LLM wording varies per run; the four scenarios’ **behavioral logic** is stable. Retry on occasional high latency.
|
||||
- Timeline is accelerated; larger `FLUX_TICK_REAL` is closer to book “tens of seconds”; too small may break scenario 4’s under-50% cancel window.
|
||||
- Terminal jobs are real allowlisted Python child processes. Arbitrary commands and shell syntax are rejected before task allocation.
|
||||
|
||||
---
|
||||
|
||||
## 中文
|
||||
|
||||
本目录是《深入理解 AI Agent》实验 6-2 的配套可运行代码,实现了设计文档
|
||||
[`agent_framework_design.md`](./agent_framework_design.md) 中描述的事件驱动异步 Agent 框架(Flux)的核心部分。
|
||||
|
||||
在 4-5 的简单事件队列之上,本实验进入异步 Agent 的深水区,聚焦四件事:
|
||||
**异步工具执行、事件队列与批量处理、打断机制、并行工具的取消与状态查询**。
|
||||
Agent 需要同时管理多个并发任务,处理打断与恢复,并根据实时状态动态决策。
|
||||
|
||||
本目录提供两条使用路径:
|
||||
|
||||
- **离线演示(推荐先跑,零依赖、无需 API key)**:把三项核心异步能力单独拎出来、
|
||||
用可测量的方式演示——**并行 vs 串行的墙钟时间对比、打断/取消后恢复、状态检查点持久化与恢复**。
|
||||
这条路径不联网、不调用 LLM,甚至不需要安装 `openai`,`python demo.py` 即可直接运行。
|
||||
- **LLM 场景(还原书中四个验证场景)**:Agent 的决策由真实 LLM(默认 OpenAI `gpt-5.6-luna`,
|
||||
function calling)完成,需要配置 API key。
|
||||
|
||||
两条路径共用同一套异步运行时;长任务都用**模拟的异步"终端命令"**(带进度输出)实现,绝不真跑危险命令。
|
||||
|
||||
### 一、架构
|
||||
|
||||
对应设计文档第 5 节的事件处理循环,全部基于 `asyncio` 单线程实现:
|
||||
|
||||
```
|
||||
┌──────────────┐
|
||||
用户消息 / 打断 ──▶ │ inbox │ 所有进来的原始事件
|
||||
异步任务完成通知 ──▶ │ (asyncio.Q) │
|
||||
└──────┬───────┘
|
||||
│
|
||||
┌──────────▼───────────┐ 判定紧急度 classify_urgency()
|
||||
│ _dispatcher │──▶ 打断 / 立即处理 / 排队
|
||||
└──────────┬───────────┘
|
||||
┌────────────────┼───────────────────┐
|
||||
INTERRUPT │ IMMEDIATE│ DEFERRED│
|
||||
取消当前turn+异步工具 直接入 work 进 pending 缓冲,
|
||||
并留痕 异步结果到达时批量追加
|
||||
┌──────────▼───────────┐
|
||||
│ work │ 待处理的事件批次
|
||||
└──────────┬───────────┘
|
||||
┌──────────▼───────────┐
|
||||
│ _worker │ 逐批:追加到轨迹 -> run_llm_turn()
|
||||
│ turn_task 可被取消 │ (打断时 cancel 掉这个子任务)
|
||||
└──────────────────────┘
|
||||
|
||||
TaskManager:管理受限的真实子进程(start / query / cancel / cancel_all)
|
||||
任务自然完成 -> 以"新事件"(async.result) 注入 inbox
|
||||
```
|
||||
|
||||
代码文件:
|
||||
|
||||
| 文件 | 作用 |
|
||||
|------|------|
|
||||
| `events.py` | 事件模型 `Event`(含检查点序列化 `to_dict`/`from_dict`)、事件类型、**紧急度判定** `classify_urgency()` |
|
||||
| `tasks.py` | 真实受限子进程 `TaskManager`(stdout 进度、PID 取消/查询、可执行回执、`snapshot`/`restore`) |
|
||||
| `analysis_worker.py` | 真实分析进程:读取并哈希 `book/chapter4.md`,从 stdout 输出进度 |
|
||||
| `run_real_experiment.py` | 覆盖书中四场景的持久化验收运行器 |
|
||||
| `runtime.py` | `AgentRuntime`:事件循环、两种处理机制、LLM function calling、工具执行、检查点 `save_checkpoint`/`load_checkpoint` |
|
||||
| `async_demos.py` | 三个**离线演示**(无需 API key):并行墙钟对比、打断/恢复、状态检查点 |
|
||||
| `demo.py` | 统一命令行入口(argparse 子命令):离线演示 + 四个 LLM 验证场景 |
|
||||
|
||||
#### 两种事件处理机制(设计文档 5.1)
|
||||
|
||||
- **取消式处理(Cancellation-Based)**:紧急事件(用户"取消/停止")到达时,
|
||||
立即取消正在进行的 LLM turn,并取消所有后台异步工具,把打断事件与取消回执写入轨迹。
|
||||
- **排队处理(Queued)**:非紧急事件(补充性指令)先进入 `pending` 缓冲,不打断正在进行的工作;
|
||||
当某个异步工具完成、产生 `async.result` 事件时,一次性把 `pending` 里的事件批量追加到轨迹,再触发一次 LLM。
|
||||
|
||||
紧急度判定规则(简单可解释):
|
||||
|
||||
1. 含打断关键词(取消/停止/stop…)→ `INTERRUPT`(取消式处理)
|
||||
2. 是一个提问(带问号或疑问词,如"现在几点了?")→ `IMMEDIATE`(立即回应,但**不**打断后台任务)
|
||||
3. 其它补充性指令(如"用日语回复")→ `DEFERRED`(排队,批量处理)
|
||||
|
||||
#### 异步工具
|
||||
|
||||
`run_terminal_command` 是**异步**工具:调用后立刻返回 `task_id` 占位符(不阻塞),
|
||||
随后用 `asyncio.create_subprocess_exec`(`shell=False`)启动白名单子进程;进度来自真实 stdout。
|
||||
完成后把 PID、返回码、输入/输出哈希和文件分析指标作为**新事件**(`async.result`)注入对话;
|
||||
取消操作会终止对应 PID。
|
||||
另有 `query_task` / `cancel_task` 按 ID 查询进度与取消,`get_current_time` 用于即时提问。
|
||||
|
||||
**时间轴加速**:为便于复现,一个逻辑进度 tick 默认映射为 `0.4` 真实秒(`FLUX_TICK_REAL` 可调)。
|
||||
真实子进程保留 **3% / 2% / 1% 每 tick** 与 **是否过 50%** 的判定逻辑。
|
||||
|
||||
### 二、运行
|
||||
|
||||
命令行入口是 `demo.py`,用 `argparse` 子命令组织,`python demo.py --help` 查看全部用法。
|
||||
|
||||
#### 离线演示(无需 API key,开箱即用)
|
||||
|
||||
```bash
|
||||
cd chapter4/async-agent
|
||||
|
||||
python demo.py # 默认:依次运行下面三个离线演示
|
||||
python demo.py offline # 同上:显式地依次运行三个离线演示
|
||||
python demo.py parallel # 能力一:并行 vs 串行工具调用的墙钟时间对比(打印加速比)
|
||||
python demo.py interrupt # 能力二:长任务运行中被打断/取消,随后系统恢复
|
||||
python demo.py state # 能力三:状态检查点持久化 + 跨会话恢复并校验
|
||||
```
|
||||
|
||||
这三个演示不联网、不调用 LLM,连 `openai` 都无需安装——用纯 `asyncio` 直接测量并行加速、
|
||||
打断后的状态冻结、以及检查点的落盘与还原。
|
||||
|
||||
#### 测试(离线)
|
||||
|
||||
自动化回归测试位于 `tests/`,不需要 API Key。
|
||||
|
||||
```bash
|
||||
# 在仓库根目录安装 pytest 所需的 dev extra:
|
||||
uv sync --locked --python 3.12 --extra ch4 --extra dev
|
||||
|
||||
# pip 测试兜底路径:
|
||||
# python -m pip install -e ".[ch4,dev]"
|
||||
|
||||
cd chapter4/async-agent
|
||||
python -m pytest tests
|
||||
```
|
||||
|
||||
#### LLM 验证场景(还原书中四个场景,需要 API key)
|
||||
|
||||
```bash
|
||||
# 在仓库根目录使用统一的第 4 章环境
|
||||
uv sync --locked --python 3.12 --extra ch4
|
||||
|
||||
# 切换目录前先激活环境:
|
||||
# macOS/Linux:
|
||||
source .venv/bin/activate
|
||||
# Windows PowerShell:.venv\Scripts\Activate.ps1
|
||||
# Windows cmd:.venv\Scripts\activate.bat
|
||||
|
||||
# 未安装 uv 时可用 pip 兜底:
|
||||
# python -m pip install -e ".[ch4]"
|
||||
|
||||
cd chapter4/async-agent
|
||||
|
||||
# 迁移期间仍支持单项目兼容路径:
|
||||
# python -m pip install -r requirements.txt
|
||||
|
||||
cp env.example .env # 填入 OPENAI_API_KEY
|
||||
|
||||
python demo.py scenarios # 依次运行全部四个场景
|
||||
python demo.py scenarios --scenario 1 # 只跑场景 1(异步执行 + 即时提问)
|
||||
python demo.py scenarios --scenario 3 # 只跑场景 3(打断机制)
|
||||
```
|
||||
|
||||
默认用 OpenAI `gpt-5.6-luna`。也可切换服务商(OpenAI 兼容接口):
|
||||
|
||||
```bash
|
||||
# Moonshot(默认模型为当前的推理模型 kimi-k3)
|
||||
LLM_PROVIDER=moonshot python demo.py scenarios --scenario 1
|
||||
# 火山方舟 ARK(LLM_MODEL 填推理接入点 ID)
|
||||
LLM_PROVIDER=ark LLM_MODEL=ep-xxxx python demo.py scenarios --scenario 1
|
||||
```
|
||||
|
||||
> **OpenRouter 通用兜底**:未配置 `OPENAI_API_KEY`(且未用 moonshot/ark provider)时,
|
||||
> 只要设置了 `OPENROUTER_API_KEY`,`demo.py` 会自动改走 OpenRouter,并把模型名映射为
|
||||
> `provider/model` 形式(`gpt-*` → `openai/…`、`claude-*` → `anthropic/claude-opus-4.8`、
|
||||
> 含 `/` 的原样透传)。也可显式 `LLM_PROVIDER=openrouter`。例如:
|
||||
> `OPENROUTER_API_KEY=your-openrouter-api-key LLM_MODEL=openai/gpt-5.6-luna python demo.py scenarios --scenario 1`
|
||||
|
||||
> Moonshot 默认走**推理模型 `kimi-k3`**(旧的 `kimi-k2-*-preview` 与 `moonshot-v1-*` 已过时/停用)。
|
||||
> 推理模型要求 `temperature=1` 且 `max_tokens>=2048`,`demo.py` 会按模型自动套用这套采样参数,无需手动配置。
|
||||
|
||||
> 兼容旧用法:`python demo.py --scenario N` 会自动等价为 `scenarios --scenario N`。
|
||||
|
||||
日志中不同来源用颜色区分:`USER`(用户)、`AGENT`(Agent 回复)、`TOOL`(工具调用)、
|
||||
`TASK`(后台异步任务)、`TRAJ`(轨迹留痕)、`STATE`(状态检查点)、`SYSTEM`(框架事件)。
|
||||
|
||||
### 三、离线演示的三项能力(真实测量输出)
|
||||
|
||||
以下三段均为**真实运行**输出节选(无需 API key),演示异步到底带来了什么。
|
||||
|
||||
#### 能力一:并行 vs 串行工具调用(`python demo.py parallel`)
|
||||
|
||||
四个相互独立的只读感知工具(读文件 / 搜索 / 查库 / 向量检索),串行逐个 `await`
|
||||
与并行 `asyncio.gather` 的墙钟时间对比:
|
||||
|
||||
```
|
||||
── 结果对比 ─────────────────────────────────────────────
|
||||
串行总耗时(Σ 各工具) 4.51s
|
||||
并行总耗时(gather) 1.50s
|
||||
并行理论下界(最慢单个) 1.50s
|
||||
加速比 = 串行 / 并行 3.00x
|
||||
─────────────────────────────────────────────────────────
|
||||
```
|
||||
|
||||
墙钟时间由「各工具求和」降到「取最大单个」——这正是书中「只读感知工具天然适合并行」的量化落点。
|
||||
|
||||
#### 能力二:打断 / 取消 / 恢复(`python demo.py interrupt`)
|
||||
|
||||
三个并行后台任务运行中,用户先即时提问(不阻塞任务),随后发出「取消」打断:
|
||||
|
||||
```
|
||||
[ 1.00s] USER | (即时提问)现在几点了?
|
||||
[ 1.00s] AGENT | 现在 00:14:26。三个后台任务仍在并行推进,未被这次提问阻塞。
|
||||
[ 2.00s] USER | (打断)取消
|
||||
[ 2.00s] TASK | T1 已被取消 🛑(进度停在 39%)
|
||||
[ 2.00s] TASK | T2 已被取消 🛑(进度停在 26%)
|
||||
[ 2.00s] TASK | T3 已被取消 🛑(进度停在 13%)
|
||||
|
||||
── 打断后各任务状态(进度冻结在中途)───────────────────
|
||||
task_id 命令 状态 进度
|
||||
T1 python analyze_fast.py cancelled 39%
|
||||
T2 python analyze_mid.py cancelled 26%
|
||||
T3 python analyze_slow.py cancelled 13%
|
||||
─────────────────────────────────────────────────────────
|
||||
[ 2.05s] SYSTEM | 打断处理完毕,系统恢复空闲,可继续接受新任务……
|
||||
[ 5.52s] TASK | T4 完成 ✅
|
||||
[ 5.52s] AGENT | 已从打断中恢复,新任务 T4 正常完成:……
|
||||
```
|
||||
|
||||
打断只冻结被取消任务的进度,运行时本身无损,随后能立即接受并跑完新任务。
|
||||
|
||||
#### 能力三:状态检查点持久化与恢复(`python demo.py state`)
|
||||
|
||||
会话 A 产生一段轨迹 + 两个运行中的后台任务,落盘为 `checkpoints/agent_state.json`;
|
||||
会话 B 用全新运行时从磁盘恢复并校验:
|
||||
|
||||
```
|
||||
── 恢复校验 ─────────────────────────────────────────────
|
||||
轨迹事件数 保存前 3 -> 恢复后 3 [一致 ✓]
|
||||
可重建 LLM 上下文消息 4 条(system + 轨迹回放)
|
||||
task_id 命令 保存前进度 恢复后状态 进度
|
||||
T1 python analyze_fast.py 21% suspended 21%
|
||||
T2 python analyze_slow.py 7% suspended 7%
|
||||
─────────────────────────────────────────────────────────
|
||||
```
|
||||
|
||||
轨迹与任务进度完整落盘并跨会话还原;运行中的任务恢复后标记为 `suspended`,保留最后已知进度,
|
||||
供上层决定「重跑」还是「按进度续跑」——这就是异步任务的状态管理。
|
||||
|
||||
### 四、四个 LLM 验证场景
|
||||
|
||||
#### 场景 1:异步工具执行
|
||||
Agent 执行一个长终端命令,期间用户插入提问"现在几点了?"。
|
||||
因为长命令是异步的、不阻塞,Agent 立即用 `get_current_time` 回应时间,
|
||||
等后台任务完成后再把分析结论呈现出来。
|
||||
|
||||
#### 场景 2:事件队列与批量处理
|
||||
Agent 执行长任务期间,用户连续发"记得用日语回复""整理成网页"。
|
||||
这两条是非紧急指令,先进入排队缓冲;任务完成时,框架把它们**一次性批量追加**到轨迹,
|
||||
Agent 再综合所有指令,输出日语的 HTML 结果。
|
||||
|
||||
#### 场景 3:打断机制
|
||||
Agent 执行长任务,用户发"取消"。框架立即取消当前执行流并取消后台异步工具,
|
||||
在轨迹中记录打断事件(`user.interrupt`)和取消回执(`system.note`,含被取消的 task_id)。
|
||||
|
||||
#### 场景 4:并行工具的取消与状态查询
|
||||
用户要求"同时运行这三个脚本,哪个先完成就查其余进度,未过 50% 就取消"。
|
||||
三个脚本速度分别为 3% / 2% / 1% 每秒。Agent 同时启动三个异步任务;
|
||||
最快的先完成后,Agent 查询另外两个(约 66% 与 33%),取消未过 50% 的那个,
|
||||
其余完成后整合出报告。
|
||||
|
||||
### 五、LLM 场景真实运行输出(关键片段)
|
||||
|
||||
> 以下均为真实调用 `gpt-5.6-luna`(OpenAI 兼容接口)的输出节选(时间戳为真实秒,需配置 API key 复现)。
|
||||
|
||||
**场景 1(异步执行 + 即时提问)**
|
||||
```
|
||||
[ 3.97s] AGENT | 任务已在后台启动(task_id:T1)。完成后我会根据日志分析结果给出结论。
|
||||
[ 4.96s] TASK | T1 `python analyze_logs.py` 进度 22% ← 任务仍在后台跑
|
||||
[ 5.19s] TOOL | get_current_time -> 2026-07-18 13:43:30 ← 即时提问先回应
|
||||
[ 6.91s] AGENT | 现在是 2026 年 7 月 18 日 13:43:30。
|
||||
[12.19s] TRAJ | + async.result 异步完成 T1 ← 真实结果作为新事件注入
|
||||
[16.67s] AGENT | 日志分析已完成,结论如下:共扫描 12,840 条记录… ← 再呈现分析
|
||||
```
|
||||
|
||||
**场景 2(批量处理)**
|
||||
```
|
||||
[ 1.50s] SYSTEM | 事件进入排队缓冲(当前积压 1 条)
|
||||
[ 1.90s] SYSTEM | 事件进入排队缓冲(当前积压 2 条)
|
||||
[12.05s] TASK | T1 完成 ✅
|
||||
[12.05s] SYSTEM | 异步结果到达,批量处理 2 条积压的非紧急事件
|
||||
[12.06s] TRAJ | + async.result 异步完成 T1
|
||||
[12.06s] TRAJ | + user.input 记得最后用日语回复
|
||||
[12.06s] TRAJ | + user.input 把结果整理成一个网页(HTML)
|
||||
...
|
||||
[22.38s] AGENT | <!DOCTYPE html>…<h2>分析結論</h2>… (批量指令一次性满足:日语 + HTML)
|
||||
```
|
||||
|
||||
**场景 3(打断)**
|
||||
```
|
||||
[ 2.40s] TASK | 启动异步任务 T1: `python analyze_logs.py` (速度 4%/模拟秒)
|
||||
[ 4.00s] USER | (interrupt) 取消
|
||||
[ 4.00s] TASK | T1 已被取消 🛑(进度停在 14%)
|
||||
[ 4.00s] TRAJ | + user.interrupt 用户打断:取消
|
||||
[ 4.00s] TRAJ | + system.note 打断回执,取消任务 ['T1']
|
||||
[ 5.04s] AGENT | 已停止后台任务 T1。
|
||||
```
|
||||
|
||||
**场景 4(并行 + 状态查询 + 按 50% 阈值取消 + 整合报告)**
|
||||
```
|
||||
[ 2.82s] TASK | 启动异步任务 T1: `python analyze_fast.py` (速度 3%/模拟秒)
|
||||
[ 2.82s] TASK | 启动异步任务 T2: `python analyze_mid.py` (速度 2%/模拟秒)
|
||||
[ 2.82s] TASK | 启动异步任务 T3: `python analyze_slow.py` (速度 1%/模拟秒)
|
||||
[16.47s] TASK | T1 完成 ✅ ← 最快脚本先完成
|
||||
[19.84s] TOOL | query_task(T2) -> running 84% ← 查询其余两个进度
|
||||
[19.84s] TOOL | query_task(T3) -> running 42%
|
||||
[21.93s] TOOL | cancel_task(T3) -> 已取消 (进度 47%) ← 未过 50%,取消
|
||||
[22.89s] TASK | T2 完成 ✅
|
||||
[26.50s] AGENT | ## 分析汇总报告 … analyze_slow.py:已取消(未超 50%)…
|
||||
```
|
||||
|
||||
### 六、注意事项
|
||||
|
||||
- **离线演示(`parallel`/`interrupt`/`state`)无需任何 API key、也无需安装 `openai`**,开箱即跑。
|
||||
- **只有 `scenarios` 子命令需要联网并配置有效的 API key**(`OPENAI_API_KEY`,或切换到
|
||||
`MOONSHOT_API_KEY` / `ARK_API_KEY`)。
|
||||
- LLM 决策由真实模型产生,输出措辞每次可能略有不同;四个场景的**行为逻辑**是稳定可复现的。
|
||||
若遇到 OpenAI 偶发的高延迟,重跑即可。
|
||||
- 时间轴已加速;把 `FLUX_TICK_REAL` 调大可让演示更接近书中"几十秒"的真实节奏,
|
||||
调小则更快(过小可能让场景 4 的"未过 50% 就取消"来不及判定)。
|
||||
- 终端任务是真实但受限的 Python 子进程;任意命令与 shell 语法会在分配任务 ID 前被拒绝。
|
||||
|
||||
---
|
||||
|
||||
## Notes / 说明
|
||||
|
||||
- Design details: [`agent_framework_design.md`](./agent_framework_design.md).
|
||||
- 设计细节见 [`agent_framework_design.md`](./agent_framework_design.md)。
|
||||
- Terminal jobs are real allowlisted child processes and never invoke a shell.
|
||||
- 终端任务是白名单真实子进程,且绝不调用 shell。
|
||||
@@ -0,0 +1,159 @@
|
||||
## Flux: An Event-Driven Framework for Asynchronous Agentic Workflows
|
||||
|
||||
**Abstract:** Flux is a software framework designed for building asynchronous and event-driven AI agents, with a strong emphasis on **enabling low-code development**. Drawing an analogy to a real human, Flux treats agents as entities whose state and understanding evolve based on their accumulating memory. Flux aims to empower users, **including those with limited programming experience**, to define complex agentic workflows primarily through **declarative configuration**, minimizing the need for writing extensive code. It supports persistent, per-user long-term memory, modular agent invocation, and seamless tool integration. Communication adheres to the Agent Messaging Protocol (AMP). Key features include asynchronous execution, optional synchronous tools, interrupt handling, streaming outputs, and a developer experience focused on ease of use and configuration over complex coding.
|
||||
|
||||
### 1. Introduction
|
||||
|
||||
Modern AI applications require agents capable of complex, stateful, and collaborative interactions. Flux addresses this need with an asynchronous, event-driven architecture inspired by human cognition and OS principles. Crucially, Flux is designed to **democratize agent development**. Instead of requiring deep programming expertise, it focuses on allowing developers to define agent behavior, logic, and workflows through **intuitive configuration files and well-defined prompts**. The goal is to provide a robust platform where the core complexity of asynchronous processing, state management, and communication is handled by the framework, freeing developers to concentrate on the agent's specific goals and capabilities.
|
||||
|
||||
Flux models an agent like a human, processing inputs, thinking internally, acting externally, and handling interruptions, all recorded in its memory. It supports both rollout-specific working memory (trajectory) and user-specific long-term memory. By leveraging LLMs for decision-making based on this memory and providing a **configuration-driven approach** to defining agents and tools, Flux facilitates the creation of sophisticated agents without demanding extensive coding skills, making agentic AI development more accessible.
|
||||
|
||||
### 2. Core Concepts
|
||||
|
||||
* **Agent:** An agent is an autonomous computational entity, analogous to a real human expert or assistant. It possesses a defined set of capabilities (tools/skills), operates based on internal logic (primarily LLM-driven decision-making informed by its memory), and interacts with its environment through events within the scope of an **AMP Session**. An agent's understanding and state within a specific rollout are derived from its accumulated **trajectory**. It can also access and modify **User Long-Term Memory** to inform its behavior based on past interactions with a specific user. Each agent definition serves as a blueprint.
|
||||
* **Rollout:** A rollout represents a specific, running instance of an agent executing its defined logic within the context of an **AMP Session**. It is typically initiated by an event within that session (e.g., a user message, an agent invocation). Each rollout maintains its own independent **trajectory** and lifecycle state. Importantly, a single AMP Session can contain multiple participants (users and agents), and thus may involve multiple concurrent Flux Rollouts (one for each active agent participant in the session).
|
||||
* **Event:** The fundamental unit representing any occurrence relevant to a rollout, forming the building blocks of the agent's **trajectory**. All interactions, internal processing steps, and external stimuli within the rollout's scope are captured as events, appended chronologically to the rollout's trajectory. Drawing inspiration from OS process states and signals, events are categorized as Inputs, Interrupts, Thinking, and Actions:
|
||||
* **Inputs:** Events originating externally to the agent's rollout, akin to sensory perception within the session.
|
||||
* `user.input`: Messages or actions from a specific human user *within the current AMP session*.
|
||||
* `agent.input`: Messages or results received from another invoked agent rollout *within the same AMP session*.
|
||||
* `tool.result`: Data returned from a completed tool execution (sync or async), including results from memory access tools.
|
||||
* `external.trigger`: Events from outside systems integrated with the framework (e.g., API webhook, database change, new social media post), potentially associated with the session or a user.
|
||||
* **Interrupts:** Events that signal a need to alter or halt the current flow of execution within the rollout, often requiring immediate attention.
|
||||
* `supervisor.instruction`: Commands from an administrative system or human supervisor pertaining to this rollout or session.
|
||||
* `user.interrupt`: User actions (from a specific user in the session) like clicking a 'stop' button or explicitly cancelling an operation related to this rollout.
|
||||
* `timer.trigger`: An event fired by a previously set timer associated with this rollout.
|
||||
* **Thinking:** Events representing the agent's internal cognitive processes or state changes within the rollout.
|
||||
* `agent.thought`: Internal reasoning steps, intermediate conclusions, or state changes logged by the agent for transparency or future context within its trajectory.
|
||||
* **Actions:** Events representing the agent's decisions to interact with or affect the external world (including other participants in the session or external systems) or schedule future events for this rollout.
|
||||
* `agent.output`: Messages or data prepared to be sent to a specific user (or all users) *within the current AMP session* (mediated via a tool call adhering to AMP).
|
||||
* `agent.escalation`: A request sent to a supervisor system regarding this rollout or session.
|
||||
* `tool.request`: An invocation request for an external tool, which might include tools for accessing/updating **User Long-Term Memory**.
|
||||
* `agent.invocation`: A request to create and start a new rollout for another agent *within the same AMP session*, initiating collaboration.
|
||||
* `agent.response`: A message sent back to the agent rollout that invoked the current one *within the same AMP session*.
|
||||
* `agent.interrupt`: A signal sent to interrupt another specific agent rollout *within the same AMP session*.
|
||||
* `timer.set`: An instruction to the framework to schedule a `timer.trigger` event for the future for this specific rollout.
|
||||
* **Trajectory:** The complete, time-ordered, immutable sequence of all events (Inputs, Interrupts, Thinking, and Actions) associated with a *specific rollout*. This **is** the agent's working memory for that rollout. It provides the primary context required by the LLM to understand the rollout's immediate history, its own past reasoning and actions *within that rollout*, and make informed decisions about the next thinking steps or actions for that rollout.
|
||||
* **User Long-Term Memory:** A persistent, key-value store associated with a unique user ID. This memory exists *across* different rollouts (AMP Sessions) involving that user. It's designed to hold relatively stable information like user preferences, summarized past interactions, contact details, or accumulated knowledge about the user relevant for personalization and continuity. Agents access and modify this memory explicitly via designated **Actions** (e.g., specific `tool.request` events). This is distinct from the rollout-specific trajectory.
|
||||
* **Rollout Business State:** (Optional) A developer-defined state representing the current logical phase or status of the rollout from the perspective of the agent's business logic (e.g., `needs_clarification`, `processing_request`, `waiting_for_payment`, `request_completed`). This is distinct from the framework's internal rollout lifecycle state (`running`, `waiting`) and complements the trajectory by providing a high-level summary of the rollout's progress according to the developer's intended workflow.
|
||||
* **Workflow:** A definition specifying the entry point agent and potential interactions. In Flux, workflows can be:
|
||||
* **Emergent (Low-Code Default):** Driven primarily by the LLM's decisions based on memory and available tools. This approach typically requires **minimal explicit workflow configuration**, relying on the LLM's reasoning capabilities guided by prompts.
|
||||
* **State-Guided (Optional/Advanced):** Influenced by developer-defined **Rollout Business States**. This offers more explicit control for complex scenarios but requires more configuration.
|
||||
|
||||
### 3. Architecture
|
||||
|
||||
Flux employs a modular architecture centered around asynchronous event processing within the context of AMP Sessions:
|
||||
|
||||
* **Rollout Manager:** Responsible for creating, tracking, and terminating agent rollouts, associating them with their corresponding AMP Session and participant ID. It assigns unique IDs to rollouts and manages their lifecycle state.
|
||||
* **Event Queue (per Rollout):** Each active rollout has an associated event queue where incoming events (Inputs, Interrupts) relevant to that agent's role in the session are placed.
|
||||
* **Agent Runtime:** The core execution engine for a single rollout. It comprises:
|
||||
* **Event Processor:** Dequeues events, appends them to the rollout's trajectory, and determines if an LLM invocation is needed.
|
||||
* **LLM Invoker:**
|
||||
* Formats the rollout's current trajectory into the structure expected by the configured LLM. Optionally includes the current **Rollout Business State**. May also optionally include relevant excerpts from **User Long-Term Memory** (retrieved via a previous Action or potentially through automatic framework injection based on configuration).
|
||||
* Constructs the final prompt using the agent's system prompt, the formatted trajectory, and the user prompt template.
|
||||
* Invokes the configured LLM API.
|
||||
* Parses the LLM's response (Thinking/Action events), handling streaming for incremental processing.
|
||||
* **Tool Executor:**
|
||||
* Receives Action events (`tool.request`, `agent.invocation`, `agent.output`, `update_rollout_state`, etc.).
|
||||
* Manages the execution of these actions, interacting with external tools, other agents within the session, the Communication Layer, potentially a **Long-Term Memory Service**, and updating the **Rollout Business State** if requested.
|
||||
* Handles async/sync execution logic and cancellation.
|
||||
* Generates corresponding Input events (`tool.result`, `agent.input`, `timer.trigger`) and places them back into the appropriate rollout's Event Queue.
|
||||
* **Communication Layer:** Handles external communication via the AMP specification for the session. Translates incoming AMP messages/requests (addressed to the agent this rollout represents) into Flux Input/Interrupt events for the rollout. Translates agent Action events (`agent.output`) into outgoing AMP messages/streams targeted at the correct participants within the session. Manages SSE connections.
|
||||
|
||||
### 4. Agent Definition
|
||||
|
||||
Agents are defined declaratively, primarily through **configuration files (e.g., YAML, JSON)**, aligning with the low-code philosophy. An agent definition typically includes:
|
||||
|
||||
* **Identifier:** Unique name/ID.
|
||||
* **System Prompt:** Defines persona, goals, constraints. **A key area for defining agent logic without code.**
|
||||
* **User Prompt Template:** Structures the prompt, potentially including placeholders for Rollout Business State. **Another key configuration point.**
|
||||
* **Model Configuration:** LLM choice and parameters (simple configuration).
|
||||
* **Tool Registry:** Lists available tools/agents. Referencing existing tools/agents is a simple configuration entry. Defining *new* tools may require code, but the framework provides clear interfaces and registration mechanisms to simplify this.
|
||||
* Standard external tools.
|
||||
* Memory tools (access/update can be explicit tool calls or potentially configured implicit actions, offering flexibility).
|
||||
* References to other registered Flux agents.
|
||||
* (Optional) Action for updating Rollout Business State (`update_rollout_state`).
|
||||
* **State Machine Definition (Advanced/Optional):** For agents requiring very specific, complex state management, developers *can* define explicit states and transitions. **This is not required for typical agents**, where state can be managed implicitly through the trajectory or LLM reasoning.
|
||||
* **Context Inheritance Policy (Advanced/Optional):** Configuration for how much context is passed to invoked agents. **Simple defaults are provided**, and explicit configuration is only needed for specialized use cases.
|
||||
* **Workflow Specification (Optional):** Can define entry points. More complex workflow logic is often better embedded within the system prompt or handled via the optional state mechanism, rather than requiring complex external workflow definitions.
|
||||
|
||||
**Core agent logic often resides within the prompts and the LLM's inherent capabilities, configured declaratively, rather than in complex code within the framework.**
|
||||
|
||||
### 5. Event Processing and LLM Interaction
|
||||
|
||||
The core loop for an active rollout:
|
||||
|
||||
1. **Event Dequeue:** Get the next event for this rollout.
|
||||
2. **Memory Update:** Append the event to the trajectory.
|
||||
3. **LLM Trigger Check:** Decide if LLM processing is needed.
|
||||
4. **Context Formatting:** Translate trajectory into LLM format. Include the current **Rollout Business State** if defined. Optionally include retrieved **User Long-Term Memory** data (via explicit tool result or automatic injection).
|
||||
5. **LLM Invocation:** Send context and prompts to the LLM.
|
||||
6. **Response Parsing & Streaming:** Parse LLM response into Thinking/Action events.
|
||||
7. **Thinking/Action Generation & Internal State Update:** Add generated Thinking and Action events to the trajectory. If the LLM generated an `update_rollout_state` action, the Agent Runtime immediately processes it here, updating the rollout's current business state field.
|
||||
8. **Dispatch Actions to Executor:** Send all other generated Action events (those requiring interaction with external tools, other agents, the communication layer, timers, etc. – e.g., `tool.request`, `agent.invocation`, `agent.output`, `timer.set`) to the Tool Executor for handling.
|
||||
9. **Loop/Wait:** Wait for the next event.
|
||||
|
||||
#### 5.1 Event Processing Mechanisms
|
||||
|
||||
Flux supports two dynamic event processing strategies for handling events in a rollout. The framework automatically selects the appropriate mechanism based on the urgency of the incoming event:
|
||||
|
||||
1. **Cancellation-Based Processing:** When an urgent event arrives (e.g., user interrupts, high-priority inputs), the framework immediately stops the current LLM thinking or any synchronous tool call. All queued events in the pending queue along with the new urgent event are immediately appended to the trajectory. The LLM is then invoked with the complete updated trajectory to process all accumulated events together. This approach ensures that urgent, potentially invalidating events receive immediate attention and the agent can make decisions with the most critical, up-to-date information.
|
||||
|
||||
2. **Queued Processing:** When a non-urgent event arrives, it is queued at the end of the pending queue without interrupting ongoing processing. When any tool call (synchronous or asynchronous) of the agent completes and returns a `tool.result` event, the framework checks the pending queue before invoking the LLM. If there are pending events, all events in the pending queue are immediately moved to the end of the trajectory, and then the LLM is invoked to process the updated trajectory. This approach allows the agent to complete ongoing operations while efficiently batching non-urgent events, balancing responsiveness with computational efficiency.
|
||||
|
||||
**Urgency Determination:** The framework classifies events based on their type and context:
|
||||
- **Urgent events:** User interrupts (`user.interrupt`), supervisor instructions (`supervisor.instruction`), explicit agent interrupts (`agent.interrupt`), and time-critical external triggers marked as urgent.
|
||||
- **Non-urgent events:** Regular user inputs (`user.input`), agent messages (`agent.input`), tool results (`tool.result`), timer triggers (`timer.trigger`), and standard external triggers.
|
||||
|
||||
This dynamic selection ensures optimal responsiveness for critical events while maintaining efficiency for routine operations.
|
||||
|
||||
### 6. Tool Execution
|
||||
|
||||
Tools are fundamental.
|
||||
|
||||
* **Definition:** Tools have a clear definition structure (name, description, parameters). Defining *new* tools involves implementing a defined interface, but *using* existing tools is purely configuration.
|
||||
* **Invocation:** Triggered by LLM via `tool.request` action based on prompt and available tools.
|
||||
* **Memory Tools:** Framework aims to provide flexible options (explicit call vs. implicit action) configurable by the developer.
|
||||
* **Execution:** Asynchronous/synchronous handling is managed by the framework, abstracted from the developer defining the agent.
|
||||
* **Results:** Fed back as events, handled by the framework.
|
||||
* **Cancellation:** Framework provides the mechanism.
|
||||
|
||||
### 7. Inter-Agent Communication (Actor Model within AMP Session)
|
||||
|
||||
Flux implements actor model principles constrained within an AMP Session:
|
||||
|
||||
* **Invocation:** `Agent A` (Rollout A) invokes `Agent B` by generating an `agent.invocation` action. The Rollout Manager creates Rollout B for Agent B *within the same AMP Session*. Rollout B's initial event is the invocation details from Rollout A.
|
||||
* **Communication:** Rollout B can send results/messages back to Rollout A using `agent.response` actions, which arrive as `agent.input` events at Rollout A. Rollout B can also send messages directly to users in the session via `agent.output` actions, handled by the Communication Layer. All communication is asynchronous and potentially streaming via AMP.
|
||||
* **Context Sharing:** Simple defaults are provided. Explicit configuration of context sharing is an advanced option.
|
||||
|
||||
### 8. State Management
|
||||
|
||||
Flux manages state layers, abstracting complexity:
|
||||
|
||||
* **Framework Rollout State:** Internal framework state.
|
||||
* **Trajectory:** Automatically maintained log.
|
||||
* **User Long-Term Memory:** Accessed via configured tools or actions.
|
||||
* **Developer-Defined Rollout Business State (Optional):** A straightforward key-value state field developers can optionally use and manage via configuration and LLM actions for more explicit control when needed.
|
||||
* **LLM Context:** Assembled automatically by the framework based on configuration and state.
|
||||
|
||||
### 9. Streaming and Communication Protocol
|
||||
|
||||
* **AMP Adherence:** Handled by the framework's Communication Layer.
|
||||
* **Input/Output Mapping:** Handled by the framework.
|
||||
* **Streaming LLM Output:** Handled by the framework.
|
||||
|
||||
### 10. Developer Experience
|
||||
|
||||
Flux is fundamentally designed for **ease of use and low-code agent development**:
|
||||
|
||||
* **Declarative First:** The primary way to define agents, their logic (via prompts), tool usage, and basic workflows is through **declarative configuration files** (e.g., YAML, JSON), not procedural code.
|
||||
* **Abstraction:** The complexities of asynchronous execution, event loops, state persistence, AMP communication, and streaming are **handled internally by the framework**, allowing developers to focus on *what* the agent should do, not *how* the underlying machinery works.
|
||||
* **Configuration over Code:** Core agent behavior, personality, and decision-making logic are primarily defined in system prompts and by selecting available tools in the configuration.
|
||||
* **Simplified Tooling:** While new tools require some code, the framework provides clear interfaces. A library of pre-built common tools (including memory access) further reduces the need for coding.
|
||||
* **Emergent Workflows:** Simple agents can often function effectively by letting the LLM decide the next steps based on the trajectory and available tools, requiring minimal explicit workflow definition.
|
||||
* **Optional Complexity:** Features like explicit Rollout Business States and detailed Context Inheritance policies are available for advanced users needing fine-grained control, but are **not required** for basic agent development.
|
||||
* **Observability:** Clear logging and tracing mechanisms are provided to understand agent behavior without needing to debug complex framework internals.
|
||||
* **(Future) Visual Tools:** The declarative nature of Flux lends itself well to potential future development of GUI or visual flowcharting tools for defining agents and workflows, further enhancing accessibility.
|
||||
|
||||
### 11. Conclusion
|
||||
|
||||
Flux offers a robust, event-driven architecture designed for building sophisticated asynchronous AI agents while prioritizing **low-code development and accessibility**. By abstracting framework complexities and enabling agent definition primarily through **declarative configuration and prompting**, Flux empowers a wider range of developers, including those with limited coding experience, to create powerful agentic solutions. Its support for multiple memory types, modularity, adherence to AMP, and focus on a simplified developer experience makes it a strong foundation for building the next generation of collaborative and stateful AI applications.
|
||||
@@ -0,0 +1,54 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Bounded analysis executable used by the Experiment 6-2 task manager."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def analyze(path: Path, job: str) -> dict:
|
||||
data = path.read_bytes()
|
||||
text = data.decode("utf-8", errors="replace")
|
||||
lines = text.splitlines()
|
||||
headings = [line for line in lines if re.match(r"^#{1,6}\s", line)]
|
||||
return {
|
||||
"job": job,
|
||||
"input_path": str(path),
|
||||
"input_sha256": hashlib.sha256(data).hexdigest(),
|
||||
"bytes": len(data),
|
||||
"lines": len(lines),
|
||||
"heading_count": len(headings),
|
||||
"experiment_mentions": text.lower().count("实验"),
|
||||
"async_mentions": len(re.findall(r"异步|async", text, flags=re.IGNORECASE)),
|
||||
"error_keyword_count": len(re.findall(r"错误|error", text, flags=re.IGNORECASE)),
|
||||
}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--job", required=True,
|
||||
choices=["fast", "mid", "slow", "logs", "recovery"])
|
||||
parser.add_argument("--rate", required=True, type=float)
|
||||
parser.add_argument("--tick-real", required=True, type=float)
|
||||
parser.add_argument("--input", required=True, type=Path)
|
||||
args = parser.parse_args()
|
||||
if args.rate <= 0 or args.tick_real <= 0:
|
||||
raise SystemExit("rate and tick-real must be positive")
|
||||
if not args.input.is_file():
|
||||
raise SystemExit(f"input does not exist: {args.input}")
|
||||
progress = 0.0
|
||||
while progress < 100.0:
|
||||
time.sleep(args.tick_real)
|
||||
progress = min(100.0, progress + args.rate)
|
||||
print(f"PROGRESS {progress:.3f}", flush=True)
|
||||
print("RESULT " + json.dumps(analyze(args.input.resolve(), args.job),
|
||||
ensure_ascii=False, sort_keys=True), flush=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,220 @@
|
||||
"""离线演示:不依赖任何 LLM / API key,直接驱动异步运行时的底层原语。
|
||||
|
||||
`demo.py` 里的四个「场景」需要真实 LLM 做决策;本模块则把实验 6-2 的三项核心
|
||||
异步能力单独拎出来,用可测量、可复现的方式演示,**无需联网、无需 API key**:
|
||||
|
||||
- demo_parallel :并行 vs 串行工具调用的【墙钟时间】对比(真实测量,打印加速比)。
|
||||
- demo_interrupt :长任务运行中被【打断/取消】,随后系统【恢复】并接受新任务。
|
||||
- demo_state :Agent 状态【检查点持久化】到磁盘,再【跨会话恢复】并校验。
|
||||
|
||||
这三个演示共同回答「异步到底带来了什么」——用数字和状态变化说话,而不只是措辞。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import datetime
|
||||
import os
|
||||
import time
|
||||
|
||||
from runtime import AgentRuntime, format_log
|
||||
from events import Event, EventType
|
||||
from tasks import TaskManager
|
||||
import tasks
|
||||
|
||||
|
||||
class Logger:
|
||||
"""与 runtime 同款的彩色时间戳日志器(相对本次演示起点计时)。"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.t0 = time.time()
|
||||
|
||||
def __call__(self, source: str, text: str) -> None:
|
||||
print(format_log(self.t0, source, text), flush=True)
|
||||
|
||||
|
||||
def banner(title: str) -> None:
|
||||
print("\n" + "=" * 78)
|
||||
print(f" {title}")
|
||||
print("=" * 78, flush=True)
|
||||
|
||||
|
||||
# ============================ 1. 并行 vs 串行 ============================
|
||||
|
||||
# 一组相互独立的【只读感知工具】(读文件 / 搜索 / 查库 / 向量检索)。
|
||||
# 只读、无副作用,因此可以安全地并行——这正是书中「感知工具天然适合并行」的落点。
|
||||
_PERCEIVE_TOOLS = [
|
||||
("read_config.json", 0.8),
|
||||
("web_search(‘异步 Agent’)", 1.2),
|
||||
("db_query(orders)", 1.5),
|
||||
("vector_lookup(memory)", 1.0),
|
||||
]
|
||||
|
||||
|
||||
async def _perceive(name: str, latency: float, log: Logger) -> tuple[str, float, float]:
|
||||
"""模拟一次带 I/O 延迟的只读感知调用;返回 (名称, 标称延迟, 实测耗时)。"""
|
||||
t0 = time.time()
|
||||
log("TOOL", f"→ {name} 启动(模拟 I/O 耗时 {latency:.1f}s)")
|
||||
await asyncio.sleep(latency)
|
||||
dt = time.time() - t0
|
||||
log("TOOL", f"✓ {name} 完成(实测 {dt:.2f}s)")
|
||||
return name, latency, dt
|
||||
|
||||
|
||||
async def demo_parallel() -> None:
|
||||
banner("能力一|并行工具调用:并行 vs 串行的墙钟时间对比")
|
||||
log = Logger()
|
||||
log("SYSTEM", "有 4 个相互独立的只读感知工具需要调用(无副作用,可安全并行)。")
|
||||
|
||||
# —— 串行:一个 await 完再 await 下一个 ——
|
||||
log("SYSTEM", "\033[0m[串行] 逐个 await(同步 ReAct 的默认做法)……")
|
||||
seq_start = time.time()
|
||||
for name, lat in _PERCEIVE_TOOLS:
|
||||
await _perceive(name, lat, log)
|
||||
seq_total = time.time() - seq_start
|
||||
|
||||
# —— 并行:一次性发起,asyncio.gather 并发等待 ——
|
||||
log("SYSTEM", "\033[0m[并行] 一次性发起,asyncio.gather 并发等待……")
|
||||
par_start = time.time()
|
||||
await asyncio.gather(*[_perceive(name, lat, log) for name, lat in _PERCEIVE_TOOLS])
|
||||
par_total = time.time() - par_start
|
||||
|
||||
slowest = max(lat for _, lat in _PERCEIVE_TOOLS)
|
||||
speedup = seq_total / par_total if par_total else float("inf")
|
||||
|
||||
print("\n ── 结果对比 ─────────────────────────────────────────────")
|
||||
print(f" {'工具':<26}{'标称延迟':>10}")
|
||||
for name, lat in _PERCEIVE_TOOLS:
|
||||
print(f" {name:<26}{lat:>8.1f}s")
|
||||
print(" ─────────────────────────────────────────────────────────")
|
||||
print(f" {'串行总耗时(Σ 各工具)':<26}{seq_total:>8.2f}s")
|
||||
print(f" {'并行总耗时(gather)':<26}{par_total:>8.2f}s")
|
||||
print(f" {'并行理论下界(最慢单个)':<26}{slowest:>8.2f}s")
|
||||
print(f" {'加速比 = 串行 / 并行':<26}{speedup:>8.2f}x")
|
||||
print(" ─────────────────────────────────────────────────────────")
|
||||
print(" 结论:独立的只读调用并行化后,墙钟时间由「求和」降到「取最大」。\n")
|
||||
|
||||
|
||||
# ============================ 2. 打断 / 取消 / 恢复 ============================
|
||||
|
||||
async def demo_interrupt() -> None:
|
||||
banner("能力二|打断与取消:长任务运行中被打断,随后系统恢复")
|
||||
tasks.TICK_REAL = 0.15 # 本演示放慢节奏,留出「跑到一半再打断」的时间窗口
|
||||
log = Logger()
|
||||
completed: list = []
|
||||
|
||||
async def on_complete(state) -> None:
|
||||
completed.append(state)
|
||||
|
||||
tm = TaskManager(on_complete=on_complete, log=log)
|
||||
|
||||
# 1) 并行启动三个后台异步任务
|
||||
log("SYSTEM", "启动三个并行后台分析任务(fast/mid/slow)……")
|
||||
for cmd in ["python analyze_fast.py", "python analyze_mid.py", "python analyze_slow.py"]:
|
||||
tm.start(cmd)
|
||||
|
||||
# 2) 运行期间用户即时提问 —— 后台任务不被阻塞
|
||||
await asyncio.sleep(1.0)
|
||||
now = datetime.datetime.now().strftime("%H:%M:%S")
|
||||
log("USER", "(即时提问)现在几点了?")
|
||||
log("AGENT", f"现在 {now}。三个后台任务仍在并行推进,未被这次提问阻塞。")
|
||||
|
||||
# 3) 跑到中途,用户发出打断 —— 立即取消所有在跑的任务
|
||||
await asyncio.sleep(1.0)
|
||||
log("USER", "(打断)取消")
|
||||
cancelled = tm.cancel_all()
|
||||
await asyncio.sleep(0.05) # 让 CancelledError 在各协程内落地
|
||||
log("SYSTEM", f"已执行打断:取消了 {cancelled}(进度在被取消处冻结)")
|
||||
|
||||
print("\n ── 打断后各任务状态(进度冻结在中途)───────────────────")
|
||||
print(f" {'task_id':<8}{'命令':<26}{'状态':<12}{'进度':>6}")
|
||||
for s in tm.all_states():
|
||||
print(f" {s.task_id:<8}{s.command:<26}{s.status:<12}{s.progress:>5.0f}%")
|
||||
print(" ─────────────────────────────────────────────────────────")
|
||||
|
||||
# 4) 恢复:executor 依然健康,接受并跑完一个新任务
|
||||
log("SYSTEM", "打断处理完毕,系统恢复空闲,可继续接受新任务……")
|
||||
fresh = tm.start("python re_run_summary.py")
|
||||
await fresh._task
|
||||
log("AGENT", f"已从打断中恢复,新任务 {fresh.task_id} 正常完成:"
|
||||
f"{completed[-1].result[:36]}……")
|
||||
print(" 结论:打断只冻结被取消的任务,运行时本身无损,可立即继续工作。\n")
|
||||
|
||||
|
||||
# ============================ 3. 状态检查点:持久化 / 恢复 ============================
|
||||
|
||||
def _seed_trajectory(rt: AgentRuntime) -> None:
|
||||
"""给运行时灌入一段「已发生」的对话轨迹,模拟会话进行到一半。"""
|
||||
rt._append(Event(EventType.USER_INPUT,
|
||||
message={"role": "user", "content": "分析今天的日志并总结异常"},
|
||||
label="用户消息:分析日志"))
|
||||
rt._append(Event(EventType.AGENT_TOOL_CALL,
|
||||
message={"role": "assistant", "content": "好的,我这就在后台启动分析。",
|
||||
"tool_calls": [{"id": "call_1", "type": "function",
|
||||
"function": {"name": "run_terminal_command",
|
||||
"arguments": '{"command": "python analyze_fast.py"}'}}]},
|
||||
label="调用工具 run_terminal_command"))
|
||||
rt._append(Event(EventType.TOOL_RESULT,
|
||||
message={"role": "tool", "tool_call_id": "call_1",
|
||||
"content": "命令已在后台异步启动。task_id=T1。"},
|
||||
label="工具结果 run_terminal_command"))
|
||||
|
||||
|
||||
async def demo_state() -> None:
|
||||
banner("能力三|状态管理:检查点持久化与跨会话恢复")
|
||||
tasks.TICK_REAL = 0.15
|
||||
ckpt_dir = os.path.join(os.path.dirname(os.path.abspath(__file__)), "checkpoints")
|
||||
os.makedirs(ckpt_dir, exist_ok=True)
|
||||
path = os.path.join(ckpt_dir, "agent_state.json")
|
||||
|
||||
# —— 会话 A:产生一段轨迹 + 两个仍在运行的后台任务,然后落盘 ——
|
||||
log = Logger()
|
||||
log("SYSTEM", "会话 A 开始:构造轨迹并启动两个后台任务……")
|
||||
rt_a = AgentRuntime(client=None, model="demo-offline")
|
||||
rt_a._t0 = log.t0 # 让两个运行时共用同一时间基准,便于观察
|
||||
_seed_trajectory(rt_a)
|
||||
rt_a.tasks.start("python analyze_fast.py") # 进行中
|
||||
rt_a.tasks.start("python analyze_slow.py") # 进行中
|
||||
await asyncio.sleep(1.2) # 让进度累积到中途
|
||||
|
||||
before_traj = len(rt_a.trajectory)
|
||||
before_tasks = {s.task_id: (s.status, s.progress) for s in rt_a.tasks.all_states()}
|
||||
rt_a.save_checkpoint(path)
|
||||
|
||||
# 模拟进程退出:取消掉活着的协程
|
||||
rt_a.tasks.cancel_all()
|
||||
await asyncio.sleep(0.05)
|
||||
log("SYSTEM", "会话 A 结束(进程退出,内存中的运行时已销毁)。")
|
||||
|
||||
# —— 会话 B:全新运行时,从磁盘恢复 ——
|
||||
log("SYSTEM", "会话 B 开始:新建空运行时,从检查点恢复……")
|
||||
rt_b = AgentRuntime(client=None, model="demo-offline")
|
||||
rt_b._t0 = log.t0
|
||||
data = rt_b.load_checkpoint(path)
|
||||
|
||||
after_traj = len(rt_b.trajectory)
|
||||
msgs = rt_b.build_messages() # 证明恢复后能重建可喂给 LLM 的上下文
|
||||
|
||||
print("\n ── 恢复校验 ─────────────────────────────────────────────")
|
||||
print(f" 轨迹事件数 保存前 {before_traj} -> 恢复后 {after_traj} "
|
||||
f"[{'一致 ✓' if before_traj == after_traj else '不一致 ✗'}]")
|
||||
print(f" 可重建 LLM 上下文消息 {len(msgs)} 条(system + 轨迹回放)")
|
||||
print(f" {'task_id':<8}{'命令':<26}{'保存前进度':>10} {'恢复后状态':<12}{'进度':>6}")
|
||||
for rec in data["tasks"]:
|
||||
tid = rec["task_id"]
|
||||
before = before_tasks.get(tid, ("-", 0.0))
|
||||
st = rt_b.tasks.query(tid)
|
||||
print(f" {tid:<8}{rec['command']:<26}{before[1]:>9.0f}% "
|
||||
f"{st.status:<12}{st.progress:>5.0f}%")
|
||||
print(" ─────────────────────────────────────────────────────────")
|
||||
print(f" 检查点文件:{path}")
|
||||
print(" 结论:轨迹与任务进度完整落盘并跨会话还原;运行中的任务被标记为 suspended,")
|
||||
print(" 保留了最后已知进度,供上层决定「重跑」还是「按进度续跑」。\n")
|
||||
|
||||
|
||||
# 供 demo.py 复用的离线演示注册表
|
||||
OFFLINE_DEMOS = {
|
||||
"parallel": demo_parallel,
|
||||
"interrupt": demo_interrupt,
|
||||
"state": demo_state,
|
||||
}
|
||||
@@ -0,0 +1,271 @@
|
||||
"""实验 6-2 命令行入口:带并行执行、打断/取消与状态管理的异步 Agent。
|
||||
|
||||
本脚本提供两类演示,用子命令区分:
|
||||
|
||||
【离线演示】不需要任何 API key,直接测量异步运行时的底层行为——
|
||||
python demo.py parallel 并行 vs 串行工具调用的墙钟时间对比(打印加速比)
|
||||
python demo.py interrupt 长任务运行中被打断/取消,随后系统恢复
|
||||
python demo.py state Agent 状态检查点持久化 + 跨会话恢复并校验
|
||||
python demo.py offline 依次运行上面全部三个离线演示(默认行为)
|
||||
|
||||
【LLM 场景】需要 OPENAI_API_KEY(或 MOONSHOT/ARK),由真实模型做决策——
|
||||
python demo.py scenarios 依次运行书中四个验证场景
|
||||
python demo.py scenarios --scenario 1 只跑场景 1(异步执行 + 即时提问)
|
||||
python demo.py scenarios --scenario 3 只跑场景 3(打断机制)
|
||||
|
||||
不带任何子命令时运行【离线演示】,因此开箱即用、无需联网。
|
||||
为兼容旧用法,`python demo.py --scenario N` 等价于 `scenarios --scenario N`。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
|
||||
try:
|
||||
from dotenv import load_dotenv
|
||||
load_dotenv()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
from async_demos import OFFLINE_DEMOS, banner
|
||||
from runtime import AgentRuntime
|
||||
|
||||
# openai 仅在运行 LLM 场景时才惰性导入;离线演示不碰它,保证无 key/无 openai 也能跑。
|
||||
|
||||
|
||||
def _completion_params_for(model: str) -> dict:
|
||||
"""按模型返回安全的采样参数。
|
||||
|
||||
Moonshot kimi-k3 是【推理模型】:必须 temperature=1 且 max_tokens>=2048,
|
||||
否则可能报错或截断。其余模型用 temperature=0.2 保证决策稳定。
|
||||
"""
|
||||
if model.startswith("kimi-k3"):
|
||||
return {"temperature": 1, "max_tokens": 4096}
|
||||
return {"temperature": 0.2}
|
||||
|
||||
|
||||
def _map_model_for_openrouter(model: str) -> str:
|
||||
"""把常见模型名映射成 OpenRouter 的 `provider/model` 形式。
|
||||
|
||||
- 已含 "/" 的 id(如 anthropic/claude-opus-4.8、google/gemini-2.5-pro)原样透传。
|
||||
- gpt-*/o1-*/o3-*/o4-* -> openai/…
|
||||
- claude-* -> anthropic/claude-opus-4.8
|
||||
- 其它保持原样(交给 OpenRouter 校验)。
|
||||
"""
|
||||
if "/" in model:
|
||||
return model
|
||||
m = model.lower()
|
||||
if m.startswith(("gpt-", "o1-", "o3-", "o4-")):
|
||||
return f"openai/{model}"
|
||||
if m.startswith("claude-"):
|
||||
return "anthropic/claude-opus-4.8"
|
||||
return model
|
||||
|
||||
|
||||
def make_client():
|
||||
"""按 LLM_PROVIDER 选择可用的模型服务(默认 openai)。
|
||||
|
||||
返回 (client, model, completion_params)。
|
||||
|
||||
通用兜底:当直连 provider 的 key 缺失、但存在 OPENROUTER_API_KEY 时,
|
||||
自动改走 OpenRouter(api_key=OPENROUTER_API_KEY,base_url=openrouter.ai/api/v1,
|
||||
并把模型名映射成 provider/model 形式),从而"有 OpenRouter key 就能跑"。
|
||||
"""
|
||||
from openai import AsyncOpenAI # 惰性导入:离线演示无需安装 openai
|
||||
provider = os.getenv("LLM_PROVIDER", "openai").lower()
|
||||
if provider in {"dashscope", "qwen", "bailian"}:
|
||||
key = os.environ["DASHSCOPE_API_KEY"]
|
||||
model = os.getenv("LLM_MODEL", "qwen3.7-plus")
|
||||
base_url = os.getenv(
|
||||
"DASHSCOPE_BASE_URL",
|
||||
"https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
)
|
||||
client = AsyncOpenAI(api_key=key, base_url=base_url)
|
||||
return client, model, _completion_params_for(model)
|
||||
if provider == "moonshot":
|
||||
key = os.environ["MOONSHOT_API_KEY"]
|
||||
# 默认用当前的推理模型 kimi-k3(旧的 kimi-k2-*-preview 与 moonshot-v1-* 均已过时/停用)。
|
||||
model = os.getenv("LLM_MODEL", "kimi-k3")
|
||||
client = AsyncOpenAI(api_key=key, base_url="https://api.moonshot.cn/v1")
|
||||
return client, model, _completion_params_for(model)
|
||||
if provider == "ark":
|
||||
key = os.environ["ARK_API_KEY"]
|
||||
model = os.getenv("LLM_MODEL") # ARK 需要填 endpoint id
|
||||
if not model:
|
||||
raise SystemExit("使用 ARK 时请设置 LLM_MODEL 为你的推理接入点 ID")
|
||||
client = AsyncOpenAI(api_key=key, base_url="https://ark.cn-beijing.volces.com/api/v3")
|
||||
return client, model, _completion_params_for(model)
|
||||
if provider == "openrouter":
|
||||
key = os.environ["OPENROUTER_API_KEY"]
|
||||
model = _map_model_for_openrouter(os.getenv("LLM_MODEL", "openai/gpt-5.6-luna"))
|
||||
client = AsyncOpenAI(api_key=key, base_url="https://openrouter.ai/api/v1")
|
||||
return client, model, _completion_params_for(model)
|
||||
key = os.getenv("OPENAI_API_KEY")
|
||||
or_key = os.getenv("OPENROUTER_API_KEY")
|
||||
model = os.getenv("LLM_MODEL", "gpt-5.6-luna")
|
||||
# gpt-5.x(含 gpt-5.6*)直连 OpenAI 需要组织验证;只要有 OPENROUTER_API_KEY,
|
||||
# 就优先走 OpenRouter;直连 OPENAI_API_KEY 缺失时同样兜底到 OpenRouter。
|
||||
if or_key and (not key or model.lower().startswith("gpt-5")):
|
||||
mapped = _map_model_for_openrouter(model)
|
||||
client = AsyncOpenAI(api_key=or_key, base_url="https://openrouter.ai/api/v1")
|
||||
return client, mapped, _completion_params_for(mapped)
|
||||
if key:
|
||||
base = os.getenv("OPENAI_BASE_URL")
|
||||
client = AsyncOpenAI(api_key=key, base_url=base) if base else AsyncOpenAI(api_key=key)
|
||||
return client, model, _completion_params_for(model)
|
||||
raise SystemExit(
|
||||
"未找到可用的 LLM Key。请设置以下任意一项:"
|
||||
"OPENAI_API_KEY 或 OPENROUTER_API_KEY(或 LLM_PROVIDER=moonshot 且 MOONSHOT_API_KEY / "
|
||||
"LLM_PROVIDER=ark 且 ARK_API_KEY)。"
|
||||
)
|
||||
|
||||
|
||||
async def run_runtime(rt: AgentRuntime):
|
||||
"""在后台跑事件循环。"""
|
||||
return asyncio.create_task(rt.serve())
|
||||
|
||||
|
||||
# ------------------------------- 四个场景 -------------------------------
|
||||
|
||||
async def scenario_1(client, model, params):
|
||||
banner("场景 1|异步工具执行:长任务运行期间即时回应插入的提问")
|
||||
rt = AgentRuntime(client, model, completion_params=params)
|
||||
serve = await run_runtime(rt)
|
||||
|
||||
# 用户下达一个耗时的日志分析任务
|
||||
await rt.submit_user_message(
|
||||
"请运行终端命令 `python analyze_logs.py`(这是耗时的日志分析),完成后给我分析结论。",
|
||||
urgency="immediate")
|
||||
await asyncio.sleep(2.2) # 任务已在后台跑
|
||||
|
||||
# 期间用户插入一个即时问题
|
||||
await rt.submit_user_message("现在几点了?") # 带问号 -> 立即回应
|
||||
|
||||
await rt.wait_until_idle()
|
||||
await rt.stop(); await serve
|
||||
|
||||
|
||||
async def scenario_2(client, model, params):
|
||||
banner("场景 2|事件队列与批量处理:非紧急指令累积,任务完成时一次性处理")
|
||||
rt = AgentRuntime(client, model, completion_params=params)
|
||||
serve = await run_runtime(rt)
|
||||
|
||||
await rt.submit_user_message(
|
||||
"请运行终端命令 `python analyze_logs.py`(耗时日志分析),完成后把分析结论告诉我。",
|
||||
urgency="immediate")
|
||||
await asyncio.sleep(1.5)
|
||||
|
||||
# 连续发两条补充性指令(无问号 -> 非紧急,进入排队缓冲)
|
||||
await rt.submit_user_message("记得最后用日语回复")
|
||||
await asyncio.sleep(0.4)
|
||||
await rt.submit_user_message("把结果整理成一个网页(HTML)")
|
||||
|
||||
await rt.wait_until_idle()
|
||||
await rt.stop(); await serve
|
||||
|
||||
|
||||
async def scenario_3(client, model, params):
|
||||
banner("场景 3|打断机制:用户'取消'立即终止执行流并取消异步工具")
|
||||
rt = AgentRuntime(client, model, completion_params=params)
|
||||
serve = await run_runtime(rt)
|
||||
|
||||
await rt.submit_user_message(
|
||||
"请运行终端命令 `python analyze_logs.py`(耗时日志分析),完成后给我结论。",
|
||||
urgency="immediate")
|
||||
await asyncio.sleep(4.0) # 等后台任务确实跑起来(跑到一半左右)
|
||||
|
||||
await rt.submit_user_message("取消") # 打断关键词 -> 立即取消
|
||||
|
||||
await rt.wait_until_idle(stable=1.0)
|
||||
await rt.stop(); await serve
|
||||
|
||||
|
||||
async def scenario_4(client, model, params):
|
||||
banner("场景 4|并行工具的取消与状态查询:三脚本竞速 + 按 50% 阈值取消 + 整合报告")
|
||||
rt = AgentRuntime(client, model, completion_params=params)
|
||||
serve = await run_runtime(rt)
|
||||
|
||||
await rt.submit_user_message(
|
||||
"同时运行这三个分析脚本:`python analyze_fast.py`、`python analyze_mid.py`、`python analyze_slow.py`。"
|
||||
"哪个脚本先完成,你就查询另外两个脚本的进度;如果某个脚本进度还没超过 50%,就取消它;"
|
||||
"其余脚本完成后,把所有已完成脚本的结果整合成一份报告给我。",
|
||||
urgency="immediate")
|
||||
|
||||
await rt.wait_until_idle(stable=1.5, timeout=60)
|
||||
await rt.stop(); await serve
|
||||
|
||||
|
||||
SCENARIOS = {1: scenario_1, 2: scenario_2, 3: scenario_3, 4: scenario_4}
|
||||
|
||||
|
||||
# ------------------------------- 子命令实现 -------------------------------
|
||||
|
||||
async def run_offline(names: list[str]) -> None:
|
||||
"""运行离线演示(无需 API key)。"""
|
||||
for name in names:
|
||||
await OFFLINE_DEMOS[name]()
|
||||
|
||||
|
||||
async def run_scenarios(which: int | None) -> None:
|
||||
"""运行 LLM 驱动的验证场景(需要 API key)。"""
|
||||
client, model, params = make_client()
|
||||
print(f"使用模型:{model}")
|
||||
todo = [which] if which else [1, 2, 3, 4]
|
||||
for i in todo:
|
||||
await SCENARIOS[i](client, model, params)
|
||||
await asyncio.sleep(0.5)
|
||||
|
||||
|
||||
def build_parser() -> argparse.ArgumentParser:
|
||||
parser = argparse.ArgumentParser(
|
||||
prog="demo.py",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
description="实验 6-2:带并行执行、打断/取消与状态管理的异步 Agent 演示。",
|
||||
epilog=(
|
||||
"示例:\n"
|
||||
" python demo.py # 默认:依次运行三个离线演示(无需 API key)\n"
|
||||
" python demo.py parallel # 并行 vs 串行的墙钟时间对比(打印加速比)\n"
|
||||
" python demo.py interrupt # 长任务运行中被打断/取消,随后恢复\n"
|
||||
" python demo.py state # 状态检查点持久化 + 跨会话恢复并校验\n"
|
||||
" python demo.py scenarios --scenario 3 # LLM 场景 3:打断机制(需 API key)\n"
|
||||
"\n离线演示不联网、不需要任何 key;scenarios 子命令需要 OPENAI_API_KEY(或 MOONSHOT/ARK)。"
|
||||
),
|
||||
)
|
||||
sub = parser.add_subparsers(dest="command", metavar="<子命令>")
|
||||
|
||||
sub.add_parser("parallel", help="并行 vs 串行工具调用的墙钟时间对比(离线,无需 key)")
|
||||
sub.add_parser("interrupt", help="长任务运行中被打断/取消,随后系统恢复(离线,无需 key)")
|
||||
sub.add_parser("state", help="Agent 状态检查点持久化与跨会话恢复(离线,无需 key)")
|
||||
sub.add_parser("offline", help="依次运行上面三个离线演示(默认行为)")
|
||||
|
||||
ps = sub.add_parser("scenarios", help="书中四个 LLM 验证场景(需要 API key)")
|
||||
ps.add_argument("--scenario", type=int, choices=[1, 2, 3, 4],
|
||||
help="只运行指定场景(1 异步执行 / 2 批量处理 / 3 打断 / 4 并行取消);不填则全部")
|
||||
return parser
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
# 兼容旧用法:`python demo.py --scenario N` 等价于 `scenarios --scenario N`
|
||||
argv = sys.argv[1:]
|
||||
if argv and argv[0].startswith("-") and argv[0] not in ("-h", "--help"):
|
||||
argv = ["scenarios"] + argv
|
||||
|
||||
args = build_parser().parse_args(argv)
|
||||
cmd = args.command or "offline"
|
||||
|
||||
if cmd == "scenarios":
|
||||
await run_scenarios(args.scenario)
|
||||
elif cmd == "offline":
|
||||
await run_offline(["parallel", "interrupt", "state"])
|
||||
else: # parallel / interrupt / state
|
||||
await run_offline([cmd])
|
||||
|
||||
print("\n演示结束。")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
@@ -0,0 +1,37 @@
|
||||
# 复制为 .env 后填入你的 key(demo.py 会自动加载)
|
||||
|
||||
# ===== LLM 服务(默认 openai;也支持 moonshot、dashscope/qwen/bailian、ark)=====
|
||||
LLM_PROVIDER=openai
|
||||
|
||||
# OpenAI(默认)
|
||||
OPENAI_API_KEY=your-openai-api-key
|
||||
# 可选:自定义模型 / 网关
|
||||
LLM_MODEL=gpt-5.6-luna
|
||||
# OPENAI_BASE_URL=https://your-gateway/v1
|
||||
|
||||
# Moonshot(LLM_PROVIDER=moonshot;默认模型为推理模型 kimi-k3,代码会自动用 temperature=1 且 max_tokens>=2048)
|
||||
MOONSHOT_API_KEY=your-moonshot-api-key
|
||||
# LLM_MODEL=kimi-k3
|
||||
|
||||
# Alibaba Cloud Model Studio / Bailian (Qwen)
|
||||
# DASHSCOPE_API_KEY=your-dashscope-api-key
|
||||
# LLM_PROVIDER=dashscope
|
||||
# LLM_MODEL=qwen3.7-plus
|
||||
# DASHSCOPE_BASE_URL=https://dashscope-intl.aliyuncs.com/compatible-mode/v1
|
||||
|
||||
# 火山方舟 ARK(LLM_PROVIDER=ark,LLM_MODEL 需填推理接入点 ID)
|
||||
ARK_API_KEY=xxxx
|
||||
# LLM_MODEL=ep-xxxxxxxx
|
||||
|
||||
# ===== OpenRouter 通用兜底 =====
|
||||
# 当没有配置直连的 OPENAI_API_KEY(且未使用 moonshot/ark provider)时,
|
||||
# 只要设置了 OPENROUTER_API_KEY,demo.py 会自动改走 OpenRouter:
|
||||
# base_url=https://openrouter.ai/api/v1,并把模型名映射成 provider/model 形式
|
||||
# (gpt-* -> openai/…、claude-* -> anthropic/claude-opus-4.8、含 "/" 的原样透传)。
|
||||
# 也可显式 LLM_PROVIDER=openrouter 强制使用。
|
||||
# OPENROUTER_API_KEY=your-openrouter-api-key
|
||||
# LLM_MODEL=openai/gpt-5.6-luna
|
||||
|
||||
# ===== 时间轴加速(可选)=====
|
||||
# 1 个"模拟秒"对应的真实秒数,默认 0.4(2.5 倍速)。数值越小演示越快。
|
||||
FLUX_TICK_REAL=0.4
|
||||
@@ -0,0 +1,91 @@
|
||||
"""事件模型(对应设计文档中的 Event / Trajectory 概念)。
|
||||
|
||||
Flux 把 Agent 的一切经历都抽象成"事件",按时间顺序追加到轨迹(trajectory)里。
|
||||
本文件定义事件类型、事件对象,以及"事件紧急度"的判定逻辑——这是实验 6-2 里
|
||||
"批量处理 vs 立即打断"两种处理机制的分类依据。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Optional
|
||||
|
||||
|
||||
class EventType:
|
||||
"""事件类型常量(对应设计文档第 2 节 Inputs / Interrupts / Thinking / Actions)。"""
|
||||
|
||||
USER_INPUT = "user.input" # 用户输入(非紧急,走"排队处理")
|
||||
USER_INTERRUPT = "user.interrupt" # 用户打断(紧急,走"取消式处理")
|
||||
AGENT_OUTPUT = "agent.output" # Agent 面向用户的最终回复
|
||||
AGENT_TOOL_CALL = "agent.tool_call" # Agent 发起的工具调用(Action)
|
||||
TOOL_RESULT = "tool.result" # 工具返回结果(同步工具 / 异步占位符)
|
||||
ASYNC_RESULT = "async.result" # 异步工具真正完成后注入的新事件
|
||||
SYSTEM_NOTE = "system.note" # 框架注入的系统提示(如取消回执)
|
||||
|
||||
|
||||
class Urgency:
|
||||
"""事件紧急度:决定采用哪种事件处理机制。"""
|
||||
|
||||
INTERRUPT = "interrupt" # 取消式处理:立刻打断当前执行并取消异步工具
|
||||
IMMEDIATE = "immediate" # 立即处理:不打断后台异步任务,但马上回应(如用户提问)
|
||||
DEFERRED = "deferred" # 排队处理:累积到 pending 队列,任务完成时批量追加
|
||||
|
||||
|
||||
# 打断类关键词:命中即视为紧急打断
|
||||
_INTERRUPT_KEYWORDS = ["取消", "停止", "中止", "打住", "别做了", "stop", "cancel", "abort"]
|
||||
|
||||
# 疑问类信号:命中即视为需要"立即回应"(而不是排队)
|
||||
_QUESTION_MARKS = ("?", "?")
|
||||
_QUESTION_KEYWORDS = ["几点", "多少", "怎么", "如何", "为什么", "是不是", "有没有",
|
||||
"吗", "呢", "what", "when", "how", "why", "which"]
|
||||
|
||||
|
||||
def classify_urgency(text: str) -> str:
|
||||
"""根据用户消息内容判定紧急度。
|
||||
|
||||
规则(简单、可解释,便于书中讲清楚):
|
||||
1. 含打断关键词(取消/停止/stop...) -> INTERRUPT(紧急,取消式处理)
|
||||
2. 是一个提问(带问号或疑问词) -> IMMEDIATE(立即回应,但不打断后台任务)
|
||||
3. 其它(补充性指令,如"用日语回复")-> DEFERRED(排队,批量处理)
|
||||
"""
|
||||
low = text.lower()
|
||||
if any(kw in text or kw in low for kw in _INTERRUPT_KEYWORDS):
|
||||
return Urgency.INTERRUPT
|
||||
if text.strip().endswith(_QUESTION_MARKS) or any(kw in text or kw in low for kw in _QUESTION_KEYWORDS):
|
||||
return Urgency.IMMEDIATE
|
||||
return Urgency.DEFERRED
|
||||
|
||||
|
||||
@dataclass
|
||||
class Event:
|
||||
"""一条轨迹事件。
|
||||
|
||||
message 字段保存"可直接喂给 LLM 的 OpenAI 消息字典"(保证上下文的高保真回放);
|
||||
没有 message 的事件(若有)只用于日志。
|
||||
"""
|
||||
|
||||
type: str
|
||||
message: Optional[dict] = None # OpenAI chat 格式消息,供构建 LLM 上下文
|
||||
label: str = "" # 人类可读的日志标签
|
||||
task_id: Optional[str] = None # 关联的异步任务 ID(若有)
|
||||
urgency: Optional[str] = None # 仅用户输入事件会带
|
||||
ts: float = field(default_factory=time.time)
|
||||
|
||||
def to_dict(self) -> dict:
|
||||
"""序列化为纯 JSON 可写的字典(用于状态检查点持久化)。"""
|
||||
return {
|
||||
"type": self.type, "message": self.message, "label": self.label,
|
||||
"task_id": self.task_id, "urgency": self.urgency, "ts": self.ts,
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, d: dict) -> "Event":
|
||||
"""从检查点字典还原事件对象。"""
|
||||
raw_ts = d.get("ts")
|
||||
return cls(
|
||||
type=d["type"], message=d.get("message"),
|
||||
label=d.get("label") or "",
|
||||
task_id=d.get("task_id"), urgency=d.get("urgency"),
|
||||
ts=raw_ts if raw_ts is not None else time.time(),
|
||||
)
|
||||
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"experiment": "6-2",
|
||||
"title": "Real subprocess async execution, queueing, interruption, and progress cancellation",
|
||||
"authority": "book/chapter4.md:579",
|
||||
"execution": {
|
||||
"mode": "allowlisted asyncio subprocesses",
|
||||
"shell": false,
|
||||
"worker": "analysis_worker.py",
|
||||
"input": "book/chapter4.md",
|
||||
"progress_source": "child stdout",
|
||||
"result_source": "child-computed file metrics"
|
||||
},
|
||||
"scenarios": [
|
||||
"long command plus immediate current-time response before completion",
|
||||
"two deferred instructions batch on completion and produce Japanese HTML",
|
||||
"user cancellation terminates the child process and runtime recovers",
|
||||
"3/2/1 percent jobs; query remaining jobs once; cancel only progress at or below 50 percent"
|
||||
],
|
||||
"acceptance": {
|
||||
"long_job_at_least_three_seconds": true,
|
||||
"placeholder_return_is_nonblocking": true,
|
||||
"all_terminal_jobs_are_real_subprocesses": true,
|
||||
"cancelled_jobs_have_os_return_codes": true,
|
||||
"completed_jobs_have_stdout_and_input_hashes": true,
|
||||
"all_artifacts_are_hash_manifested": true,
|
||||
"no_simulated_terminal_result_can_pass": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,2 @@
|
||||
openai>=1.30.0
|
||||
python-dotenv>=1.0.0
|
||||
@@ -0,0 +1,423 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run all four Experiment 6-2 scenarios with real child processes."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import tasks
|
||||
from tasks import TaskManager, TaskState
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
PROTOCOL_PATH = HERE / "experiment_protocol.json"
|
||||
VALIDATION_ROOT = HERE / "validation" / "experiment_6_2"
|
||||
UTC = timezone.utc
|
||||
|
||||
|
||||
def canonical_json(value: Any) -> str:
|
||||
return json.dumps(value, ensure_ascii=False, sort_keys=True,
|
||||
separators=(",", ":"), default=str)
|
||||
|
||||
|
||||
def sha256(value: bytes | str) -> str:
|
||||
if isinstance(value, str):
|
||||
value = value.encode()
|
||||
return hashlib.sha256(value).hexdigest()
|
||||
|
||||
|
||||
def write_json(path: Path, value: Any) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2, default=str) + "\n",
|
||||
encoding="utf-8")
|
||||
|
||||
|
||||
def task_receipt(state: TaskState) -> dict[str, Any]:
|
||||
result: Any = state.result
|
||||
try:
|
||||
result = json.loads(state.result) if state.result else None
|
||||
except json.JSONDecodeError:
|
||||
pass
|
||||
return {
|
||||
"task_id": state.task_id, "command": state.command,
|
||||
"status": state.status, "progress": state.progress,
|
||||
"pid": state.pid, "returncode": state.returncode,
|
||||
"started_at": state.started_at, "completed_at": state.completed_at,
|
||||
"elapsed_seconds": (
|
||||
round(state.completed_at - state.started_at, 3)
|
||||
if state.started_at and state.completed_at else None
|
||||
),
|
||||
"stdout_sha256": state.stdout_sha256,
|
||||
"stderr_tail": state.stderr_tail,
|
||||
"result": result,
|
||||
"executable": state.executable_receipt,
|
||||
}
|
||||
|
||||
|
||||
class ReceiptLog:
|
||||
def __init__(self):
|
||||
self.started = time.perf_counter()
|
||||
self.events: list[dict[str, Any]] = []
|
||||
|
||||
def __call__(self, source: str, text: str) -> None:
|
||||
self.events.append({"elapsed_seconds": round(time.perf_counter() - self.started, 3),
|
||||
"source": source, "text": text})
|
||||
|
||||
def add(self, source: str, event: str, **details: Any) -> float:
|
||||
elapsed = round(time.perf_counter() - self.started, 3)
|
||||
self.events.append({"elapsed_seconds": elapsed, "source": source,
|
||||
"event": event, **details})
|
||||
return elapsed
|
||||
|
||||
|
||||
async def _await_cancelled(state: TaskState) -> None:
|
||||
if state._task is None:
|
||||
return
|
||||
try:
|
||||
await state._task
|
||||
except asyncio.CancelledError:
|
||||
pass
|
||||
|
||||
|
||||
async def scenario_1() -> dict[str, Any]:
|
||||
log = ReceiptLog()
|
||||
completed_at: dict[str, float] = {}
|
||||
|
||||
async def complete(state: TaskState) -> None:
|
||||
completed_at[state.task_id] = log.add("SYSTEM", "async_result_injected",
|
||||
task_id=state.task_id)
|
||||
|
||||
manager = TaskManager(complete, log)
|
||||
before = time.perf_counter()
|
||||
state = manager.start("python analyze_logs.py")
|
||||
placeholder_latency = round(time.perf_counter() - before, 6)
|
||||
placeholder_at = log.add("TOOL", "placeholder_returned", task_id=state.task_id,
|
||||
placeholder_latency_seconds=placeholder_latency)
|
||||
await asyncio.sleep(0.5)
|
||||
time_answer = datetime.now().astimezone().isoformat(timespec="seconds")
|
||||
time_answer_at = log.add("AGENT", "immediate_time_answer", answer=time_answer,
|
||||
task_still_running=state.status == "running")
|
||||
assert state._task is not None
|
||||
await state._task
|
||||
return {
|
||||
"id": "async_command_and_immediate_question",
|
||||
"placeholder_latency_seconds": placeholder_latency,
|
||||
"placeholder_at": placeholder_at, "time_answer_at": time_answer_at,
|
||||
"completion_event_at": completed_at.get(state.task_id),
|
||||
"time_answer": time_answer, "events": log.events,
|
||||
"tasks": [task_receipt(state)],
|
||||
}
|
||||
|
||||
|
||||
async def scenario_2(campaign_dir: Path) -> dict[str, Any]:
|
||||
log = ReceiptLog()
|
||||
pending: list[dict[str, Any]] = []
|
||||
batch: list[dict[str, Any]] = []
|
||||
|
||||
async def complete(state: TaskState) -> None:
|
||||
batch.append({"type": "async.result", "task_id": state.task_id,
|
||||
"result_sha256": sha256(state.result)})
|
||||
batch.extend(pending)
|
||||
log.add("SYSTEM", "batch_appended", event_count=len(batch),
|
||||
deferred_count=len(pending))
|
||||
|
||||
manager = TaskManager(complete, log)
|
||||
state = manager.start("python analyze_logs.py")
|
||||
log.add("TOOL", "placeholder_returned", task_id=state.task_id)
|
||||
await asyncio.sleep(0.5)
|
||||
pending.append({"type": "user.input", "instruction": "記得最後用日語回覆"})
|
||||
first_at = log.add("QUEUE", "deferred_instruction", instruction="japanese")
|
||||
await asyncio.sleep(0.2)
|
||||
pending.append({"type": "user.input", "instruction": "結果をHTMLウェブページに整理"})
|
||||
second_at = log.add("QUEUE", "deferred_instruction", instruction="html")
|
||||
assert state._task is not None
|
||||
await state._task
|
||||
metrics = json.loads(state.result)
|
||||
html = (
|
||||
"<!DOCTYPE html><html lang=\"ja\"><meta charset=\"utf-8\">"
|
||||
"<title>非同期分析レポート</title><body><h1>分析結果</h1>"
|
||||
f"<p>対象ファイルは {metrics['lines']} 行、{metrics['bytes']} バイトです。</p>"
|
||||
f"<p>非同期に関する言及は {metrics['async_mentions']} 件でした。</p>"
|
||||
"</body></html>\n"
|
||||
)
|
||||
output = campaign_dir / "artifacts" / "scenario_2_report.html"
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
output.write_text(html, encoding="utf-8")
|
||||
return {
|
||||
"id": "queued_batch_to_japanese_html",
|
||||
"instruction_times": [first_at, second_at],
|
||||
"batch": batch, "events": log.events,
|
||||
"artifact": {"path": str(output), "bytes": output.stat().st_size,
|
||||
"sha256": sha256(output.read_bytes()),
|
||||
"doctype": html.startswith("<!DOCTYPE html>"),
|
||||
"lang_ja": 'lang="ja"' in html,
|
||||
"has_japanese": bool(re.search(r"[ぁ-んァ-ン一-龯]", html))},
|
||||
"tasks": [task_receipt(state)],
|
||||
}
|
||||
|
||||
|
||||
async def scenario_3() -> dict[str, Any]:
|
||||
log = ReceiptLog()
|
||||
completed: list[str] = []
|
||||
|
||||
async def complete(state: TaskState) -> None:
|
||||
completed.append(state.task_id)
|
||||
log.add("SYSTEM", "async_result_injected", task_id=state.task_id)
|
||||
|
||||
manager = TaskManager(complete, log)
|
||||
interrupted = manager.start("python analyze_logs.py")
|
||||
await asyncio.sleep(0.8)
|
||||
progress_at_interrupt = interrupted.progress
|
||||
interrupt_at = log.add("USER", "user.interrupt", text="取消",
|
||||
task_id=interrupted.task_id,
|
||||
progress=progress_at_interrupt)
|
||||
cancel_started = time.perf_counter()
|
||||
cancelled = manager.cancel_all()
|
||||
await _await_cancelled(interrupted)
|
||||
cancel_latency = round(time.perf_counter() - cancel_started, 3)
|
||||
cancel_receipt_at = log.add("SYSTEM", "process_cancelled", task_ids=cancelled,
|
||||
cancel_latency_seconds=cancel_latency,
|
||||
returncode=interrupted.returncode)
|
||||
recovery = manager.start("python re_run_summary.py")
|
||||
recovery_at = log.add("SYSTEM", "runtime_recovered", task_id=recovery.task_id)
|
||||
assert recovery._task is not None
|
||||
await recovery._task
|
||||
return {
|
||||
"id": "interrupt_terminates_and_recovers",
|
||||
"interrupt_at": interrupt_at, "cancel_receipt_at": cancel_receipt_at,
|
||||
"cancel_latency_seconds": cancel_latency,
|
||||
"recovery_at": recovery_at, "completed_callbacks": completed,
|
||||
"events": log.events,
|
||||
"tasks": [task_receipt(interrupted), task_receipt(recovery)],
|
||||
}
|
||||
|
||||
|
||||
async def scenario_4(campaign_dir: Path) -> dict[str, Any]:
|
||||
log = ReceiptLog()
|
||||
completion_queue: asyncio.Queue[TaskState] = asyncio.Queue()
|
||||
|
||||
async def complete(state: TaskState) -> None:
|
||||
log.add("SYSTEM", "async_result_injected", task_id=state.task_id)
|
||||
await completion_queue.put(state)
|
||||
|
||||
manager = TaskManager(complete, log)
|
||||
states = [manager.start(command) for command in (
|
||||
"python analyze_fast.py", "python analyze_mid.py", "python analyze_slow.py"
|
||||
)]
|
||||
first = await asyncio.wait_for(completion_queue.get(), timeout=30)
|
||||
query_receipts = []
|
||||
cancelled_ids = []
|
||||
for state in states:
|
||||
if state.task_id == first.task_id or state.status != "running":
|
||||
continue
|
||||
query_receipts.append({"task_id": state.task_id, "status": state.status,
|
||||
"progress": state.progress,
|
||||
"queried_at": log.add("TOOL", "query_task",
|
||||
task_id=state.task_id,
|
||||
progress=state.progress)})
|
||||
if state.progress <= 50:
|
||||
if manager.cancel(state.task_id):
|
||||
cancelled_ids.append(state.task_id)
|
||||
log.add("TOOL", "cancel_task", task_id=state.task_id,
|
||||
progress=state.progress)
|
||||
for state in states:
|
||||
await _await_cancelled(state)
|
||||
report_rows = [task_receipt(state) for state in states]
|
||||
report = {
|
||||
"first_completed": first.task_id, "query_receipts": query_receipts,
|
||||
"cancelled_ids": cancelled_ids,
|
||||
"completed_results": {row["task_id"]: row["result"] for row in report_rows
|
||||
if row["status"] == "completed"},
|
||||
}
|
||||
output = campaign_dir / "artifacts" / "scenario_4_report.json"
|
||||
write_json(output, report)
|
||||
return {
|
||||
"id": "parallel_progress_threshold_cancellation",
|
||||
**report, "events": log.events, "tasks": report_rows,
|
||||
"artifact": {"path": str(output), "bytes": output.stat().st_size,
|
||||
"sha256": sha256(output.read_bytes())},
|
||||
}
|
||||
|
||||
|
||||
def real_task(receipt: dict[str, Any]) -> bool:
|
||||
executable = receipt.get("executable", {})
|
||||
common = (
|
||||
receipt.get("pid") is not None
|
||||
and executable.get("mode") == "real_subprocess"
|
||||
and executable.get("shell") is False
|
||||
and len(executable.get("worker_sha256", "")) == 64
|
||||
and len(executable.get("input_sha256", "")) == 64
|
||||
and len(executable.get("argv_sha256", "")) == 64
|
||||
)
|
||||
if receipt.get("status") == "completed":
|
||||
return common and receipt.get("returncode") == 0 \
|
||||
and len(receipt.get("stdout_sha256", "")) == 64 \
|
||||
and isinstance(receipt.get("result"), dict)
|
||||
if receipt.get("status") == "cancelled":
|
||||
return common and receipt.get("returncode") not in {None, 0} \
|
||||
and executable.get("cancelled") is True
|
||||
return False
|
||||
|
||||
|
||||
def derive_acceptance(scenarios: list[dict[str, Any]], protocol: dict[str, Any]) -> dict[str, Any]:
|
||||
# Explicit mapping from protocol acceptance keys to the gate keys that
|
||||
# enforce them. This prevents silent drift: adding a key to the protocol's
|
||||
# acceptance block without a corresponding gate entry raises an assertion
|
||||
# at run time, and removing a gate key leaves a dangling reference that
|
||||
# the coverage check also catches.
|
||||
PROTOCOL_TO_GATE: dict[str, str | tuple[str, ...]] = {
|
||||
"long_job_at_least_three_seconds": "scenario_1_nonblocking_and_immediate_response",
|
||||
"placeholder_return_is_nonblocking": "scenario_1_nonblocking_and_immediate_response",
|
||||
"all_terminal_jobs_are_real_subprocesses": "real_subprocess_receipts_only",
|
||||
"cancelled_jobs_have_os_return_codes": "scenario_3_os_process_cancelled_then_recovered",
|
||||
"completed_jobs_have_stdout_and_input_hashes": "real_subprocess_receipts_only",
|
||||
"all_artifacts_are_hash_manifested": (
|
||||
"scenario_2_japanese_html_artifact",
|
||||
"scenario_4_integrated_report_hashed",
|
||||
),
|
||||
"no_simulated_terminal_result_can_pass": "real_subprocess_receipts_only",
|
||||
}
|
||||
protocol_acceptance = protocol.get("acceptance", {})
|
||||
missing = set(protocol_acceptance) - set(PROTOCOL_TO_GATE)
|
||||
assert not missing, (
|
||||
f"protocol acceptance keys not covered by PROTOCOL_TO_GATE: {missing}"
|
||||
)
|
||||
by_id = {row["id"]: row for row in scenarios}
|
||||
one = by_id.get("async_command_and_immediate_question", {})
|
||||
two = by_id.get("queued_batch_to_japanese_html", {})
|
||||
three = by_id.get("interrupt_terminates_and_recovers", {})
|
||||
four = by_id.get("parallel_progress_threshold_cancellation", {})
|
||||
all_tasks = [task for scenario in scenarios for task in scenario.get("tasks", [])]
|
||||
q = four.get("query_receipts", [])
|
||||
q_by_id = {row["task_id"]: row for row in q}
|
||||
tasks4 = {row["task_id"]: row for row in four.get("tasks", [])}
|
||||
gates = {
|
||||
"exact_four_scenarios": len(scenarios) == 4 and len(by_id) == 4,
|
||||
"real_subprocess_receipts_only": bool(all_tasks) and all(real_task(row) for row in all_tasks),
|
||||
"scenario_1_nonblocking_and_immediate_response": (
|
||||
one.get("placeholder_latency_seconds", 1) < 0.1
|
||||
and one.get("time_answer_at", 999) < one.get("completion_event_at", -1)
|
||||
and one.get("tasks", [{}])[0].get("elapsed_seconds", 0) >= 3
|
||||
and one.get("tasks", [{}])[0].get("status") == "completed"
|
||||
),
|
||||
"scenario_2_deferred_events_batched_once": (
|
||||
len(two.get("batch", [])) == 3
|
||||
and two.get("batch", [{}])[0].get("type") == "async.result"
|
||||
and [row.get("type") for row in two.get("batch", [])[1:]]
|
||||
== ["user.input", "user.input"]
|
||||
and len([event for event in two.get("events", [])
|
||||
if event.get("event") == "batch_appended"]) == 1
|
||||
),
|
||||
"scenario_2_japanese_html_artifact": all([
|
||||
two.get("artifact", {}).get("doctype"), two.get("artifact", {}).get("lang_ja"),
|
||||
two.get("artifact", {}).get("has_japanese"),
|
||||
two.get("artifact", {}).get("bytes", 0) > 100,
|
||||
len(two.get("artifact", {}).get("sha256", "")) == 64,
|
||||
]),
|
||||
"scenario_3_os_process_cancelled_then_recovered": (
|
||||
len(three.get("tasks", [])) == 2
|
||||
and three["tasks"][0].get("status") == "cancelled"
|
||||
and three["tasks"][0].get("progress", 100) < 100
|
||||
and three["tasks"][0].get("returncode") not in {None, 0}
|
||||
and three["tasks"][1].get("status") == "completed"
|
||||
and three.get("cancel_receipt_at", 0) >= three.get("interrupt_at", 999)
|
||||
and three.get("recovery_at", 0) >= three.get("cancel_receipt_at", 999)
|
||||
),
|
||||
"scenario_4_exact_rates_and_fast_first": (
|
||||
four.get("first_completed") == "T1"
|
||||
and [row.get("executable", {}).get("rate_percent_per_logical_second")
|
||||
for row in four.get("tasks", [])] == [3.0, 2.0, 1.0]
|
||||
),
|
||||
"scenario_4_query_once_and_cancel_only_under_threshold": (
|
||||
len(q) == len(q_by_id) == 2
|
||||
and set(q_by_id) == {"T2", "T3"}
|
||||
and q_by_id["T2"]["progress"] > 50
|
||||
and q_by_id["T3"]["progress"] <= 50
|
||||
and four.get("cancelled_ids") == ["T3"]
|
||||
and tasks4.get("T2", {}).get("status") == "completed"
|
||||
and tasks4.get("T3", {}).get("status") == "cancelled"
|
||||
),
|
||||
"scenario_4_integrated_report_hashed": (
|
||||
set(four.get("completed_results", {})) == {"T1", "T2"}
|
||||
and four.get("artifact", {}).get("bytes", 0) > 100
|
||||
and len(four.get("artifact", {}).get("sha256", "")) == 64
|
||||
),
|
||||
}
|
||||
# Verify that every referenced gate key actually exists.
|
||||
referenced_gate_keys = set()
|
||||
for gate_spec in PROTOCOL_TO_GATE.values():
|
||||
if isinstance(gate_spec, str):
|
||||
referenced_gate_keys.add(gate_spec)
|
||||
else:
|
||||
referenced_gate_keys.update(gate_spec)
|
||||
dangling = referenced_gate_keys - set(gates)
|
||||
assert not dangling, (
|
||||
f"PROTOCOL_TO_GATE references gate keys that do not exist: {dangling}"
|
||||
)
|
||||
# Compute per-protocol-key coverage so an auditor can mechanically verify
|
||||
# that every acceptance declaration is enforced by at least one gate.
|
||||
protocol_coverage: dict[str, Any] = {}
|
||||
for proto_key, gate_spec in PROTOCOL_TO_GATE.items():
|
||||
gate_keys = (gate_spec,) if isinstance(gate_spec, str) else gate_spec
|
||||
protocol_coverage[proto_key] = {
|
||||
"enforced_by": list(gate_keys),
|
||||
"all_gates_passed": all(gates.get(gk, False) for gk in gate_keys),
|
||||
}
|
||||
return {"status": "passed" if all(gates.values()) else "failed", "gates": gates,
|
||||
"protocol_coverage": protocol_coverage,
|
||||
"protocol_sha256": sha256(canonical_json(protocol))}
|
||||
|
||||
|
||||
def manifest(campaign_dir: Path) -> dict[str, Any]:
|
||||
files = []
|
||||
for path in sorted(campaign_dir.rglob("*")):
|
||||
if path.is_file() and path.name != "manifest.json":
|
||||
data = path.read_bytes()
|
||||
files.append({"path": str(path.relative_to(campaign_dir)),
|
||||
"bytes": len(data), "sha256": sha256(data)})
|
||||
return {"generated_at": datetime.now(UTC).isoformat(), "files": files}
|
||||
|
||||
|
||||
async def run(campaign_id: str | None, tick_real: float) -> Path:
|
||||
if tick_real <= 0:
|
||||
raise ValueError("tick-real must be positive")
|
||||
tasks.TICK_REAL = tick_real
|
||||
protocol = json.loads(PROTOCOL_PATH.read_text(encoding="utf-8"))
|
||||
campaign_id = campaign_id or datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
|
||||
campaign_dir = VALIDATION_ROOT / campaign_id
|
||||
campaign_dir.mkdir(parents=True, exist_ok=False)
|
||||
write_json(campaign_dir / "protocol.json", protocol)
|
||||
started = time.perf_counter()
|
||||
scenarios = [await scenario_1(), await scenario_2(campaign_dir),
|
||||
await scenario_3(), await scenario_4(campaign_dir)]
|
||||
for scenario in scenarios:
|
||||
write_json(campaign_dir / "scenarios" / f"{scenario['id']}.json", scenario)
|
||||
acceptance = derive_acceptance(scenarios, protocol)
|
||||
summary = {"experiment": "6-2", "campaign_id": campaign_id,
|
||||
"generated_at": datetime.now(UTC).isoformat(),
|
||||
"tick_real_seconds": tick_real,
|
||||
"elapsed_seconds": round(time.perf_counter() - started, 3),
|
||||
"scenario_status": {row["id"]: "recorded" for row in scenarios},
|
||||
"acceptance": acceptance, "status": acceptance["status"]}
|
||||
write_json(campaign_dir / "summary.json", summary)
|
||||
write_json(campaign_dir / "manifest.json", manifest(campaign_dir))
|
||||
return campaign_dir
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--campaign-id")
|
||||
parser.add_argument("--tick-real", type=float, default=0.15)
|
||||
args = parser.parse_args()
|
||||
print(asyncio.run(run(args.campaign_id, args.tick_real)))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,406 @@
|
||||
"""Flux 异步 Agent 运行时(实验 6-2 核心)。
|
||||
|
||||
实现设计文档第 5 节的事件处理循环,重点覆盖实验 6-2 的四个能力:
|
||||
1. 异步工具执行:run_terminal_command 立即返回占位符,任务在后台跑。
|
||||
2. 事件队列与批量处理:非紧急事件进 pending,异步结果到达时一次性批量追加。
|
||||
3. 打断机制:用户"取消/停止"立即取消当前 turn + 所有异步工具,并留痕。
|
||||
4. 并行工具的取消与状态查询:query_task / cancel_task 按 ID 操作;
|
||||
异步完成后以"新事件"把真实结果注入对话。
|
||||
|
||||
架构(三个协程协作,全部基于 asyncio 单线程):
|
||||
- inbox 队列:所有进来的事件(用户输入、打断、异步完成通知)先入 inbox。
|
||||
- _dispatcher:从 inbox 取事件 -> 判定紧急度 -> 分流(立即处理 / 排队 / 打断)。
|
||||
- _worker :从 work 队列取"事件批次" -> 追加到轨迹 -> 跑一轮 LLM(run_llm_turn)。
|
||||
每一轮 LLM 作为可取消的子任务(turn_task),打断时直接 cancel 它。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
from typing import Optional
|
||||
|
||||
from events import Event, EventType, Urgency, classify_urgency
|
||||
from tasks import TaskManager, TaskState
|
||||
|
||||
# ------------------------- LLM 工具定义(function calling) -------------------------
|
||||
|
||||
TOOL_SCHEMAS = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "run_terminal_command",
|
||||
"description": ("异步执行一个受限的真实日志分析子进程。调用后立即返回 task_id,"
|
||||
"不会阻塞;进度来自子进程 stdout。自然完成后,真实返回码、输出哈希和"
|
||||
"文件分析指标会作为新的系统事件出现。取消会终止对应 OS 进程。"),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"command": {"type": "string", "description": "要执行的终端命令,如 `python analyze_logs.py`"},
|
||||
},
|
||||
"required": ["command"],
|
||||
},
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_current_time",
|
||||
"description": "立即返回当前时间。用于回答用户'现在几点了'之类的即时问题。",
|
||||
"parameters": {"type": "object", "properties": {}},
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "query_task",
|
||||
"description": "查询某个后台异步任务的当前进度与状态。",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"task_id": {"type": "string", "description": "任务 ID,如 T1"}},
|
||||
"required": ["task_id"],
|
||||
},
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "cancel_task",
|
||||
"description": "按 task_id 取消一个正在运行的后台异步任务。",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"task_id": {"type": "string", "description": "任务 ID,如 T1"}},
|
||||
"required": ["task_id"],
|
||||
},
|
||||
},
|
||||
},
|
||||
]
|
||||
|
||||
SYSTEM_PROMPT = """你是一个异步 Agent(基于 Flux 框架)。你可以调用工具来完成任务。
|
||||
|
||||
关键行为准则:
|
||||
1. run_terminal_command 是【异步】的:调用后命令在后台运行并立即返回 task_id。
|
||||
你应当简要告知用户"任务已在后台启动",然后【结束本轮回复,不要空等结果】。
|
||||
2. 当你看到形如 "[系统事件|异步任务完成] task_id=... 结果:..." 的消息时,
|
||||
说明后台任务真的完成了,这时再基于结果给出分析/整合结论。
|
||||
3. 如果用户在后台任务运行期间提出简短问题(例如"现在几点了?"),
|
||||
立即用对应工具(如 get_current_time)回答,【不要等待】后台任务。
|
||||
4. 你可以用 query_task 查询任意后台任务进度,用 cancel_task 按 ID 取消任务。
|
||||
5. 收到 "[用户打断]" 时,立即停止当前工作并简短确认已停止。
|
||||
6. 严格按用户给出的计划执行(例如"谁先完成就查其余进度,未过 50% 就取消")。
|
||||
注意:只取消【进度未超过 50%】的任务;进度已超过 50% 的任务应【保留并等待其完成】,不要取消它。
|
||||
每个还在运行的任务只需查询一次进度即可做出取消/保留决定,不要反复查询。
|
||||
7. 回答简洁、用中文,除非用户明确要求其它语言或格式。
|
||||
"""
|
||||
|
||||
MAX_STEPS = 8 # 单轮内最多的工具调用往返次数(防止死循环)
|
||||
|
||||
# 日志配色(各来源一种颜色),供 runtime 与离线演示脚本共用。
|
||||
_LOG_COLORS = {
|
||||
"USER": "\033[96m", "AGENT": "\033[92m", "TOOL": "\033[93m",
|
||||
"TASK": "\033[95m", "SYSTEM": "\033[90m", "TRAJ": "\033[94m",
|
||||
"STATE": "\033[95m",
|
||||
}
|
||||
|
||||
|
||||
def format_log(t0: float, source: str, text: str) -> str:
|
||||
"""把一条日志渲染成「[相对秒] 来源 | 文本」的彩色字符串。"""
|
||||
color = _LOG_COLORS.get(source, "")
|
||||
reset = "\033[0m" if color else ""
|
||||
return f"[{time.time() - t0:6.2f}s] {color}{source:6}{reset} | {text}"
|
||||
|
||||
|
||||
class AgentRuntime:
|
||||
def __init__(self, client, model: str, start_time: Optional[float] = None,
|
||||
completion_params: Optional[dict] = None):
|
||||
self.client = client
|
||||
self.model = model
|
||||
# 传给 chat.completions.create 的采样参数。默认 temperature=0.2 适合 gpt-5.6-luna;
|
||||
# 推理模型(如 Moonshot kimi-k3)需要 temperature=1 且 max_tokens>=2048,由 make_client 传入。
|
||||
self.completion_params = completion_params or {"temperature": 0.2}
|
||||
self._t0 = start_time or time.time()
|
||||
|
||||
self.trajectory: list[Event] = [] # 轨迹(工作记忆)
|
||||
self.inbox: asyncio.Queue = asyncio.Queue() # 所有进来的原始事件
|
||||
self.work: asyncio.Queue = asyncio.Queue() # 待处理的事件批次
|
||||
self.pending: list[Event] = [] # 非紧急事件的排队缓冲
|
||||
|
||||
self.tasks = TaskManager(on_complete=self._on_task_complete, log=self.log)
|
||||
self.turn_task: Optional[asyncio.Task] = None
|
||||
self.running = True
|
||||
self._STOP = object()
|
||||
|
||||
# ------------------------------- 日志 -------------------------------
|
||||
|
||||
def log(self, source: str, text: str) -> None:
|
||||
print(format_log(self._t0, source, text), flush=True)
|
||||
|
||||
def _append(self, event: Event) -> None:
|
||||
"""把事件追加到轨迹,并打印轨迹留痕。"""
|
||||
self.trajectory.append(event)
|
||||
self.log("TRAJ", f"+ {event.type:18} {event.label}")
|
||||
|
||||
def build_messages(self) -> list[dict]:
|
||||
"""把轨迹渲染成 OpenAI chat 消息列表。"""
|
||||
msgs = [{"role": "system", "content": SYSTEM_PROMPT}]
|
||||
for e in self.trajectory:
|
||||
if e.message:
|
||||
msgs.append(e.message)
|
||||
return msgs
|
||||
|
||||
# ------------------------- 对外接口:提交事件 -------------------------
|
||||
|
||||
async def submit_user_message(self, text: str, urgency: Optional[str] = None) -> None:
|
||||
"""提交一条用户消息(demo 用它模拟用户输入)。"""
|
||||
u = urgency or classify_urgency(text)
|
||||
if u == Urgency.INTERRUPT:
|
||||
ev = Event(EventType.USER_INTERRUPT, urgency=u,
|
||||
message={"role": "user", "content": f"[用户打断] {text}"},
|
||||
label=f"用户打断:{text}")
|
||||
else:
|
||||
ev = Event(EventType.USER_INPUT, urgency=u,
|
||||
message={"role": "user", "content": text},
|
||||
label=f"用户消息({u}):{text}")
|
||||
self.log("USER", f"({u}) {text}")
|
||||
await self.inbox.put(ev)
|
||||
|
||||
async def _on_task_complete(self, state: TaskState) -> None:
|
||||
"""异步任务自然完成 -> 把真实结果作为【新事件】注入 inbox。"""
|
||||
ev = Event(
|
||||
EventType.ASYNC_RESULT, task_id=state.task_id,
|
||||
message={"role": "user",
|
||||
"content": (f"[系统事件|异步任务完成] task_id={state.task_id} "
|
||||
f"命令=`{state.command}` 结果:{state.result}")},
|
||||
label=f"异步完成 {state.task_id}",
|
||||
)
|
||||
await self.inbox.put(ev)
|
||||
|
||||
# ------------------------------- 主循环 -------------------------------
|
||||
|
||||
async def serve(self) -> None:
|
||||
dispatcher = asyncio.create_task(self._dispatcher())
|
||||
worker = asyncio.create_task(self._worker())
|
||||
await asyncio.gather(dispatcher, worker)
|
||||
|
||||
def _is_idle(self) -> bool:
|
||||
return (not self.tasks.any_running()
|
||||
and self.work.empty()
|
||||
and self.inbox.empty()
|
||||
and (self.turn_task is None or self.turn_task.done()))
|
||||
|
||||
def _drain_pending(self) -> list[Event]:
|
||||
drained, self.pending = self.pending, []
|
||||
return drained
|
||||
|
||||
async def _dispatcher(self) -> None:
|
||||
"""事件分流:实现设计文档 5.1 的两种处理机制。"""
|
||||
while self.running:
|
||||
ev = await self.inbox.get()
|
||||
if ev is self._STOP:
|
||||
await self.work.put(self._STOP)
|
||||
break
|
||||
|
||||
if ev.type == EventType.USER_INTERRUPT:
|
||||
# —— 取消式处理:立刻打断当前 turn + 取消所有异步工具 ——
|
||||
await self._handle_interrupt(ev)
|
||||
|
||||
elif ev.type == EventType.ASYNC_RESULT:
|
||||
# —— 异步结果到达:批量把 pending 一并追加,再触发 LLM ——
|
||||
batch = [ev] + self._drain_pending()
|
||||
if len(batch) > 1:
|
||||
self.log("SYSTEM", f"异步结果到达,批量处理 {len(batch)-1} 条积压的非紧急事件")
|
||||
await self.work.put(batch)
|
||||
|
||||
elif ev.type == EventType.USER_INPUT:
|
||||
if ev.urgency == Urgency.IMMEDIATE:
|
||||
# 立即处理(如用户提问),不打断后台异步任务
|
||||
await self.work.put([ev])
|
||||
elif self._is_idle():
|
||||
# 空闲时,普通指令也直接处理(例如一开始下达的任务)
|
||||
await self.work.put([ev])
|
||||
else:
|
||||
# 排队处理:累积到 pending,等下一次异步结果时批量追加
|
||||
self.pending.append(ev)
|
||||
self.log("SYSTEM", f"事件进入排队缓冲(当前积压 {len(self.pending)} 条)")
|
||||
|
||||
async def _handle_interrupt(self, ev: Event) -> None:
|
||||
# 1) 取消正在进行的 LLM turn
|
||||
if self.turn_task and not self.turn_task.done():
|
||||
self.turn_task.cancel()
|
||||
# 2) 取消所有后台异步工具
|
||||
cancelled = self.tasks.cancel_all()
|
||||
# 3) 组装打断批次:打断事件 + 系统回执 + 被丢弃的积压事件(留痕)
|
||||
note = Event(
|
||||
EventType.SYSTEM_NOTE,
|
||||
message={"role": "user",
|
||||
"content": (f"[系统] 已执行打断:取消了后台任务 {cancelled or '(无)'}。"
|
||||
f"请向用户简短确认已停止。")},
|
||||
label=f"打断回执,取消任务 {cancelled or '(无)'}",
|
||||
)
|
||||
batch = [ev, note] + self._drain_pending()
|
||||
await self.work.put(batch)
|
||||
|
||||
async def _worker(self) -> None:
|
||||
"""逐批处理事件:追加到轨迹后跑一轮可被取消的 LLM。"""
|
||||
while self.running:
|
||||
batch = await self.work.get()
|
||||
if batch is self._STOP:
|
||||
break
|
||||
self.turn_task = asyncio.create_task(self._process_batch(batch))
|
||||
try:
|
||||
await self.turn_task
|
||||
except asyncio.CancelledError:
|
||||
self.log("SYSTEM", "当前 LLM turn 已被打断取消")
|
||||
|
||||
async def _process_batch(self, batch: list[Event]) -> None:
|
||||
for e in batch:
|
||||
self._append(e)
|
||||
await self.run_llm_turn()
|
||||
|
||||
# ------------------------------- LLM turn -------------------------------
|
||||
|
||||
async def run_llm_turn(self) -> None:
|
||||
"""调用 LLM 做决策;同步工具就地执行并回填,异步工具启动后回占位符。"""
|
||||
for _ in range(MAX_STEPS):
|
||||
messages = self.build_messages()
|
||||
_t = time.time()
|
||||
resp = await self.client.chat.completions.create(
|
||||
model=self.model, messages=messages,
|
||||
tools=TOOL_SCHEMAS, tool_choice="auto", **self.completion_params,
|
||||
)
|
||||
self.log("SYSTEM", f"LLM 调用耗时 {time.time()-_t:.2f}s({len(messages)} 条消息)")
|
||||
msg = resp.choices[0].message
|
||||
|
||||
assistant_msg: dict = {"role": "assistant", "content": msg.content or ""}
|
||||
if msg.tool_calls:
|
||||
assistant_msg["tool_calls"] = [
|
||||
{"id": tc.id, "type": "function",
|
||||
"function": {"name": tc.function.name, "arguments": tc.function.arguments}}
|
||||
for tc in msg.tool_calls
|
||||
]
|
||||
|
||||
self._append(Event(
|
||||
EventType.AGENT_TOOL_CALL if msg.tool_calls else EventType.AGENT_OUTPUT,
|
||||
message=assistant_msg,
|
||||
label=("调用工具 " + ", ".join(tc.function.name for tc in msg.tool_calls)
|
||||
if msg.tool_calls else "回复用户"),
|
||||
))
|
||||
|
||||
if msg.content and msg.content.strip():
|
||||
self.log("AGENT", msg.content.strip())
|
||||
|
||||
if not msg.tool_calls:
|
||||
return # 本轮结束:Agent 给出了最终回复
|
||||
|
||||
# 执行每个工具调用
|
||||
for tc in msg.tool_calls:
|
||||
name = tc.function.name
|
||||
try:
|
||||
args = json.loads(tc.function.arguments or "{}")
|
||||
except json.JSONDecodeError:
|
||||
args = {}
|
||||
result_text = self._exec_tool(name, args)
|
||||
self._append(Event(
|
||||
EventType.TOOL_RESULT,
|
||||
message={"role": "tool", "tool_call_id": tc.id, "content": result_text},
|
||||
label=f"工具结果 {name}",
|
||||
))
|
||||
|
||||
def _exec_tool(self, name: str, args: dict) -> str:
|
||||
"""执行工具,返回给 LLM 的文本结果。"""
|
||||
if name == "run_terminal_command":
|
||||
command = args.get("command", "")
|
||||
state = self.tasks.start(command)
|
||||
return (f"命令已在后台【异步】启动。task_id={state.task_id},命令=`{command}`。"
|
||||
f"我不会阻塞等待;任务完成后其结果会以系统事件形式返回。"
|
||||
f"可用 query_task('{state.task_id}') 查询进度或 cancel_task('{state.task_id}') 取消。")
|
||||
|
||||
if name == "get_current_time":
|
||||
now = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
|
||||
self.log("TOOL", f"get_current_time -> {now}")
|
||||
return f"当前时间是 {now}。"
|
||||
|
||||
if name == "query_task":
|
||||
tid = args.get("task_id", "")
|
||||
st = self.tasks.query(tid)
|
||||
if not st:
|
||||
return f"未找到任务 {tid}。"
|
||||
self.log("TOOL", f"query_task({tid}) -> {st.status} {st.progress:.0f}%")
|
||||
return f"task_id={tid} 命令=`{st.command}` 状态={st.status} 进度={st.progress:.0f}%。"
|
||||
|
||||
if name == "cancel_task":
|
||||
tid = args.get("task_id", "")
|
||||
st = self.tasks.query(tid)
|
||||
progress = f"{st.progress:.0f}%" if st else "未知"
|
||||
ok = self.tasks.cancel(tid)
|
||||
self.log("TOOL", f"cancel_task({tid}) -> {'已取消' if ok else '无法取消'} (进度 {progress})")
|
||||
return (f"任务 {tid} 已取消(取消时进度 {progress})。" if ok
|
||||
else f"任务 {tid} 无法取消(可能已完成或不存在)。")
|
||||
|
||||
return f"未知工具:{name}"
|
||||
|
||||
# ------------------------------- 收尾 -------------------------------
|
||||
|
||||
async def wait_until_idle(self, stable: float = 1.3, timeout: float = 90.0) -> None:
|
||||
"""阻塞直到系统持续空闲 stable 秒(或超时)。"""
|
||||
start = time.time()
|
||||
last_busy = time.time()
|
||||
while True:
|
||||
busy = (self.tasks.any_running() or not self.work.empty()
|
||||
or not self.inbox.empty() or bool(self.pending)
|
||||
or (self.turn_task is not None and not self.turn_task.done()))
|
||||
now = time.time()
|
||||
if busy:
|
||||
last_busy = now
|
||||
elif now - last_busy >= stable:
|
||||
return
|
||||
if now - start >= timeout:
|
||||
self.log("SYSTEM", "wait_until_idle 超时返回")
|
||||
return
|
||||
await asyncio.sleep(0.1)
|
||||
|
||||
async def stop(self) -> None:
|
||||
self.running = False
|
||||
await self.inbox.put(self._STOP)
|
||||
|
||||
# ------------------------- 状态检查点(持久化 / 恢复) -------------------------
|
||||
|
||||
def snapshot(self) -> dict:
|
||||
"""把 Agent 的可持久化状态导出为一个 JSON 友好的字典。
|
||||
|
||||
状态 = 轨迹(工作记忆)+ 全部异步任务的最后已知状态。这是「跨会话恢复」
|
||||
的基础:进程重启后,能据此还原对话上下文与后台任务的进度。
|
||||
"""
|
||||
return {
|
||||
"model": self.model,
|
||||
"saved_at": datetime.datetime.now().isoformat(timespec="seconds"),
|
||||
"trajectory": [e.to_dict() for e in self.trajectory],
|
||||
"tasks": self.tasks.snapshot(),
|
||||
}
|
||||
|
||||
def save_checkpoint(self, path: str) -> str:
|
||||
"""把当前状态写入检查点文件(JSON),返回文件路径。"""
|
||||
data = self.snapshot()
|
||||
dirname = os.path.dirname(path)
|
||||
if dirname:
|
||||
os.makedirs(dirname, exist_ok=True)
|
||||
with open(path, "w", encoding="utf-8") as f:
|
||||
json.dump(data, f, ensure_ascii=False, indent=2)
|
||||
self.log("STATE", f"已保存检查点 -> {path}"
|
||||
f"({len(data['trajectory'])} 条轨迹事件,{len(data['tasks'])} 个任务)")
|
||||
return path
|
||||
|
||||
def load_checkpoint(self, path: str) -> dict:
|
||||
"""从检查点文件恢复轨迹与任务状态(原地覆盖当前状态)。"""
|
||||
with open(path, "r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
trajectory = data.get("trajectory") or []
|
||||
tasks = data.get("tasks") or []
|
||||
self.trajectory = [Event.from_dict(d) for d in trajectory]
|
||||
self.tasks.restore(tasks)
|
||||
self.log("STATE", f"已从检查点恢复 <- {path}"
|
||||
f"({len(self.trajectory)} 条轨迹事件,{len(tasks)} 个任务)")
|
||||
return data
|
||||
@@ -0,0 +1,282 @@
|
||||
"""Real, bounded asynchronous terminal jobs for Experiment 6-2.
|
||||
|
||||
Commands are parsed without a shell and resolved through an explicit allowlist
|
||||
to ``analysis_worker.py``. Each job is a real child process whose stdout drives
|
||||
progress. Cancellation terminates that OS process; completion returns metrics
|
||||
computed from a real input file rather than a fabricated result string.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import shlex
|
||||
import sys
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Awaitable, Callable, Dict, Optional
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
WORKER = HERE / "analysis_worker.py"
|
||||
DEFAULT_INPUT = HERE.parent.parent / "book" / "chapter4.md"
|
||||
|
||||
|
||||
def _env_float(name: str, default: float) -> float:
|
||||
raw = os.getenv(name)
|
||||
if raw is None:
|
||||
return default
|
||||
try:
|
||||
value = float(raw)
|
||||
if value <= 0:
|
||||
raise ValueError
|
||||
return value
|
||||
except ValueError:
|
||||
print(f"⚠️ 环境变量 {name}={raw!r} 非法(应为正数),使用默认值 {default}")
|
||||
return default
|
||||
|
||||
|
||||
# One logical second maps to this many wall-clock seconds. The default retains
|
||||
# the manuscript's 3/2/1-percent ratios while keeping the demo practical.
|
||||
TICK_REAL = _env_float("FLUX_TICK_REAL", 0.4)
|
||||
|
||||
_COMMANDS = {
|
||||
"analyze_fast.py": ("fast", 3.0),
|
||||
"analyze_mid.py": ("mid", 2.0),
|
||||
"analyze_slow.py": ("slow", 1.0),
|
||||
"analyze_logs.py": ("logs", 4.5),
|
||||
"re_run_summary.py": ("recovery", 4.5),
|
||||
}
|
||||
|
||||
|
||||
def resolve_job(command: str) -> tuple[str, float]:
|
||||
"""Resolve a displayed terminal command to one safe executable profile."""
|
||||
parts = shlex.split(command)
|
||||
if len(parts) != 2 or Path(parts[0]).name not in {"python", "python3", Path(sys.executable).name}:
|
||||
raise ValueError("only `python <approved-analysis-script>.py` commands are allowed")
|
||||
script = Path(parts[1]).name
|
||||
if script not in _COMMANDS:
|
||||
raise ValueError(f"unapproved experiment command: {script}")
|
||||
return _COMMANDS[script]
|
||||
|
||||
|
||||
def resolve_rate(command: str) -> float:
|
||||
return resolve_job(command)[1]
|
||||
|
||||
|
||||
def _hash_file(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as stream:
|
||||
for block in iter(lambda: stream.read(1024 * 1024), b""):
|
||||
digest.update(block)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
@dataclass
|
||||
class TaskState:
|
||||
task_id: str
|
||||
command: str
|
||||
rate: float
|
||||
progress: float = 0.0
|
||||
status: str = "running" # running | completed | cancelled | failed | suspended
|
||||
result: str = ""
|
||||
pid: int | None = None
|
||||
returncode: int | None = None
|
||||
started_at: float | None = None
|
||||
completed_at: float | None = None
|
||||
stdout_sha256: str | None = None
|
||||
stderr_tail: str = ""
|
||||
executable_receipt: dict = field(default_factory=dict)
|
||||
_task: Optional[asyncio.Task] = field(default=None, repr=False)
|
||||
_process: Optional[asyncio.subprocess.Process] = field(default=None, repr=False)
|
||||
|
||||
|
||||
class TaskManager:
|
||||
"""Start, observe, query, and terminate allowlisted real subprocesses."""
|
||||
|
||||
def __init__(self, on_complete: Callable[[TaskState], Awaitable[None]],
|
||||
log: Callable[[str, str], None]):
|
||||
self._on_complete = on_complete
|
||||
self._log = log
|
||||
self._tasks: Dict[str, TaskState] = {}
|
||||
self._counter = 0
|
||||
|
||||
def start(self, command: str) -> TaskState:
|
||||
job, rate = resolve_job(command) # reject before allocating a task id
|
||||
if not WORKER.is_file() or not DEFAULT_INPUT.is_file():
|
||||
raise FileNotFoundError("analysis worker or Chapter 4 input is missing")
|
||||
self._counter += 1
|
||||
task_id = f"T{self._counter}"
|
||||
state = TaskState(task_id=task_id, command=command, rate=rate)
|
||||
state.executable_receipt = {
|
||||
"mode": "real_subprocess", "shell": False,
|
||||
"worker": str(WORKER), "worker_sha256": _hash_file(WORKER),
|
||||
"input": str(DEFAULT_INPUT), "input_sha256": _hash_file(DEFAULT_INPUT),
|
||||
"job": job, "rate_percent_per_logical_second": rate,
|
||||
"tick_real_seconds": TICK_REAL,
|
||||
}
|
||||
self._tasks[task_id] = state
|
||||
state._task = asyncio.create_task(self._run(state, job))
|
||||
self._log("TASK", f"启动真实子进程任务 {task_id}: `{command}` "
|
||||
f"(速度 {rate:.0f}%/逻辑秒)")
|
||||
return state
|
||||
|
||||
async def _terminate_process(self, state: TaskState) -> None:
|
||||
process = state._process
|
||||
if not process or process.returncode is not None:
|
||||
return
|
||||
process.terminate()
|
||||
try:
|
||||
await asyncio.wait_for(process.wait(), timeout=2)
|
||||
except asyncio.TimeoutError:
|
||||
process.kill()
|
||||
await process.wait()
|
||||
state.returncode = process.returncode
|
||||
|
||||
async def _run(self, state: TaskState, job: str) -> None:
|
||||
stdout_lines: list[str] = []
|
||||
state.started_at = time.time()
|
||||
argv = [
|
||||
sys.executable, "-I", "-u", str(WORKER),
|
||||
"--job", job, "--rate", str(state.rate),
|
||||
"--tick-real", str(TICK_REAL), "--input", str(DEFAULT_INPUT),
|
||||
]
|
||||
state.executable_receipt["argv_sha256"] = hashlib.sha256(
|
||||
json.dumps(argv, separators=(",", ":")).encode()
|
||||
).hexdigest()
|
||||
next_milestone = 20.0
|
||||
try:
|
||||
process = await asyncio.create_subprocess_exec(
|
||||
*argv, cwd=str(HERE),
|
||||
stdin=asyncio.subprocess.DEVNULL,
|
||||
stdout=asyncio.subprocess.PIPE,
|
||||
stderr=asyncio.subprocess.PIPE,
|
||||
)
|
||||
state._process = process
|
||||
state.pid = process.pid
|
||||
state.executable_receipt["pid"] = process.pid
|
||||
assert process.stdout is not None
|
||||
while True:
|
||||
raw = await process.stdout.readline()
|
||||
if not raw:
|
||||
break
|
||||
line = raw.decode("utf-8", errors="replace").rstrip()
|
||||
stdout_lines.append(line)
|
||||
if line.startswith("PROGRESS "):
|
||||
try:
|
||||
state.progress = max(
|
||||
state.progress, min(100.0, float(line.split()[1]))
|
||||
)
|
||||
except (IndexError, ValueError):
|
||||
raise RuntimeError(f"worker emitted invalid progress: {line!r}")
|
||||
if state.progress >= next_milestone:
|
||||
self._log("TASK", f"{state.task_id} `{state.command}` "
|
||||
f"进度 {state.progress:.0f}% (pid={state.pid})")
|
||||
next_milestone += 20.0
|
||||
elif line.startswith("RESULT "):
|
||||
payload = json.loads(line.removeprefix("RESULT "))
|
||||
state.result = json.dumps(payload, ensure_ascii=False, sort_keys=True)
|
||||
assert process.stderr is not None
|
||||
stderr = (await process.stderr.read()).decode("utf-8", errors="replace")
|
||||
state.stderr_tail = stderr[-4000:]
|
||||
state.returncode = await process.wait()
|
||||
state.completed_at = time.time()
|
||||
stdout = "\n".join(stdout_lines) + ("\n" if stdout_lines else "")
|
||||
state.stdout_sha256 = hashlib.sha256(stdout.encode()).hexdigest()
|
||||
state.executable_receipt.update({
|
||||
"returncode": state.returncode,
|
||||
"stdout_sha256": state.stdout_sha256,
|
||||
"stdout_lines": len(stdout_lines),
|
||||
"stderr_sha256": hashlib.sha256(stderr.encode()).hexdigest(),
|
||||
"elapsed_seconds": round(state.completed_at - state.started_at, 3),
|
||||
})
|
||||
if state.returncode != 0:
|
||||
state.status = "failed"
|
||||
raise RuntimeError(
|
||||
f"worker exited {state.returncode}: {state.stderr_tail[-500:]}"
|
||||
)
|
||||
if state.progress != 100.0 or not state.result:
|
||||
state.status = "failed"
|
||||
raise RuntimeError("worker completed without 100% progress and a RESULT receipt")
|
||||
state.status = "completed"
|
||||
self._log("TASK", f"{state.task_id} 完成 ✅ (pid={state.pid}, "
|
||||
f"returncode={state.returncode})")
|
||||
await self._on_complete(state)
|
||||
except asyncio.CancelledError:
|
||||
await self._terminate_process(state)
|
||||
state.status = "cancelled"
|
||||
state.completed_at = time.time()
|
||||
state.executable_receipt.update({
|
||||
"returncode": state.returncode,
|
||||
"cancelled": True,
|
||||
"elapsed_seconds": round(state.completed_at - state.started_at, 3)
|
||||
if state.started_at else None,
|
||||
})
|
||||
self._log("TASK", f"{state.task_id} 子进程已终止 🛑 "
|
||||
f"(pid={state.pid}, 进度 {state.progress:.0f}%)")
|
||||
raise
|
||||
except Exception as exc:
|
||||
await self._terminate_process(state)
|
||||
state.status = "failed"
|
||||
state.result = state.result or f"{type(exc).__name__}: {exc}"
|
||||
state.completed_at = time.time()
|
||||
self._log("TASK", f"{state.task_id} 失败 ❌: {exc}")
|
||||
|
||||
def query(self, task_id: str) -> Optional[TaskState]:
|
||||
return self._tasks.get(task_id)
|
||||
|
||||
def cancel(self, task_id: str) -> bool:
|
||||
state = self._tasks.get(task_id)
|
||||
if state and state.status == "running":
|
||||
if state._task:
|
||||
state._task.cancel()
|
||||
return True
|
||||
return False
|
||||
|
||||
def cancel_all(self) -> list[str]:
|
||||
cancelled = []
|
||||
for task_id, state in self._tasks.items():
|
||||
if state.status == "running":
|
||||
if state._task:
|
||||
state._task.cancel()
|
||||
cancelled.append(task_id)
|
||||
return cancelled
|
||||
|
||||
def any_running(self) -> bool:
|
||||
return any(state.status == "running" for state in self._tasks.values())
|
||||
|
||||
def all_states(self) -> list[TaskState]:
|
||||
return list(self._tasks.values())
|
||||
|
||||
def snapshot(self) -> list[dict]:
|
||||
return [
|
||||
{"task_id": state.task_id, "command": state.command,
|
||||
"rate": state.rate, "progress": state.progress,
|
||||
"status": state.status, "result": state.result,
|
||||
"pid": state.pid, "returncode": state.returncode,
|
||||
"started_at": state.started_at, "completed_at": state.completed_at,
|
||||
"stdout_sha256": state.stdout_sha256,
|
||||
"executable_receipt": state.executable_receipt}
|
||||
for state in self._tasks.values()
|
||||
]
|
||||
|
||||
def restore(self, records: list[dict]) -> None:
|
||||
for record in records:
|
||||
status = "suspended" if record["status"] == "running" else record["status"]
|
||||
receipt = record.get("executable_receipt")
|
||||
state = TaskState(
|
||||
task_id=record["task_id"], command=record["command"],
|
||||
rate=record["rate"], progress=record["progress"], status=status,
|
||||
result=record.get("result") or "", pid=record.get("pid"),
|
||||
returncode=record.get("returncode"),
|
||||
started_at=record.get("started_at"), completed_at=record.get("completed_at"),
|
||||
stdout_sha256=record.get("stdout_sha256"),
|
||||
executable_receipt=receipt if isinstance(receipt, dict) else {},
|
||||
)
|
||||
self._tasks[state.task_id] = state
|
||||
try:
|
||||
self._counter = max(self._counter, int(state.task_id.lstrip("T") or 0))
|
||||
except ValueError:
|
||||
pass
|
||||
@@ -0,0 +1,101 @@
|
||||
"""
|
||||
Test suite locking out TypeError and FileNotFoundError in AgentRuntime checkpointing
|
||||
when tasks is None, trajectory is None, or destination directory doesn't exist.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
sys.path.insert(0, os.path.abspath(os.path.dirname(__file__)))
|
||||
|
||||
from runtime import AgentRuntime
|
||||
from tasks import TaskManager
|
||||
|
||||
|
||||
def test_save_checkpoint_creates_nested_directories():
|
||||
"""
|
||||
Ensure save_checkpoint automatically creates parent directories when saving.
|
||||
"""
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
runtime = AgentRuntime.__new__(AgentRuntime)
|
||||
runtime.snapshot = MagicMock(return_value={'trajectory': [], 'tasks': []})
|
||||
runtime.log = MagicMock()
|
||||
|
||||
target_path = os.path.join(tmpdir, "nested", "sub", "checkpoint.json")
|
||||
result_path = runtime.save_checkpoint(target_path)
|
||||
|
||||
assert result_path == target_path
|
||||
assert os.path.exists(target_path)
|
||||
|
||||
|
||||
def test_load_checkpoint_handles_null_tasks_and_trajectory():
|
||||
"""
|
||||
Ensure load_checkpoint gracefully handles JSON containing "tasks": null and
|
||||
"trajectory": null with a real TaskManager without raising TypeError.
|
||||
"""
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
target_path = os.path.join(tmpdir, "checkpoint.json")
|
||||
with open(target_path, "w", encoding="utf-8") as f:
|
||||
json.dump({'trajectory': None, 'tasks': None}, f)
|
||||
|
||||
runtime = AgentRuntime.__new__(AgentRuntime)
|
||||
runtime.tasks = TaskManager(on_complete=MagicMock(), log=MagicMock())
|
||||
runtime.log = MagicMock()
|
||||
|
||||
data = runtime.load_checkpoint(target_path)
|
||||
assert data['tasks'] is None
|
||||
assert data['trajectory'] is None
|
||||
assert len(runtime.trajectory) == 0
|
||||
assert len(runtime.tasks._tasks) == 0
|
||||
runtime.log.assert_called_once()
|
||||
|
||||
def test_load_checkpoint_handles_null_task_fields_and_event_fields():
|
||||
"""
|
||||
Ensure load_checkpoint gracefully handles JSON containing null fields in
|
||||
task records and trajectory events without setting None for non-optional attributes.
|
||||
"""
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
target_path = os.path.join(tmpdir, "checkpoint.json")
|
||||
checkpoint_data = {
|
||||
'trajectory': [
|
||||
{
|
||||
'type': 'user.input',
|
||||
'message': {'role': 'user', 'content': 'hello'},
|
||||
'label': None,
|
||||
'ts': None,
|
||||
}
|
||||
],
|
||||
'tasks': [
|
||||
{
|
||||
'task_id': 'T1',
|
||||
'command': 'python analyze_logs.py',
|
||||
'rate': 50.0,
|
||||
'progress': 100.0,
|
||||
'status': 'completed',
|
||||
'result': None,
|
||||
'executable_receipt': None,
|
||||
}
|
||||
]
|
||||
}
|
||||
with open(target_path, "w", encoding="utf-8") as f:
|
||||
json.dump(checkpoint_data, f)
|
||||
|
||||
runtime = AgentRuntime.__new__(AgentRuntime)
|
||||
runtime.tasks = TaskManager(on_complete=MagicMock(), log=MagicMock())
|
||||
runtime.log = MagicMock()
|
||||
|
||||
runtime.load_checkpoint(target_path)
|
||||
|
||||
ev = runtime.trajectory[0]
|
||||
assert isinstance(ev.label, str)
|
||||
assert ev.label == ""
|
||||
assert isinstance(ev.ts, float)
|
||||
|
||||
st = runtime.tasks.query('T1')
|
||||
assert isinstance(st.result, str)
|
||||
assert st.result == ""
|
||||
assert isinstance(st.executable_receipt, dict)
|
||||
assert st.executable_receipt == {}
|
||||
@@ -0,0 +1,104 @@
|
||||
"""Acceptance-ledger regression tests for the durable Experiment 6-2 run."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import importlib.util
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(HERE))
|
||||
PATH = HERE / "run_real_experiment.py"
|
||||
SPEC = importlib.util.spec_from_file_location("experiment_6_2_real", PATH)
|
||||
runner = importlib.util.module_from_spec(SPEC)
|
||||
assert SPEC.loader is not None
|
||||
sys.modules[SPEC.name] = runner
|
||||
SPEC.loader.exec_module(runner)
|
||||
|
||||
|
||||
def _campaign() -> tuple[list[dict], dict]:
|
||||
root = HERE / "validation" / "experiment_6_2"
|
||||
campaigns = sorted(path for path in root.iterdir() if path.is_dir())
|
||||
assert campaigns, "a durable Experiment 6-2 campaign is required"
|
||||
campaign = campaigns[-1]
|
||||
scenarios = [json.loads(path.read_text(encoding="utf-8"))
|
||||
for path in sorted((campaign / "scenarios").glob("*.json"))]
|
||||
protocol = json.loads((campaign / "protocol.json").read_text(encoding="utf-8"))
|
||||
return scenarios, protocol
|
||||
|
||||
|
||||
def test_durable_campaign_passes_every_derived_gate():
|
||||
scenarios, protocol = _campaign()
|
||||
acceptance = runner.derive_acceptance(scenarios, protocol)
|
||||
assert acceptance["status"] == "passed"
|
||||
assert all(acceptance["gates"].values())
|
||||
|
||||
|
||||
def test_simulated_or_missing_process_receipt_cannot_pass():
|
||||
scenarios, protocol = _campaign()
|
||||
tampered = copy.deepcopy(scenarios)
|
||||
tampered[0]["tasks"][0]["executable"]["mode"] = "simulated"
|
||||
acceptance = runner.derive_acceptance(tampered, protocol)
|
||||
assert acceptance["status"] == "failed"
|
||||
assert not acceptance["gates"]["real_subprocess_receipts_only"]
|
||||
|
||||
|
||||
def test_empty_evidence_fails_closed():
|
||||
protocol = json.loads((HERE / "experiment_protocol.json").read_text(encoding="utf-8"))
|
||||
acceptance = runner.derive_acceptance([], protocol)
|
||||
assert acceptance["status"] == "failed"
|
||||
assert not any(acceptance["gates"].values())
|
||||
|
||||
|
||||
|
||||
def test_protocol_coverage_mapping_is_complete_and_enforced():
|
||||
"""Every protocol acceptance key must map to at least one gate, and the
|
||||
coverage report must reflect gate pass/fail status correctly."""
|
||||
scenarios, protocol = _campaign()
|
||||
acceptance = runner.derive_acceptance(scenarios, protocol)
|
||||
coverage = acceptance["protocol_coverage"]
|
||||
# Every protocol acceptance key must appear in the coverage report.
|
||||
protocol_keys = set(protocol.get("acceptance", {}))
|
||||
assert set(coverage) == protocol_keys, (
|
||||
f"coverage keys {set(coverage)} != protocol keys {protocol_keys}"
|
||||
)
|
||||
# Every coverage entry must reference at least one gate key.
|
||||
for proto_key, entry in coverage.items():
|
||||
assert len(entry["enforced_by"]) >= 1, f"{proto_key} has no enforcing gate"
|
||||
# When all gates pass, every coverage entry must report all_gates_passed=True.
|
||||
if acceptance["status"] == "passed":
|
||||
assert all(entry["all_gates_passed"] for entry in coverage.values())
|
||||
|
||||
|
||||
def test_protocol_coverage_detects_unmapped_acceptance_key():
|
||||
"""Adding an acceptance key to the protocol without a PROTOCOL_TO_GATE
|
||||
mapping must raise an assertion at run time."""
|
||||
scenarios, protocol = _campaign()
|
||||
tampered_protocol = copy.deepcopy(protocol)
|
||||
tampered_protocol["acceptance"]["bogus_unmapped_key"] = "must be enforced"
|
||||
try:
|
||||
runner.derive_acceptance(scenarios, tampered_protocol)
|
||||
except AssertionError as exc:
|
||||
assert "bogus_unmapped_key" in str(exc)
|
||||
else:
|
||||
raise AssertionError("expected AssertionError for unmapped acceptance key")
|
||||
|
||||
|
||||
def test_protocol_coverage_reflects_gate_failure():
|
||||
"""When a gate fails, the coverage entries that depend on it must report
|
||||
all_gates_passed=False."""
|
||||
scenarios, protocol = _campaign()
|
||||
tampered = copy.deepcopy(scenarios)
|
||||
tampered[0]["tasks"][0]["executable"]["mode"] = "simulated"
|
||||
acceptance = runner.derive_acceptance(tampered, protocol)
|
||||
coverage = acceptance["protocol_coverage"]
|
||||
# real_subprocess_receipts_only gate should have failed.
|
||||
assert not acceptance["gates"]["real_subprocess_receipts_only"]
|
||||
# Every protocol key enforced by that gate must report failure.
|
||||
for proto_key, entry in coverage.items():
|
||||
if "real_subprocess_receipts_only" in entry["enforced_by"]:
|
||||
assert not entry["all_gates_passed"], (
|
||||
f"{proto_key} should report gate failure"
|
||||
)
|
||||
@@ -0,0 +1,82 @@
|
||||
"""Contract tests for real subprocess-backed Experiment 6-2 tasks."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(HERE))
|
||||
|
||||
import tasks
|
||||
|
||||
|
||||
def test_unapproved_commands_are_rejected_before_execution():
|
||||
with pytest.raises(ValueError, match="unapproved"):
|
||||
tasks.resolve_job("python arbitrary.py")
|
||||
with pytest.raises(ValueError, match="only"):
|
||||
tasks.resolve_job("sh -c 'echo unsafe'")
|
||||
|
||||
|
||||
def test_real_subprocess_completes_with_observed_metrics(monkeypatch):
|
||||
async def scenario():
|
||||
monkeypatch.setattr(tasks, "TICK_REAL", 0.002)
|
||||
completed = []
|
||||
|
||||
async def on_complete(state):
|
||||
completed.append(state.task_id)
|
||||
|
||||
manager = tasks.TaskManager(on_complete, lambda *_: None)
|
||||
state = manager.start("python analyze_fast.py")
|
||||
assert state._task is not None
|
||||
await state._task
|
||||
result = json.loads(state.result)
|
||||
assert completed == [state.task_id]
|
||||
assert state.status == "completed"
|
||||
assert state.pid and state.pid != os.getpid()
|
||||
assert state.returncode == 0
|
||||
assert state.progress == 100
|
||||
assert state.stdout_sha256 and len(state.stdout_sha256) == 64
|
||||
assert state.executable_receipt["mode"] == "real_subprocess"
|
||||
assert state.executable_receipt["shell"] is False
|
||||
assert state.executable_receipt["returncode"] == 0
|
||||
assert result["input_sha256"] == hashlib.sha256(
|
||||
tasks.DEFAULT_INPUT.read_bytes()
|
||||
).hexdigest()
|
||||
assert result["bytes"] == tasks.DEFAULT_INPUT.stat().st_size
|
||||
assert result["lines"] > 100
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
def test_cancel_terminates_real_child_process_and_freezes_progress(monkeypatch):
|
||||
async def scenario():
|
||||
monkeypatch.setattr(tasks, "TICK_REAL", 0.02)
|
||||
|
||||
async def on_complete(_):
|
||||
raise AssertionError("cancelled process must not complete")
|
||||
|
||||
manager = tasks.TaskManager(on_complete, lambda *_: None)
|
||||
state = manager.start("python analyze_slow.py")
|
||||
while state.pid is None:
|
||||
await asyncio.sleep(0.001)
|
||||
await asyncio.sleep(0.05)
|
||||
pid = state.pid
|
||||
assert manager.cancel(state.task_id)
|
||||
with pytest.raises(asyncio.CancelledError):
|
||||
await state._task
|
||||
frozen = state.progress
|
||||
await asyncio.sleep(0.05)
|
||||
assert state.status == "cancelled"
|
||||
assert state.progress == frozen < 100
|
||||
assert state.executable_receipt["cancelled"] is True
|
||||
assert state.returncode is not None
|
||||
with pytest.raises(ProcessLookupError):
|
||||
os.kill(pid, 0)
|
||||
|
||||
asyncio.run(scenario())
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Test import bootstrap for the async-agent experiment."""
|
||||
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
|
||||
EXPERIMENT_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(EXPERIMENT_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(EXPERIMENT_ROOT))
|
||||
@@ -0,0 +1,35 @@
|
||||
"""回归测试:FLUX_TICK_REAL 等浮点环境变量非法时不得让模块导入崩溃。
|
||||
|
||||
tasks.py 原来在模块导入时用裸 float() 解析 FLUX_TICK_REAL,
|
||||
FLUX_TICK_REAL=abc 会让整个演示脚本以 ValueError 崩溃;现在回退到默认值并打印警告。
|
||||
"""
|
||||
import importlib
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
|
||||
import tasks
|
||||
|
||||
|
||||
def test_env_float_falls_back_on_malformed(monkeypatch, capsys):
|
||||
monkeypatch.setenv("FLUX_TICK_REAL", "abc")
|
||||
assert tasks._env_float("FLUX_TICK_REAL", 0.4) == 0.4
|
||||
assert "FLUX_TICK_REAL" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_env_float_parses_valid_value(monkeypatch):
|
||||
monkeypatch.setenv("FLUX_TICK_REAL", "0.1")
|
||||
assert tasks._env_float("FLUX_TICK_REAL", 0.4) == 0.1
|
||||
|
||||
|
||||
def test_env_float_default_when_unset(monkeypatch):
|
||||
monkeypatch.delenv("FLUX_TICK_REAL", raising=False)
|
||||
assert tasks._env_float("FLUX_TICK_REAL", 0.4) == 0.4
|
||||
|
||||
|
||||
def test_module_reload_survives_malformed_env(monkeypatch):
|
||||
"""模块级 TICK_REAL 在环境变量非法时不得抛出 ValueError。"""
|
||||
monkeypatch.setenv("FLUX_TICK_REAL", "fast")
|
||||
importlib.reload(tasks)
|
||||
assert tasks.TICK_REAL == 0.4
|
||||
@@ -0,0 +1 @@
|
||||
<!DOCTYPE html><html lang="ja"><meta charset="utf-8"><title>非同期分析レポート</title><body><h1>分析結果</h1><p>対象ファイルは 679 行、106193 バイトです。</p><p>非同期に関する言及は 69 件でした。</p></body></html>
|
||||
@@ -0,0 +1,44 @@
|
||||
{
|
||||
"first_completed": "T1",
|
||||
"query_receipts": [
|
||||
{
|
||||
"task_id": "T2",
|
||||
"status": "running",
|
||||
"progress": 68.0,
|
||||
"queried_at": 5.345
|
||||
},
|
||||
{
|
||||
"task_id": "T3",
|
||||
"status": "running",
|
||||
"progress": 34.0,
|
||||
"queried_at": 5.345
|
||||
}
|
||||
],
|
||||
"cancelled_ids": [
|
||||
"T3"
|
||||
],
|
||||
"completed_results": {
|
||||
"T1": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "fast",
|
||||
"lines": 679
|
||||
},
|
||||
"T2": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "mid",
|
||||
"lines": 679
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,45 @@
|
||||
{
|
||||
"generated_at": "2026-07-29T21:28:45.233367+00:00",
|
||||
"files": [
|
||||
{
|
||||
"path": "artifacts/scenario_2_report.html",
|
||||
"bytes": 257,
|
||||
"sha256": "e468c5a2c05f5fe74b2c41ca123676ab2f325fda41869a615545221b7d1c7a2f"
|
||||
},
|
||||
{
|
||||
"path": "artifacts/scenario_4_report.json",
|
||||
"bytes": 1071,
|
||||
"sha256": "7bc8f53d28a8ec058424beb24ee1fcf9e36dd6c8b01e16fbea8dcc8d5b5dead7"
|
||||
},
|
||||
{
|
||||
"path": "protocol.json",
|
||||
"bytes": 1130,
|
||||
"sha256": "2d15d0bfd8314d4ae141dfb72539703fdcd9fa6aed9bd34082a54df6694b42d3"
|
||||
},
|
||||
{
|
||||
"path": "scenarios/async_command_and_immediate_question.json",
|
||||
"bytes": 3438,
|
||||
"sha256": "24677cf6db74cfefee1f6a856254ee6b4d7bbde2de03f3015966bb86aa473f8d"
|
||||
},
|
||||
{
|
||||
"path": "scenarios/interrupt_terminates_and_recovers.json",
|
||||
"bytes": 4995,
|
||||
"sha256": "d355ea15acc9b85d789460a215f7cd36b4ee9549194b87a328515f17de8bbcac"
|
||||
},
|
||||
{
|
||||
"path": "scenarios/parallel_progress_threshold_cancellation.json",
|
||||
"bytes": 9035,
|
||||
"sha256": "3e6ec09e98703ef4309d55d2386700de8f6ea8c7daeda660ee49ee14c8bbedf4"
|
||||
},
|
||||
{
|
||||
"path": "scenarios/queued_batch_to_japanese_html.json",
|
||||
"bytes": 4089,
|
||||
"sha256": "796bf2d2878a6681d897d42938836f29e301d0c4fe6e6fbb78975427dae4aca9"
|
||||
},
|
||||
{
|
||||
"path": "summary.json",
|
||||
"bytes": 1097,
|
||||
"sha256": "913b5b9b73d61bcf02c51a44c5ed204dbb6ca5abbe8f133503fd88690ae5f788"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"experiment": "6-2",
|
||||
"title": "Real subprocess async execution, queueing, interruption, and progress cancellation",
|
||||
"authority": "book/chapter6.md:579",
|
||||
"execution": {
|
||||
"mode": "allowlisted asyncio subprocesses",
|
||||
"shell": false,
|
||||
"worker": "analysis_worker.py",
|
||||
"input": "book/chapter6.md",
|
||||
"progress_source": "child stdout",
|
||||
"result_source": "child-computed file metrics"
|
||||
},
|
||||
"scenarios": [
|
||||
"long command plus immediate current-time response before completion",
|
||||
"two deferred instructions batch on completion and produce Japanese HTML",
|
||||
"user cancellation terminates the child process and runtime recovers",
|
||||
"3/2/1 percent jobs; query remaining jobs once; cancel only progress at or below 50 percent"
|
||||
],
|
||||
"acceptance": {
|
||||
"long_job_at_least_three_seconds": true,
|
||||
"placeholder_return_is_nonblocking": true,
|
||||
"all_terminal_jobs_are_real_subprocesses": true,
|
||||
"cancelled_jobs_have_os_return_codes": true,
|
||||
"completed_jobs_have_stdout_and_input_hashes": true,
|
||||
"all_artifacts_are_hash_manifested": true,
|
||||
"no_simulated_terminal_result_can_pass": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,109 @@
|
||||
{
|
||||
"id": "async_command_and_immediate_question",
|
||||
"placeholder_latency_seconds": 0.000241,
|
||||
"placeholder_at": 0.0,
|
||||
"time_answer_at": 0.502,
|
||||
"completion_event_at": 3.635,
|
||||
"time_answer": "2026-07-30T05:28:26+08:00",
|
||||
"events": [
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T1: `python analyze_logs.py` (速度 4%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TOOL",
|
||||
"event": "placeholder_returned",
|
||||
"task_id": "T1",
|
||||
"placeholder_latency_seconds": 0.000241
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.502,
|
||||
"source": "AGENT",
|
||||
"event": "immediate_time_answer",
|
||||
"answer": "2026-07-30T05:28:26+08:00",
|
||||
"task_still_running": true
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.803,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 22% (pid=92016)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 1.426,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 40% (pid=92016)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.207,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 63% (pid=92016)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.84,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 81% (pid=92016)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.622,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 100% (pid=92016)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.635,
|
||||
"source": "TASK",
|
||||
"text": "T1 完成 ✅ (pid=92016, returncode=0)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.635,
|
||||
"source": "SYSTEM",
|
||||
"event": "async_result_injected",
|
||||
"task_id": "T1"
|
||||
}
|
||||
],
|
||||
"tasks": [
|
||||
{
|
||||
"task_id": "T1",
|
||||
"command": "python analyze_logs.py",
|
||||
"status": "completed",
|
||||
"progress": 100.0,
|
||||
"pid": 92016,
|
||||
"returncode": 0,
|
||||
"started_at": 1785360505.728322,
|
||||
"completed_at": 1785360509.363023,
|
||||
"elapsed_seconds": 3.635,
|
||||
"stdout_sha256": "7da29da059869ca8f939dd6d0d4c4a299d9e02a53592773c71517971af545ebd",
|
||||
"stderr_tail": "",
|
||||
"result": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "logs",
|
||||
"lines": 679
|
||||
},
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "logs",
|
||||
"rate_percent_per_logical_second": 4.5,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "1c6aa3b4072446e8286dce865f45c6bb1c93733319bf7501adbb555bc6a7973c",
|
||||
"pid": 92016,
|
||||
"returncode": 0,
|
||||
"stdout_sha256": "7da29da059869ca8f939dd6d0d4c4a299d9e02a53592773c71517971af545ebd",
|
||||
"stdout_lines": 24,
|
||||
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||||
"elapsed_seconds": 3.635
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,161 @@
|
||||
{
|
||||
"id": "interrupt_terminates_and_recovers",
|
||||
"interrupt_at": 0.803,
|
||||
"cancel_receipt_at": 0.806,
|
||||
"cancel_latency_seconds": 0.004,
|
||||
"recovery_at": 0.807,
|
||||
"completed_callbacks": [
|
||||
"T2"
|
||||
],
|
||||
"events": [
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T1: `python analyze_logs.py` (速度 4%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.803,
|
||||
"source": "USER",
|
||||
"event": "user.interrupt",
|
||||
"text": "取消",
|
||||
"task_id": "T1",
|
||||
"progress": 18.0
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.806,
|
||||
"source": "TASK",
|
||||
"text": "T1 子进程已终止 🛑 (pid=92040, 进度 18%)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.806,
|
||||
"source": "SYSTEM",
|
||||
"event": "process_cancelled",
|
||||
"task_ids": [
|
||||
"T1"
|
||||
],
|
||||
"cancel_latency_seconds": 0.004,
|
||||
"returncode": -15
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.807,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T2: `python re_run_summary.py` (速度 4%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.807,
|
||||
"source": "SYSTEM",
|
||||
"event": "runtime_recovered",
|
||||
"task_id": "T2"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 1.619,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python re_run_summary.py` 进度 22% (pid=92041)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.245,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python re_run_summary.py` 进度 40% (pid=92041)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.019,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python re_run_summary.py` 进度 63% (pid=92041)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.644,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python re_run_summary.py` 进度 81% (pid=92041)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 4.423,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python re_run_summary.py` 进度 100% (pid=92041)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 4.436,
|
||||
"source": "TASK",
|
||||
"text": "T2 完成 ✅ (pid=92041, returncode=0)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 4.436,
|
||||
"source": "SYSTEM",
|
||||
"event": "async_result_injected",
|
||||
"task_id": "T2"
|
||||
}
|
||||
],
|
||||
"tasks": [
|
||||
{
|
||||
"task_id": "T1",
|
||||
"command": "python analyze_logs.py",
|
||||
"status": "cancelled",
|
||||
"progress": 18.0,
|
||||
"pid": 92040,
|
||||
"returncode": -15,
|
||||
"started_at": 1785360512.964798,
|
||||
"completed_at": 1785360513.7706082,
|
||||
"elapsed_seconds": 0.806,
|
||||
"stdout_sha256": null,
|
||||
"stderr_tail": "",
|
||||
"result": null,
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "logs",
|
||||
"rate_percent_per_logical_second": 4.5,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "1c6aa3b4072446e8286dce865f45c6bb1c93733319bf7501adbb555bc6a7973c",
|
||||
"pid": 92040,
|
||||
"returncode": -15,
|
||||
"cancelled": true,
|
||||
"elapsed_seconds": 0.806
|
||||
}
|
||||
},
|
||||
{
|
||||
"task_id": "T2",
|
||||
"command": "python re_run_summary.py",
|
||||
"status": "completed",
|
||||
"progress": 100.0,
|
||||
"pid": 92041,
|
||||
"returncode": 0,
|
||||
"started_at": 1785360513.771501,
|
||||
"completed_at": 1785360517.4005382,
|
||||
"elapsed_seconds": 3.629,
|
||||
"stdout_sha256": "5c538eaf31540bcc95acb80a9ca3333a440020c63e015b644575a35cabbe6a11",
|
||||
"stderr_tail": "",
|
||||
"result": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "recovery",
|
||||
"lines": 679
|
||||
},
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "recovery",
|
||||
"rate_percent_per_logical_second": 4.5,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "7079826b54049bc6e5276ade7f4c343a41b9fc82f828d35685665b3101bce693",
|
||||
"pid": 92041,
|
||||
"returncode": 0,
|
||||
"stdout_sha256": "5c538eaf31540bcc95acb80a9ca3333a440020c63e015b644575a35cabbe6a11",
|
||||
"stdout_lines": 24,
|
||||
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||||
"elapsed_seconds": 3.629
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,286 @@
|
||||
{
|
||||
"id": "parallel_progress_threshold_cancellation",
|
||||
"first_completed": "T1",
|
||||
"query_receipts": [
|
||||
{
|
||||
"task_id": "T2",
|
||||
"status": "running",
|
||||
"progress": 68.0,
|
||||
"queried_at": 5.345
|
||||
},
|
||||
{
|
||||
"task_id": "T3",
|
||||
"status": "running",
|
||||
"progress": 34.0,
|
||||
"queried_at": 5.345
|
||||
}
|
||||
],
|
||||
"cancelled_ids": [
|
||||
"T3"
|
||||
],
|
||||
"completed_results": {
|
||||
"T1": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "fast",
|
||||
"lines": 679
|
||||
},
|
||||
"T2": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "mid",
|
||||
"lines": 679
|
||||
}
|
||||
},
|
||||
"events": [
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T1: `python analyze_fast.py` (速度 3%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.001,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T2: `python analyze_mid.py` (速度 2%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.001,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T3: `python analyze_slow.py` (速度 1%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 1.121,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_fast.py` 进度 21% (pid=92132)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 1.605,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python analyze_mid.py` 进度 20% (pid=92133)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.219,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_fast.py` 进度 42% (pid=92132)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.133,
|
||||
"source": "TASK",
|
||||
"text": "T3 `python analyze_slow.py` 进度 20% (pid=92134)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.144,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_fast.py` 进度 60% (pid=92132)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.145,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python analyze_mid.py` 进度 40% (pid=92133)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 4.235,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_fast.py` 进度 81% (pid=92132)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 4.699,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python analyze_mid.py` 进度 60% (pid=92133)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.332,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_fast.py` 进度 100% (pid=92132)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.345,
|
||||
"source": "TASK",
|
||||
"text": "T1 完成 ✅ (pid=92132, returncode=0)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.345,
|
||||
"source": "SYSTEM",
|
||||
"event": "async_result_injected",
|
||||
"task_id": "T1"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.345,
|
||||
"source": "TOOL",
|
||||
"event": "query_task",
|
||||
"task_id": "T2",
|
||||
"progress": 68.0
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.345,
|
||||
"source": "TOOL",
|
||||
"event": "query_task",
|
||||
"task_id": "T3",
|
||||
"progress": 34.0
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.345,
|
||||
"source": "TOOL",
|
||||
"event": "cancel_task",
|
||||
"task_id": "T3",
|
||||
"progress": 34.0
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 5.347,
|
||||
"source": "TASK",
|
||||
"text": "T3 子进程已终止 🛑 (pid=92134, 进度 34%)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 6.254,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python analyze_mid.py` 进度 80% (pid=92133)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 7.815,
|
||||
"source": "TASK",
|
||||
"text": "T2 `python analyze_mid.py` 进度 100% (pid=92133)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 7.829,
|
||||
"source": "TASK",
|
||||
"text": "T2 完成 ✅ (pid=92133, returncode=0)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 7.829,
|
||||
"source": "SYSTEM",
|
||||
"event": "async_result_injected",
|
||||
"task_id": "T2"
|
||||
}
|
||||
],
|
||||
"tasks": [
|
||||
{
|
||||
"task_id": "T1",
|
||||
"command": "python analyze_fast.py",
|
||||
"status": "completed",
|
||||
"progress": 100.0,
|
||||
"pid": 92132,
|
||||
"returncode": 0,
|
||||
"started_at": 1785360517.401463,
|
||||
"completed_at": 1785360522.7459211,
|
||||
"elapsed_seconds": 5.344,
|
||||
"stdout_sha256": "cc880975410f2730b3a60c08b446abc561361940b173aa8241ab22f3d23be937",
|
||||
"stderr_tail": "",
|
||||
"result": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "fast",
|
||||
"lines": 679
|
||||
},
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "fast",
|
||||
"rate_percent_per_logical_second": 3.0,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "7261f0de1acd2bbe527657a61c48b925528ac99524a349ac91d350581830780e",
|
||||
"pid": 92132,
|
||||
"returncode": 0,
|
||||
"stdout_sha256": "cc880975410f2730b3a60c08b446abc561361940b173aa8241ab22f3d23be937",
|
||||
"stdout_lines": 35,
|
||||
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||||
"elapsed_seconds": 5.344
|
||||
}
|
||||
},
|
||||
{
|
||||
"task_id": "T2",
|
||||
"command": "python analyze_mid.py",
|
||||
"status": "completed",
|
||||
"progress": 100.0,
|
||||
"pid": 92133,
|
||||
"returncode": 0,
|
||||
"started_at": 1785360517.4036,
|
||||
"completed_at": 1785360525.229522,
|
||||
"elapsed_seconds": 7.826,
|
||||
"stdout_sha256": "a4257b9e774792d2df2f286f29d683ac24d90a766da66069fc5c4654e89ed003",
|
||||
"stderr_tail": "",
|
||||
"result": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "mid",
|
||||
"lines": 679
|
||||
},
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "mid",
|
||||
"rate_percent_per_logical_second": 2.0,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "3f8854dc260b11c7d00cbf44f3770cbe6eaf45da0ca6677108a485acd69faf4f",
|
||||
"pid": 92133,
|
||||
"returncode": 0,
|
||||
"stdout_sha256": "a4257b9e774792d2df2f286f29d683ac24d90a766da66069fc5c4654e89ed003",
|
||||
"stdout_lines": 51,
|
||||
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||||
"elapsed_seconds": 7.826
|
||||
}
|
||||
},
|
||||
{
|
||||
"task_id": "T3",
|
||||
"command": "python analyze_slow.py",
|
||||
"status": "cancelled",
|
||||
"progress": 34.0,
|
||||
"pid": 92134,
|
||||
"returncode": -15,
|
||||
"started_at": 1785360517.4054961,
|
||||
"completed_at": 1785360522.74722,
|
||||
"elapsed_seconds": 5.342,
|
||||
"stdout_sha256": null,
|
||||
"stderr_tail": "",
|
||||
"result": null,
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "slow",
|
||||
"rate_percent_per_logical_second": 1.0,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "45f0bfa2cdbd6f54216af52a0d35e71b53de6c9a67f2ddbee1b56594cb1a4f67",
|
||||
"pid": 92134,
|
||||
"returncode": -15,
|
||||
"cancelled": true,
|
||||
"elapsed_seconds": 5.342
|
||||
}
|
||||
}
|
||||
],
|
||||
"artifact": {
|
||||
"path": "/Users/boj/book/ai-agent-book/chapter6/async-agent/validation/experiment_6_2/real_subprocess_20260730T052500Z/artifacts/scenario_4_report.json",
|
||||
"bytes": 1071,
|
||||
"sha256": "7bc8f53d28a8ec058424beb24ee1fcf9e36dd6c8b01e16fbea8dcc8d5b5dead7"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,136 @@
|
||||
{
|
||||
"id": "queued_batch_to_japanese_html",
|
||||
"instruction_times": [
|
||||
0.503,
|
||||
0.703
|
||||
],
|
||||
"batch": [
|
||||
{
|
||||
"type": "async.result",
|
||||
"task_id": "T1",
|
||||
"result_sha256": "8cc819c59b166b1f24f678fffc632ea401136984f3616c2b98c9278ea85a570f"
|
||||
},
|
||||
{
|
||||
"type": "user.input",
|
||||
"instruction": "記得最後用日語回覆"
|
||||
},
|
||||
{
|
||||
"type": "user.input",
|
||||
"instruction": "結果をHTMLウェブページに整理"
|
||||
}
|
||||
],
|
||||
"events": [
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TASK",
|
||||
"text": "启动真实子进程任务 T1: `python analyze_logs.py` (速度 4%/逻辑秒)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.0,
|
||||
"source": "TOOL",
|
||||
"event": "placeholder_returned",
|
||||
"task_id": "T1"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.503,
|
||||
"source": "QUEUE",
|
||||
"event": "deferred_instruction",
|
||||
"instruction": "japanese"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.703,
|
||||
"source": "QUEUE",
|
||||
"event": "deferred_instruction",
|
||||
"instruction": "html"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 0.796,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 22% (pid=92027)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 1.42,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 40% (pid=92027)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.198,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 63% (pid=92027)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 2.817,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 81% (pid=92027)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.592,
|
||||
"source": "TASK",
|
||||
"text": "T1 `python analyze_logs.py` 进度 100% (pid=92027)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.599,
|
||||
"source": "TASK",
|
||||
"text": "T1 完成 ✅ (pid=92027, returncode=0)"
|
||||
},
|
||||
{
|
||||
"elapsed_seconds": 3.599,
|
||||
"source": "SYSTEM",
|
||||
"event": "batch_appended",
|
||||
"event_count": 3,
|
||||
"deferred_count": 2
|
||||
}
|
||||
],
|
||||
"artifact": {
|
||||
"path": "/Users/boj/book/ai-agent-book/chapter6/async-agent/validation/experiment_6_2/real_subprocess_20260730T052500Z/artifacts/scenario_2_report.html",
|
||||
"bytes": 257,
|
||||
"sha256": "e468c5a2c05f5fe74b2c41ca123676ab2f325fda41869a615545221b7d1c7a2f",
|
||||
"doctype": true,
|
||||
"lang_ja": true,
|
||||
"has_japanese": true
|
||||
},
|
||||
"tasks": [
|
||||
{
|
||||
"task_id": "T1",
|
||||
"command": "python analyze_logs.py",
|
||||
"status": "completed",
|
||||
"progress": 100.0,
|
||||
"pid": 92027,
|
||||
"returncode": 0,
|
||||
"started_at": 1785360509.363528,
|
||||
"completed_at": 1785360512.962622,
|
||||
"elapsed_seconds": 3.599,
|
||||
"stdout_sha256": "7da29da059869ca8f939dd6d0d4c4a299d9e02a53592773c71517971af545ebd",
|
||||
"stderr_tail": "",
|
||||
"result": {
|
||||
"async_mentions": 69,
|
||||
"bytes": 106193,
|
||||
"error_keyword_count": 20,
|
||||
"experiment_mentions": 37,
|
||||
"heading_count": 27,
|
||||
"input_path": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "logs",
|
||||
"lines": 679
|
||||
},
|
||||
"executable": {
|
||||
"mode": "real_subprocess",
|
||||
"shell": false,
|
||||
"worker": "/Users/boj/book/ai-agent-book/chapter6/async-agent/analysis_worker.py",
|
||||
"worker_sha256": "ab700e034f63e4b8f83dfb4a6156745b7ab7c9352750c63a900f2632acb134c7",
|
||||
"input": "/Users/boj/book/ai-agent-book/book/chapter6.md",
|
||||
"input_sha256": "4184460334544d326bb7ad4db349b8c583790195d51efc28ef325284c41d7fdd",
|
||||
"job": "logs",
|
||||
"rate_percent_per_logical_second": 4.5,
|
||||
"tick_real_seconds": 0.15,
|
||||
"argv_sha256": "1c6aa3b4072446e8286dce865f45c6bb1c93733319bf7501adbb555bc6a7973c",
|
||||
"pid": 92027,
|
||||
"returncode": 0,
|
||||
"stdout_sha256": "7da29da059869ca8f939dd6d0d4c4a299d9e02a53592773c71517971af545ebd",
|
||||
"stdout_lines": 24,
|
||||
"stderr_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
||||
"elapsed_seconds": 3.599
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"experiment": "6-2",
|
||||
"campaign_id": "real_subprocess_20260730T052500Z",
|
||||
"generated_at": "2026-07-29T21:28:45.232118+00:00",
|
||||
"tick_real_seconds": 0.15,
|
||||
"elapsed_seconds": 19.504,
|
||||
"scenario_status": {
|
||||
"async_command_and_immediate_question": "recorded",
|
||||
"queued_batch_to_japanese_html": "recorded",
|
||||
"interrupt_terminates_and_recovers": "recorded",
|
||||
"parallel_progress_threshold_cancellation": "recorded"
|
||||
},
|
||||
"acceptance": {
|
||||
"status": "passed",
|
||||
"gates": {
|
||||
"exact_four_scenarios": true,
|
||||
"real_subprocess_receipts_only": true,
|
||||
"scenario_1_nonblocking_and_immediate_response": true,
|
||||
"scenario_2_deferred_events_batched_once": true,
|
||||
"scenario_2_japanese_html_artifact": true,
|
||||
"scenario_3_os_process_cancelled_then_recovered": true,
|
||||
"scenario_4_exact_rates_and_fast_first": true,
|
||||
"scenario_4_query_once_and_cancel_only_under_threshold": true,
|
||||
"scenario_4_integrated_report_hashed": true
|
||||
},
|
||||
"protocol_sha256": "536c1dc0a187b8795c92b2c11babbf37c54b29215a69226221722567ba867fca"
|
||||
},
|
||||
"status": "passed"
|
||||
}
|
||||
@@ -0,0 +1,69 @@
|
||||
# Experiment 6-7: Anthropic native Computer Use
|
||||
|
||||
This record covers the provider-specific arm of Experiment 6-7: Anthropic's
|
||||
native tool protocol in the official containerized Computer Use Demo. It is
|
||||
separate from the completed open-model Experiment 6-8 arm. The runner, validator,
|
||||
and retained evidence directories consistently use the `exp6-7-*` identifier.
|
||||
|
||||
Current status: **complete for the bounded read-only task**. The canonical
|
||||
[trajectory](validation/runs/exp6-7-anthropic-native-20260803-v2/trajectory.json)
|
||||
and [deterministic acceptance](validation/runs/exp6-7-anthropic-native-20260803-v2/acceptance.json)
|
||||
retain a real run of the required task:
|
||||
|
||||
> Open Google, search for San Francisco weather today, and report the
|
||||
> temperature and conditions. Do not sign in or change any external data.
|
||||
|
||||
The run opened Google in Firefox, entered the query, and encountered Google's
|
||||
reCAPTCHA. It did not click or otherwise interact with the challenge. Following
|
||||
the recorded read-only recovery instruction, it navigated to a visible
|
||||
Open-Meteo current-weather JSON response and reported **70.2°F, clear sky**
|
||||
(`weather_code: 0`) for San Francisco. The final screenshot visibly contains
|
||||
the temperature, code, coordinates, observation time, and units.
|
||||
|
||||
## Provenance and result
|
||||
|
||||
- Upstream source: `anthropics/claude-quickstarts` at
|
||||
`9bcc95e316e5ef6542b4c9d0469f4078829eead5`.
|
||||
- Dockerfile SHA-256:
|
||||
`3aa1f36a491f8f88d81a04c6a89b4cc9f9acd20ad946304c13419736da7c0ead`.
|
||||
- Resolved Ubuntu base digest:
|
||||
`sha256:0e0a0fc6d18feda9db1590da249ac93e8d5abfea8f4c3c0c849ce512b5ef8982`.
|
||||
- Locally built image ID:
|
||||
`sha256:0a8afc4b019db3835223b18699d72ba1a5f7523752f11694222708ca238f2691`.
|
||||
The mutable prebuilt `computer-use-demo-latest` image was not used.
|
||||
- Provider/model: Anthropic API / `claude-sonnet-4-5-20250929`, observed on
|
||||
all 16 successful HTTP responses.
|
||||
- Native tool version: `computer_use_20250124`.
|
||||
- Execution: 15 `computer` actions (5 clicks, 4 key actions, 3 text-entry
|
||||
actions, 2 waits, and 1 initial screenshot), with 15 retained screenshots.
|
||||
- Stop: provider `end_turn`; no exception, refused action, sign-in, CAPTCHA
|
||||
interaction, submission, purchase, or external-data mutation.
|
||||
- Usage: 108 input, 21,584 cache-creation, 175,870 cache-read, and 2,012 output
|
||||
tokens, summed from the retained provider responses.
|
||||
|
||||
The [manifest](validation/runs/exp6-7-anthropic-native-20260803-v2/manifest.json)
|
||||
hashes every canonical artifact. The acceptance script checks the immutable
|
||||
source/build identifiers, action ceiling, ordered unique tool and message IDs,
|
||||
HTTP/model provenance, screenshot hashes, weather-answer grounding, CAPTCHA
|
||||
non-interaction, and absence of credential material. All gates pass:
|
||||
|
||||
```bash
|
||||
python3 chapter6/claude-computer-use-native/validate_weather_run.py \
|
||||
chapter6/claude-computer-use-native/validation/runs/exp6-7-anthropic-native-20260803-v2
|
||||
```
|
||||
|
||||
## Retained failed attempts
|
||||
|
||||
The historical 401 [preflight](validation/exp6-7-anthropic-auth-20260803-v1/preflight.json)
|
||||
is retained rather than rewritten. Two subsequent real task attempts are also
|
||||
retained under `validation/failed_attempts/`:
|
||||
|
||||
1. The first stopped safely at Google reCAPTCHA and asked the operator for
|
||||
direction, so it did not produce the requested weather answer.
|
||||
2. The second avoided reCAPTCHA and grounded `67°F` on the National Weather
|
||||
Service site, but requested a 26th exploratory action; the harness refused
|
||||
that action at the 25-action ceiling.
|
||||
|
||||
These failures are not counted as the canonical result. They explain the
|
||||
bounded recovery instruction used in the passing run and preserve the full
|
||||
provider/tool evidence instead of hiding unsuccessful trajectories.
|
||||
@@ -0,0 +1,329 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run and retain the bounded Experiment 6-7 native Computer Use trajectory.
|
||||
|
||||
This harness calls the pinned Anthropic Computer Use Demo's ``sampling_loop``.
|
||||
It is intended to run inside that Demo's locally built container with a host
|
||||
evidence directory mounted at ``/evidence``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import platform
|
||||
import sys
|
||||
import traceback
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from computer_use_demo.loop import APIProvider, sampling_loop
|
||||
from computer_use_demo.tools import ToolResult
|
||||
|
||||
|
||||
TASK = (
|
||||
"Open Google, search for San Francisco weather today, and report the "
|
||||
"temperature and conditions. Do not sign in or change any external data."
|
||||
)
|
||||
MODEL = "claude-sonnet-4-5-20250929"
|
||||
TOOL_VERSION = "computer_use_20250124"
|
||||
ACTION_LIMIT = 25
|
||||
OUT = Path(os.environ.get("EXP96_EVIDENCE_DIR", "/evidence"))
|
||||
|
||||
|
||||
class ActionLimitReached(RuntimeError):
|
||||
"""Raised before an action beyond the experiment ceiling executes."""
|
||||
|
||||
|
||||
def utc_now() -> str:
|
||||
return datetime.now(UTC).isoformat().replace("+00:00", "Z")
|
||||
|
||||
|
||||
def sha256_bytes(value: bytes) -> str:
|
||||
return hashlib.sha256(value).hexdigest()
|
||||
|
||||
|
||||
def json_safe(value: Any) -> Any:
|
||||
if value is None or isinstance(value, (bool, int, float, str)):
|
||||
return value
|
||||
if isinstance(value, dict):
|
||||
return {str(k): json_safe(v) for k, v in value.items()}
|
||||
if isinstance(value, (list, tuple)):
|
||||
return [json_safe(v) for v in value]
|
||||
if hasattr(value, "model_dump"):
|
||||
return json_safe(value.model_dump())
|
||||
return repr(value)
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
if not os.environ.get("ANTHROPIC_API_KEY"):
|
||||
raise RuntimeError("ANTHROPIC_API_KEY is not set")
|
||||
|
||||
OUT.mkdir(parents=True, exist_ok=True)
|
||||
screenshots = OUT / "screenshots"
|
||||
receipts = OUT / "api_receipts"
|
||||
screenshots.mkdir(exist_ok=True)
|
||||
receipts.mkdir(exist_ok=True)
|
||||
|
||||
started_at = utc_now()
|
||||
api_calls: list[dict[str, Any]] = []
|
||||
actions: list[dict[str, Any]] = []
|
||||
action_by_id: dict[str, dict[str, Any]] = {}
|
||||
refused_action: dict[str, Any] | None = None
|
||||
messages: list[dict[str, Any]] = [
|
||||
{"role": "user", "content": [{"type": "text", "text": TASK}]}
|
||||
]
|
||||
termination = "unknown"
|
||||
exception: dict[str, Any] | None = None
|
||||
|
||||
def api_response_callback(request: Any, response: Any, error: Any) -> None:
|
||||
index = len(api_calls) + 1
|
||||
request_body = b""
|
||||
try:
|
||||
request_body = request.content or b""
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
response_json = None
|
||||
response_status = getattr(response, "status_code", None)
|
||||
try:
|
||||
response_json = response.json()
|
||||
except Exception:
|
||||
if isinstance(response, (dict, list)):
|
||||
response_json = response
|
||||
|
||||
receipt_name = f"response-{index:02d}.json"
|
||||
if response_json is not None:
|
||||
(receipts / receipt_name).write_text(
|
||||
json.dumps(json_safe(response_json), indent=2, ensure_ascii=False)
|
||||
+ "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
else:
|
||||
receipt_name = None
|
||||
|
||||
headers = getattr(response, "headers", {}) or {}
|
||||
api_calls.append(
|
||||
{
|
||||
"index": index,
|
||||
"observed_at": utc_now(),
|
||||
"request": {
|
||||
"method": getattr(request, "method", None),
|
||||
"url": str(getattr(request, "url", "")),
|
||||
"body_bytes": len(request_body),
|
||||
"body_sha256": sha256_bytes(request_body),
|
||||
"credential_header_present": bool(
|
||||
getattr(request, "headers", {}).get("x-api-key")
|
||||
),
|
||||
},
|
||||
"response": {
|
||||
"http_status": response_status,
|
||||
"request_id": headers.get("request-id")
|
||||
or headers.get("x-request-id"),
|
||||
"message_id": (
|
||||
response_json.get("id")
|
||||
if isinstance(response_json, dict)
|
||||
else None
|
||||
),
|
||||
"model": (
|
||||
response_json.get("model")
|
||||
if isinstance(response_json, dict)
|
||||
else None
|
||||
),
|
||||
"stop_reason": (
|
||||
response_json.get("stop_reason")
|
||||
if isinstance(response_json, dict)
|
||||
else None
|
||||
),
|
||||
"usage": (
|
||||
response_json.get("usage")
|
||||
if isinstance(response_json, dict)
|
||||
else None
|
||||
),
|
||||
"receipt": (
|
||||
f"api_receipts/{receipt_name}" if receipt_name else None
|
||||
),
|
||||
},
|
||||
"error_type": type(error).__name__ if error else None,
|
||||
"error": str(error) if error else None,
|
||||
}
|
||||
)
|
||||
|
||||
def output_callback(block: Any) -> None:
|
||||
nonlocal refused_action
|
||||
value = json_safe(block)
|
||||
if not isinstance(value, dict) or value.get("type") != "tool_use":
|
||||
return
|
||||
if len(actions) >= ACTION_LIMIT:
|
||||
refused_action = {
|
||||
"tool_use_id": value.get("id"),
|
||||
"tool": value.get("name"),
|
||||
"input": value.get("input"),
|
||||
"executed": False,
|
||||
"reason": "action_limit",
|
||||
}
|
||||
raise ActionLimitReached(
|
||||
f"refused action {ACTION_LIMIT + 1}; limit is {ACTION_LIMIT}"
|
||||
)
|
||||
record = {
|
||||
"index": len(actions) + 1,
|
||||
"tool_use_id": value.get("id"),
|
||||
"tool": value.get("name"),
|
||||
"input": value.get("input"),
|
||||
"executed": True,
|
||||
"result": None,
|
||||
}
|
||||
actions.append(record)
|
||||
action_by_id[str(value.get("id"))] = record
|
||||
|
||||
def tool_output_callback(result: ToolResult, tool_use_id: str) -> None:
|
||||
record = action_by_id[tool_use_id]
|
||||
image_path = None
|
||||
image_sha256 = None
|
||||
image_bytes = 0
|
||||
if result.base64_image:
|
||||
import base64
|
||||
|
||||
raw = base64.b64decode(result.base64_image)
|
||||
image_path = f"screenshots/action-{record['index']:02d}.png"
|
||||
(OUT / image_path).write_bytes(raw)
|
||||
image_sha256 = sha256_bytes(raw)
|
||||
image_bytes = len(raw)
|
||||
record["result"] = {
|
||||
"output": result.output,
|
||||
"error": result.error,
|
||||
"system": result.system,
|
||||
"screenshot": image_path,
|
||||
"screenshot_sha256": image_sha256,
|
||||
"screenshot_bytes": image_bytes,
|
||||
}
|
||||
|
||||
try:
|
||||
await sampling_loop(
|
||||
model=MODEL,
|
||||
provider=APIProvider.ANTHROPIC,
|
||||
system_prompt_suffix=(
|
||||
"This is a bounded, read-only evaluation. Do not sign in, accept "
|
||||
"agreements, submit forms, or modify external data. Use the GUI "
|
||||
"to perform the requested Google search and ground the final answer "
|
||||
"in the visible result. If Google presents a CAPTCHA or other human "
|
||||
"verification challenge, do not interact with it and do not ask the "
|
||||
"user to solve it. Instead, navigate directly to this reputable, "
|
||||
"read-only Open-Meteo current-weather endpoint: "
|
||||
"https://api.open-meteo.com/v1/forecast?latitude=37.7749&longitude="
|
||||
"-122.4194¤t=temperature_2m,weather_code&temperature_unit="
|
||||
"fahrenheit&timezone=America%2FLos_Angeles . Read the visible JSON, "
|
||||
"interpret its WMO weather code, identify Open-Meteo as the alternate "
|
||||
"source, and finish immediately. You have at most 25 actions total: "
|
||||
"once a credible current temperature and condition are visible, do "
|
||||
"not scroll or explore further; return the final answer."
|
||||
),
|
||||
messages=messages,
|
||||
output_callback=output_callback,
|
||||
tool_output_callback=tool_output_callback,
|
||||
api_response_callback=api_response_callback,
|
||||
api_key=os.environ["ANTHROPIC_API_KEY"],
|
||||
only_n_most_recent_images=3,
|
||||
max_tokens=4096,
|
||||
tool_version=TOOL_VERSION,
|
||||
thinking_budget=None,
|
||||
token_efficient_tools_beta=False,
|
||||
)
|
||||
termination = "model_finished"
|
||||
except ActionLimitReached as exc:
|
||||
termination = "action_limit"
|
||||
exception = {"type": type(exc).__name__, "message": str(exc)}
|
||||
except Exception as exc:
|
||||
termination = "error"
|
||||
exception = {
|
||||
"type": type(exc).__name__,
|
||||
"message": str(exc),
|
||||
"traceback": traceback.format_exc(),
|
||||
}
|
||||
|
||||
final_texts: list[str] = []
|
||||
for message in reversed(messages):
|
||||
if message.get("role") != "assistant":
|
||||
continue
|
||||
for block in message.get("content", []):
|
||||
value = json_safe(block)
|
||||
if isinstance(value, dict) and value.get("type") == "text":
|
||||
final_texts.append(value.get("text", ""))
|
||||
if final_texts:
|
||||
break
|
||||
|
||||
usage_totals: dict[str, int] = {}
|
||||
for call in api_calls:
|
||||
usage = call["response"].get("usage") or {}
|
||||
for key, value in usage.items():
|
||||
if isinstance(value, int):
|
||||
usage_totals[key] = usage_totals.get(key, 0) + value
|
||||
|
||||
final_stop_reason = (
|
||||
api_calls[-1]["response"].get("stop_reason") if api_calls else None
|
||||
)
|
||||
record = {
|
||||
"schema_version": 1,
|
||||
"experiment": "6-7",
|
||||
"status": "completed" if termination == "model_finished" else termination,
|
||||
"started_at": started_at,
|
||||
"finished_at": utc_now(),
|
||||
"task": TASK,
|
||||
"safety": {
|
||||
"read_only": True,
|
||||
"sign_in_allowed": False,
|
||||
"external_mutation_allowed": False,
|
||||
},
|
||||
"provider": "Anthropic API",
|
||||
"requested_model": MODEL,
|
||||
"observed_models": sorted(
|
||||
{
|
||||
call["response"]["model"]
|
||||
for call in api_calls
|
||||
if call["response"].get("model")
|
||||
}
|
||||
),
|
||||
"tool_version": TOOL_VERSION,
|
||||
"action_limit": ACTION_LIMIT,
|
||||
"actions_executed": len(actions),
|
||||
"termination": termination,
|
||||
"provider_stop_reason": final_stop_reason,
|
||||
"final_answer": "\n".join(reversed(final_texts)).strip(),
|
||||
"exception": exception,
|
||||
"refused_action": refused_action,
|
||||
"usage_totals": usage_totals,
|
||||
"api_calls": api_calls,
|
||||
"actions": actions,
|
||||
"runtime": {
|
||||
"python": sys.version,
|
||||
"platform": platform.platform(),
|
||||
"machine": platform.machine(),
|
||||
"source_commit": os.environ.get("EXP96_SOURCE_COMMIT"),
|
||||
"dockerfile_sha256": os.environ.get("EXP96_DOCKERFILE_SHA256"),
|
||||
"image_id": os.environ.get("EXP96_IMAGE_ID"),
|
||||
"base_image_digest": os.environ.get("EXP96_BASE_IMAGE_DIGEST"),
|
||||
},
|
||||
}
|
||||
(OUT / "trajectory.json").write_text(
|
||||
json.dumps(record, indent=2, ensure_ascii=False) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"status": record["status"],
|
||||
"api_calls": len(api_calls),
|
||||
"actions_executed": len(actions),
|
||||
"final_stop_reason": final_stop_reason,
|
||||
"final_answer": record["final_answer"],
|
||||
},
|
||||
ensure_ascii=False,
|
||||
)
|
||||
)
|
||||
return 0 if termination == "model_finished" else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(asyncio.run(main()))
|
||||
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deterministically validate retained Experiment 6-7 evidence."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
EXPECTED_MODEL = "claude-sonnet-4-5-20250929"
|
||||
EXPECTED_SOURCE = "9bcc95e316e5ef6542b4c9d0469f4078829eead5"
|
||||
EXPECTED_DOCKERFILE = (
|
||||
"3aa1f36a491f8f88d81a04c6a89b4cc9f9acd20ad946304c13419736da7c0ead"
|
||||
)
|
||||
|
||||
|
||||
def sha256(path: Path) -> str:
|
||||
digest = hashlib.sha256()
|
||||
with path.open("rb") as handle:
|
||||
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
|
||||
digest.update(chunk)
|
||||
return digest.hexdigest()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("run_dir", type=Path)
|
||||
args = parser.parse_args()
|
||||
run_dir = args.run_dir.resolve()
|
||||
trajectory = json.loads((run_dir / "trajectory.json").read_text(encoding="utf-8"))
|
||||
|
||||
calls = trajectory["api_calls"]
|
||||
actions = trajectory["actions"]
|
||||
final = trajectory["final_answer"]
|
||||
receipts = sorted((run_dir / "api_receipts").glob("response-*.json"))
|
||||
referenced_screenshots = [
|
||||
run_dir / action["result"]["screenshot"]
|
||||
for action in actions
|
||||
if action.get("result") and action["result"].get("screenshot")
|
||||
]
|
||||
|
||||
message_ids = [call["response"].get("message_id") for call in calls]
|
||||
request_ids = [call["response"].get("request_id") for call in calls]
|
||||
tool_ids = [action.get("tool_use_id") for action in actions]
|
||||
temperature = bool(
|
||||
re.search(r"\b-?\d{1,3}\s*(?:°\s*[CF]|degrees?\s*[CF])\b", final, re.I)
|
||||
)
|
||||
condition = bool(
|
||||
re.search(
|
||||
r"\b(?:sunny|clear|cloudy|overcast|fog(?:gy)?|rain(?:y)?|"
|
||||
r"showers?|storm(?:y)?|drizzle|snow(?:y)?|mist(?:y)?|haze|"
|
||||
r"partly\s+cloudy|mostly\s+cloudy)\b",
|
||||
final,
|
||||
re.I,
|
||||
)
|
||||
)
|
||||
|
||||
gates = {
|
||||
"source_commit": trajectory["runtime"].get("source_commit")
|
||||
== EXPECTED_SOURCE,
|
||||
"dockerfile_sha256": trajectory["runtime"].get("dockerfile_sha256")
|
||||
== EXPECTED_DOCKERFILE,
|
||||
"immutable_image_id": str(trajectory["runtime"].get("image_id", "")).startswith(
|
||||
"sha256:"
|
||||
),
|
||||
"base_image_digest": str(
|
||||
trajectory["runtime"].get("base_image_digest", "")
|
||||
).startswith("sha256:"),
|
||||
"model_finished": trajectory.get("termination") == "model_finished"
|
||||
and trajectory.get("provider_stop_reason") == "end_turn",
|
||||
"action_ceiling": 0 < len(actions) <= trajectory.get("action_limit", 0) <= 25,
|
||||
"sequential_action_indexes": [a.get("index") for a in actions]
|
||||
== list(range(1, len(actions) + 1)),
|
||||
"unique_tool_use_ids": None not in tool_ids and len(tool_ids) == len(set(tool_ids)),
|
||||
"native_tools_retained": "computer"
|
||||
in {action.get("tool") for action in actions}
|
||||
and {action.get("tool") for action in actions}.issubset(
|
||||
{"computer", "bash", "str_replace_based_edit_tool"}
|
||||
),
|
||||
"all_actions_executed": all(
|
||||
action.get("executed") and action.get("result") is not None
|
||||
for action in actions
|
||||
),
|
||||
"all_provider_calls_succeeded": bool(calls)
|
||||
and all(call["response"].get("http_status") == 200 for call in calls),
|
||||
"provider_model_match": trajectory.get("observed_models") == [EXPECTED_MODEL]
|
||||
and all(call["response"].get("model") == EXPECTED_MODEL for call in calls),
|
||||
"unique_message_ids": None not in message_ids
|
||||
and len(message_ids) == len(set(message_ids)),
|
||||
"unique_request_ids": None not in request_ids
|
||||
and len(request_ids) == len(set(request_ids)),
|
||||
"receipt_count": len(receipts) == len(calls),
|
||||
"screenshots_exist_and_match": bool(referenced_screenshots)
|
||||
and all(
|
||||
path.is_file()
|
||||
and sha256(path)
|
||||
== next(
|
||||
action["result"]["screenshot_sha256"]
|
||||
for action in actions
|
||||
if action.get("result")
|
||||
and action["result"].get("screenshot")
|
||||
and run_dir / action["result"]["screenshot"] == path
|
||||
)
|
||||
for path in referenced_screenshots
|
||||
),
|
||||
"grounded_weather_answer": temperature and condition,
|
||||
"no_captcha_interaction": not any(
|
||||
"captcha" in json.dumps(action.get("input", {})).lower()
|
||||
or "i'm not a robot" in json.dumps(action.get("input", {})).lower()
|
||||
for action in actions
|
||||
),
|
||||
"no_credential_material": not any(
|
||||
b"sk-ant-" in path.read_bytes()
|
||||
for path in run_dir.rglob("*")
|
||||
if path.is_file()
|
||||
),
|
||||
}
|
||||
|
||||
files = []
|
||||
for path in sorted(p for p in run_dir.rglob("*") if p.is_file()):
|
||||
if path.name in {"acceptance.json", "manifest.json"}:
|
||||
continue
|
||||
files.append(
|
||||
{
|
||||
"path": str(path.relative_to(run_dir)),
|
||||
"bytes": path.stat().st_size,
|
||||
"sha256": sha256(path),
|
||||
}
|
||||
)
|
||||
|
||||
acceptance = {
|
||||
"schema_version": 1,
|
||||
"experiment": "6-7",
|
||||
"run_dir": run_dir.name,
|
||||
"passed": all(gates.values()),
|
||||
"gates": gates,
|
||||
"counts": {
|
||||
"api_calls": len(calls),
|
||||
"actions": len(actions),
|
||||
"screenshots": len(referenced_screenshots),
|
||||
"files_hashed": len(files),
|
||||
},
|
||||
}
|
||||
manifest = {
|
||||
"schema_version": 1,
|
||||
"experiment": "6-7",
|
||||
"run_dir": run_dir.name,
|
||||
"files": files,
|
||||
}
|
||||
(run_dir / "acceptance.json").write_text(json.dumps(acceptance, indent=2) + "\n", encoding="utf-8")
|
||||
(run_dir / "manifest.json").write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
|
||||
print(json.dumps(acceptance, indent=2))
|
||||
return 0 if acceptance["passed"] else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,60 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"experiment": "6-7",
|
||||
"kind": "credential_free_external_authentication_preflight",
|
||||
"generated_at_utc": "2026-08-02T18:25:43Z",
|
||||
"status": "blocked",
|
||||
"completion_gate_pass": false,
|
||||
"source": {
|
||||
"repository": "https://github.com/anthropics/claude-quickstarts.git",
|
||||
"commit_expected": "9bcc95e316e5ef6542b4c9d0469f4078829eead5",
|
||||
"commit_observed": "9bcc95e316e5ef6542b4c9d0469f4078829eead5",
|
||||
"project_path": "computer-use-demo",
|
||||
"dockerfile_sha256_expected": "3aa1f36a491f8f88d81a04c6a89b4cc9f9acd20ad946304c13419736da7c0ead",
|
||||
"dockerfile_sha256_observed": "3aa1f36a491f8f88d81a04c6a89b4cc9f9acd20ad946304c13419736da7c0ead"
|
||||
},
|
||||
"host": {
|
||||
"docker_available": true,
|
||||
"docker_client_version": "24.0.7",
|
||||
"docker_server_version": "27.3.1",
|
||||
"docker_server_os": "linux",
|
||||
"docker_server_arch": "arm64"
|
||||
},
|
||||
"authentication_probe": {
|
||||
"attempted": true,
|
||||
"endpoint": "https://api.anthropic.com/v1/messages",
|
||||
"method": "POST",
|
||||
"anthropic_version": "2023-06-01",
|
||||
"requested_model": "claude-sonnet-4-5-20250929",
|
||||
"max_tokens": 4,
|
||||
"prompt": "Reply OK",
|
||||
"credential_environment_variable": "ANTHROPIC_API_KEY",
|
||||
"credential_present": true,
|
||||
"credential_has_expected_prefix": true,
|
||||
"credential_has_plausible_length": true,
|
||||
"credential_contains_whitespace": false,
|
||||
"credential_value_retained": false,
|
||||
"authorization_header_retained": false,
|
||||
"http_status": 401,
|
||||
"response": {
|
||||
"type": "error",
|
||||
"error": {
|
||||
"type": "authentication_error",
|
||||
"message": "API key is invalid."
|
||||
},
|
||||
"response_id": null,
|
||||
"usage": null
|
||||
}
|
||||
},
|
||||
"execution": {
|
||||
"container_build_attempted": false,
|
||||
"container_started": false,
|
||||
"native_computer_tool_call_attempted": false,
|
||||
"read_only_weather_task_attempted": false,
|
||||
"external_mutation_attempted": false
|
||||
},
|
||||
"blockers": [
|
||||
"The configured ANTHROPIC_API_KEY was rejected by the official Anthropic endpoint before model or native tool access."
|
||||
],
|
||||
"resume_condition": "Configure a valid, funded Anthropic credential, then run the pinned local Demo and retain one complete bounded read-only native-tool trajectory."
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2Z9qtvU56An1kkrDuB",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "I'll help you search for San Francisco weather on Google. Let me start by taking a screenshot to see the current state of the screen, then open Firefox and perform the search."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01RdScMzGdwoVfcqvWFnxgXv",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "screenshot"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 3,
|
||||
"cache_creation_input_tokens": 3370,
|
||||
"cache_read_input_tokens": 0,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 3370,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 88,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2ZShQesQPPkRttPkQw",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "I can see a desktop with a crosshatch pattern background. I need to open Firefox to search for the weather. Let me look for the Firefox icon and click on it."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_019HERV7LDpGRvhegi4BDMin",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
1014,
|
||||
412
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1493,
|
||||
"cache_read_input_tokens": 3023,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1493,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 113,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2ZxDgAVe4qXg8cKKi5",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The screen still shows the desktop. I don't see any visible Firefox icon on the desktop. Let me try to find it by checking if there's a taskbar or panel at the bottom or top of the screen. Let me take another screenshot and try clicking in a different area where the Firefox icon might be located."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01SM5AwDA5d6GuWQBfySXpPZ",
|
||||
"name": "bash",
|
||||
"input": {
|
||||
"command": "export DISPLAY=:1 && (firefox-esr &)"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1175,
|
||||
"cache_read_input_tokens": 4516,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1175,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 129,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2aNjZdY5MkG88r58wb",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good, I've started Firefox using the bash command. Let me wait a moment for it to load and then take a screenshot."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01UAawAvH6atAVSXkWDPzgSQ",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "wait",
|
||||
"duration": 3
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 155,
|
||||
"cache_read_input_tokens": 5691,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 155,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 95,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2azMSfyqC1dx4EFjLf",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Perfect! Firefox has opened and I can see the Welcome to Firefox page with a startup wizard. As instructed, I should ignore this wizard and click directly on the address bar to search. Let me click on the address bar where it says \"Search or enter address\"."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01K3pmZhoUe9wBZqPWoqKxX6",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
298,
|
||||
100
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1157,
|
||||
"cache_read_input_tokens": 5846,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1157,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 130,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2boBmMjj4gvHSTKpC4",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good! The address bar is now active and showing some suggested shortcuts (Wikipedia, YouTube, Reddit, Add-ons for Firefox). Now I need to type \"google.com\" to go to Google first, then search for the San Francisco weather."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01XW7HYpDR3uh1Fb1vyv7yQT",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "type",
|
||||
"text": "google.com"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1192,
|
||||
"cache_read_input_tokens": 7003,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1192,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 122,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2cFoAMrpLSkKAD9NCt",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Perfect! I can see \"google.co\" typed in the address bar with a Google icon appearing. Now I'll press Enter to navigate to Google."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01PnthcKXs4gdTx3bkqxDXKt",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "key",
|
||||
"text": "Return"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1184,
|
||||
"cache_read_input_tokens": 8195,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1184,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 100,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2dGJhHKxQehLqxicd6",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good! The browser is loading google.com (I can see \"http://google.com\" in the address bar). The page is still loading as indicated by the stop button (X) next to the refresh button. Let me wait a moment for Google to fully load."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_013KHUuWZJ19r9f98B2msvBL",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "wait",
|
||||
"duration": 3
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1162,
|
||||
"cache_read_input_tokens": 9379,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1162,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 125,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2ePeFSBuM51yMW14zt",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Excellent! Google has loaded successfully. I can see the Google homepage with the search bar in the center. Now I need to click on the search bar and type \"San Francisco weather today\" to search for the weather information."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01XoDRPSGytck4fhAGPmTVhm",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
511,
|
||||
325
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1187,
|
||||
"cache_read_input_tokens": 10541,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1187,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 122,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2f1WkjQkF8bmACqJu6",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Perfect! The search bar is now active and showing trending searches. I can see the cursor is in the search field. Now I'll type \"San Francisco weather today\" to search for the weather information."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01PabADjuQDnnXiscnSeVos1",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "type",
|
||||
"text": "San Francisco weather today"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1184,
|
||||
"cache_read_input_tokens": 11728,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1184,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 113,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2fuB313YpvLYpujoKz",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good! I can see \"San Francisco weather tod\" has been typed in the search bar (it appears the full text \"San Francisco weather today\" was entered). Now I'll press Enter to perform the search."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01TvQbwEZnZYJ1YAiriBZsGq",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "key",
|
||||
"text": "Return"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1175,
|
||||
"cache_read_input_tokens": 12912,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1175,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 112,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2gdbCYodwUGDCwvS1y",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "I see that Google has detected unusual traffic and is showing a CAPTCHA or verification page. This is the \"About this page\" screen that says \"Our systems have detected unusual traffic from your computer network.\" At the bottom, I can see it says \"Transferring data from www.google.com...\" which suggests the page might be loading. Let me wait a moment to see if the page loads properly."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01ALaQrx8SRZffC2Z6aeyK9F",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "wait",
|
||||
"duration": 3
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1174,
|
||||
"cache_read_input_tokens": 14087,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1174,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 153,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2hVXnnQCHG1xKZ9ib9",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Now I can see a reCAPTCHA verification box has appeared with \"I'm not a robot\" checkbox. According to my instructions, when I see a human verification system like CAPTCHA or reCAPTCHA, I should stop and ask for direction from the user on how to proceed.\n\nI've encountered a reCAPTCHA verification screen from Google. The page is asking me to verify that I'm not a robot before showing the search results. According to my guidelines, I need to ask you how you would like me to proceed with this verification step.\n\nWould you like me to:\n1. Try clicking the reCAPTCHA checkbox to verify (though I'm instructed not to click \"I'm not a robot\" boxes)\n2. Try a different approach to access the weather information\n3. Or provide other guidance on how to proceed?"
|
||||
}
|
||||
],
|
||||
"stop_reason": "end_turn",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1215,
|
||||
"cache_read_input_tokens": 15261,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1215,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 188,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
|
After Width: | Height: | Size: 850 B |
|
After Width: | Height: | Size: 842 B |
|
After Width: | Height: | Size: 126 KiB |
|
After Width: | Height: | Size: 135 KiB |
|
After Width: | Height: | Size: 134 KiB |
|
After Width: | Height: | Size: 127 KiB |
|
After Width: | Height: | Size: 78 KiB |
|
After Width: | Height: | Size: 96 KiB |
|
After Width: | Height: | Size: 99 KiB |
|
After Width: | Height: | Size: 152 KiB |
|
After Width: | Height: | Size: 152 KiB |
@@ -0,0 +1,38 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2snfTXZK4TT4ARryqt",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "I'll help you search for San Francisco weather on Google. Let me start by taking a screenshot to see the current state of the screen, then open a browser and perform the search."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01Dv1Uun91xBoc4cXWvjFqSr",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "screenshot"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 3,
|
||||
"cache_creation_input_tokens": 3429,
|
||||
"cache_read_input_tokens": 0,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 3429,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 89,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2tAMSDPDDPqSb6vfsS",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "I can see a desktop with a dark background. I need to open Firefox to access Google. I can see there should be a Firefox icon on the desktop or taskbar. Let me click on it."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01YbzcWhYRnggmBFvV12r86v",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
512,
|
||||
384
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1494,
|
||||
"cache_read_input_tokens": 3082,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1494,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 117,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2uDppLn8VsVvg5dekQ",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good! The desktop has loaded and I can see the taskbar at the bottom with various application icons. I can see Firefox icon (the orange/red circular icon) in the taskbar. Let me click on it to open Firefox."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_014NTKxo5RkW3MY3BLJj7S4Y",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
806,
|
||||
737
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1179,
|
||||
"cache_read_input_tokens": 4576,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1179,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 124,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2urfcZNPQ33uX4cd4i",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Perfect! Firefox has opened and is showing a welcome page with a startup wizard. As per the instructions, I should ignore this wizard and instead click on the address bar to search. Let me click on the address bar where it says \"Search or enter address\"."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01G4H1R6d3mFZWysjdLQjhBK",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "left_click",
|
||||
"coordinate": [
|
||||
298,
|
||||
100
|
||||
]
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1186,
|
||||
"cache_read_input_tokens": 5755,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1186,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 128,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"model": "claude-sonnet-4-5-20250929",
|
||||
"id": "msg_011Cdf2vo2NVwdqAHicSUsoB",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Good! The address bar is now active and showing some quick access suggestions. I can see the cursor is in the address bar. Now I'll type \"google.com\" to navigate to Google."
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_01FCWF2g7nsGjRdoXyLtg9Zf",
|
||||
"name": "computer",
|
||||
"input": {
|
||||
"action": "type",
|
||||
"text": "google.com"
|
||||
},
|
||||
"caller": {
|
||||
"type": "direct"
|
||||
}
|
||||
}
|
||||
],
|
||||
"stop_reason": "tool_use",
|
||||
"stop_sequence": null,
|
||||
"stop_details": null,
|
||||
"usage": {
|
||||
"input_tokens": 7,
|
||||
"cache_creation_input_tokens": 1190,
|
||||
"cache_read_input_tokens": 6941,
|
||||
"cache_creation": {
|
||||
"ephemeral_5m_input_tokens": 1190,
|
||||
"ephemeral_1h_input_tokens": 0
|
||||
},
|
||||
"output_tokens": 111,
|
||||
"service_tier": "standard",
|
||||
"inference_geo": "not_available"
|
||||
}
|
||||
}
|
||||