Files
ai-agent-book/chapter10/book-translation/validation/real_20260730T061500Z_v4/judge_receipts/unit-11.json
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

43 lines
35 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n```javascript\n// ═══ Request constructed by the Agent framework (1st call) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← Written by developer\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← User input\n \"content\": \"What's the current time and weather in Vancouver?\"\n }\n ],\n \"tools\": [ // ← Tools defined by developer\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": { \"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\" }\n }\n }\n }\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": { \"type\": \"string\", \"description\": \"City name\" },\n \"unit\": { \"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"] }\n }\n }\n }\n }\n ]\n}\n```\n\n**Model returns a tool call request (not a final reply):**\n\n```javascript\n// ═══ Response returned by the API (model decides to call tools) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← Generated by model\n \"content\": null, // No text response\n \"tool_calls\": [ // Model requests two tool calls\n {\n \"id\": \"call_abc123\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\"\n }\n },\n {\n \"id\": \"call_def456\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\"\n }\n }\n ]\n }\n }]\n}\n```\n\nThe model does not answer the user's question yet. Instead, it returns two **tool call requests**: one for the current time and one for the weather. Because these requests are independent, the Agent framework can execute them in parallel. **The model issues the call requests; the Agent framework performs the actual execution.** This division of responsibility is central to Agent architecture: the model decides which tool to call and what arguments to pass, while the framework calls APIs, runs code, and returns the results.\n\n**The Agent framework executes the tools and then initiates a second API call:**\n\nAfter receiving the model's tool call requests, the Agent framework executes the two tools (for example, by calling a time API and a weather API), then sends the **complete conversation history along with the tool execution results** back to the model:\n\n```javascript\n// ═══ Request constructed by the Agent framework (2nd call) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← Same as 1st call\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← Same as 1st call\n \"content\": \"What's the current time and weather in Vancouver?\"\n },\n {\n \"role\": \"assistant\", // ← Model output from 1st call, included verbatim\n \"content\": null,\n \"tool_calls\": [\n { \"id\": \"call_abc123\", \"function\": { \"name\": \"get_current_time\", \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\" } },\n { \"id\": \"call_def456\", \"function\": { \"name\": \"get_weather\", \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\" } }\n ]\n },\n {\n \"role\": \"tool\", // ← Generated by Agent framework (tool execution result)\n \"tool_call_id\": \"call_abc123\",\n \"content\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\", \\\"datetime\\\": \\\"2025-09-13T05:18:47\\\", \\\"day_of_week\\\": \\\"Saturday\\\"}\"\n },\n {\n \"role\": \"tool\", // ← Generated by Agent framework (tool execution result)\n \"tool_call_id\": \"call_def456\",\n \"content\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"temperature\\\": 13.2, \\\"unit\\\": \\\"celsius\\\", \\\"conditions\\\": \\\"clear\\\", \\\"humidity\\\": 93}\"\n }\n ],\n \"tools\": [ ... ] // ← Same tool definitions as above, omitted\n}\n```\n\nThere are three key details here:\n\n1. **The second request includes the full conversation history from the first request** — the system message, the user message, the assistant message containing tool calls, and the newly added tool results. This illustrates the stateless nature of the API: the Agent framework must include the relevant history in every request.\n2. **The first assistant message is inserted back into the message list verbatim** — this gives the next model call access to the tool-call decisions made in the previous call.\n3. **Tool messages are linked to their corresponding tool calls via `tool_call_id`** — this tells the model which result belongs to which requested call.\n\n**The model generates the final response based on the tool results:**\n\n```javascript\n// ═══ Response returned by the API (final reply) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← Generated by model\n \"content\": \"It's currently 5:18 AM on Saturday, September 13, 2025 in Vancouver.\\n\\nWeather: 13.2°C with clear skies and 93% humidity. It's quite cool this morning - you might want to grab a jacket.\"\n }\n }]\n}\n```\n\nThis time, the model does not return `tool_calls`; it returns a text response because the tool results provide enough information to answer the user's question. If more information is needed (for example, if the user asks \"What about Tokyo?\"), the model can return `tool_calls` again, and the Agent framework repeats the same cycle: execute the tools, send back the results, and call the model again. **This \"request → tool call → execution → return results → next request\" cycle is the API-level implementation of the ReAct loop introduced in Chapter 1.**\n\n### Implementing the Agent's Core Loop in Code\n\nNow that the JSON structure is clear, we can connect the steps above in Python. The following is a minimal Agent implementation built around a single loop:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI()\n\n# ── Tool definitions ──\ntools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": {\"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\"}\n },\n },\n },\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": {\"type\": \"string\", \"description\": \"City name\"},\n \"unit\": {\"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"]},\n },\n },\n },\n },\n]\n\n# ── Tool execution function (stub with canned results; a real implementation\n# must parse the JSON `arguments` and call actual APIs) ──\ndef execute_tool(name, arguments):\n if name == \"get_current_time\":\n return '{\"datetime\": \"2025-09-13T05:18:47\", \"day_of_week\": \"Saturday\"}'\n elif name == \"get_weather\":\n return '{\"temperature\": 13.2, \"unit\": \"celsius\", \"conditions\": \"clear\", \"humidity\": 93}'\n\n# ── Initial message list ──\nmessages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant. Use tools to get real-time information when needed.\"},\n {\"role\": \"user\", \"content\": \"What's the current time and weather in Vancouver?\"},\n]\n\n# ── Agent core loop ──\n# Production code needs a max_iterations cap here: as discussed later in\n# this chapter, Agents can become stuck repeating the same tool calls forever\nwhile True:\n response = client.chat.completions.create(\n model=\"Qwen3-0.6B\", messages=messages, tools=tools\n )\n assistant_message = response.choices[0].message\n\n # Append model's response to message list (whether text or tool calls)\n messages.append(assistant_message)\n\n # If no tool calls requested, the model has produced its final response\n if not assistant_message.tool_calls:\n print(assistant_message.content)\n break\n\n # Execute each tool requested by the model, append results to message list\n for tool_call in assistant_message.tool_calls:\n result = execute_tool(tool_call.function.name, tool_call.function.arguments)\n messages.append({\n \"role\": \"tool\",\n \"tool_call_id\": tool_call.id,\n \"content\": result,\n })\n # Return to top of loop, call model again with updated message list\n```\n\nThe loop has one main branch: **if the model returns `tool_calls`, execute the tools and continue; otherwise, output the result and exit.** During this process, the `messages` list keeps growing as each round appends the model's reply and any tool execution results.\n\nThe `messages` list changes across rounds as follows:\n\n**Initial state (before the first call):**\n```\nmessages = [\n { role: \"system\", content: \"You are a helpful assistant...\" }, # Written by developer\n { role: \"user\", content: \"What's the current time and weather in Vancouver?\" }, # User input\n]\n```\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程 [第2部分/共17部分]\n\n```javascript\n// ═══ 由代理框架构造的请求(第一次调用) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 由开发者编写\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← 用户输入\n \"content\": \"What's the current time and weather in Vancouver?\"\n }\n ],\n \"tools\": [ // ← 由开发者定义的工具\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": { \"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\" }\n }\n }\n }\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": { \"type\": \"string\", \"description\": \"City name\" },\n \"unit\": { \"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"] }\n }\n }\n }\n }\n ]\n}\n```\n\n**模型返回工具调用请求(不是最终回复):**\n\n```javascript\n// ═══ API返回的响应(模型决定调用工具) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 由模型生成\n \"content\": null, // 无文本响应\n \"tool_calls\": [ // 模型请求两个工具调用\n {\n \"id\": \"call_abc123\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\"\n }\n },\n {\n \"id\": \"call_def456\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\"\n }\n }\n ]\n }\n }]\n}\n```\n\n模型尚未回答用户的问题。相反,它返回两个**工具调用请求**:一个用于当前时间,一个用于天气。由于这些请求是独立的,代理框架可以并行执行它们。**模型发出调用请求;代理框架执行实际的执行。** 这种责任划分是代理架构的核心:模型决定调用哪个工具以及传递什么参数,而框架调用API、运行代码并返回结果。\n\n**代理框架执行工具,然后发起第二次API调用:**\n\n在收到模型的工具调用请求后,代理框架执行两个工具(例如,调用时间API和天气API),然后将**完整的对话历史以及工具执行结果**发送回模型:\n\n```javascript\n// ═══ 由代理框架构造的请求(第二次调用) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 与第一次调用相同\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← 与第一次调用相同\n \"content\": \"What's the current time and weather in Vancouver?\"\n },\n {\n \"role\": \"assistant\", // ← 第一次调用的模型输出,逐字包含\n \"content\": null,\n \"tool_calls\": [\n { \"id\": \"call_abc123\", \"function\": { \"name\": \"get_current_time\", \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\" } },\n { \"id\": \"call_def456\", \"function\": { \"name\": \"get_weather\", \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\" } }\n ]\n },\n {\n \"role\": \"tool\", // ← 由代理框架生成(工具执行结果)\n \"tool_call_id\": \"call_abc123\",\n \"content\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\", \\\"datetime\\\": \\\"2025-09-13T05:18:47\\\", \\\"day_of_week\\\": \\\"Saturday\\\"}\"\n },\n {\n \"role\": \"tool\", // ← 由代理框架生成(工具执行结果)\n \"tool_call_id\": \"call_def456\",\n \"content\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"temperature\\\": 13.2, \\\"unit\\\": \\\"celsius\\\", \\\"conditions\\\": \\\"clear\\\", \\\"humidity\\\": 93}\"\n }\n ],\n \"tools\": [ ... ] // ← 与上述相同的工具定义,省略\n}\n```\n\n这里有三个关键细节:\n\n1. **第二个请求包含来自第一个请求的完整对话历史** — 系统消息、用户消息、包含工具调用的助手消息以及新添加的工具结果。这说明了API的无状态性质:代理框架必须在每个请求中包含相关的历史记录。\n2. **第一个助手消息逐字插入消息列表中** — 这使下一个模型调用能够访问上一个调用中做出的工具调用决策。\n3. **工具消息通过`tool_call_id`链接到其对应的工具调用** — 这告诉模型哪个结果属于哪个请求的调用。\n\n**模型根据工具结果生成最终响应:**\n\n```javascript\n// ═══ API返回的响应(最终回复) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 由模型生成\n \"content\": \"It's currently 5:18 AM on Saturday, September 13, 2025 in Vancouver.\\n\\nWeather: 13.2°C with clear skies and 93% humidity. It's quite cool this morning - you might want to grab a jacket.\"\n }\n }]\n}\n```\n\n这次,模型不返回`tool_calls`;它返回文本响应,因为工具结果提供了足够的信息来回答用户的问题。如果需要更多信息(例如,如果用户问“东京呢?”),模型可以再次返回`tool_calls`,代理框架重复相同的循环:执行工具、发送结果并再次调用模型。**这个“请求→工具调用→执行→返回结果→下一个请求”循环是第1章中介绍的ReAct循环的API级实现。**\n\n### 在代码中实现代理的核心循环\n\n现在JSON结构清晰了,我们可以在Python中连接上述步骤。以下是围绕单个循环构建的最小代理实现:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI()\n\n# ── 工具定义 ──\ntools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": {\"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\"}\n },\n },\n },\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": {\"type\": \"string\", \"description\": \"City name\"},\n \"unit\": {\"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"]},\n },\n },\n },\n },\n]\n\n# ── 工具执行函数(带有固定结果的存根;实际实现必须解析JSON `arguments`并调用实际API ──\ndef execute_tool(name, arguments):\n if name == \"get_current_time\":\n return '{\"datetime\": \"2025-09-13T05:18:47\", \"day_of_week\": \"Saturday\"}'\n elif name == \"get_weather\":\n return '{\"temperature\": 13.2, \"unit\": \"celsius\", \"conditions\": \"clear\", \"humidity\": 93}'\n\n# ── 初始消息列表 ──\nmessages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant. Use tools to get real-time information when needed.\"},\n {\"role\": \"user\", \"content\": \"What's the current time and weather in Vancouver?\"},\n]\n\n# ── 代理核心循环 ──\n# 生产代码需要在此处设置max_iterations限制:如本章后面所述,代理可能会永远重复相同的工具调用而陷入困境\nwhile True:\n response = client.chat.completions.create(\n model=\"Qwen3-0.6B\", messages=messages, tools=tools\n )\n assistant_message = response.choices[0].message\n\n # 将模型的响应附加到消息列表中(无论是文本还是工具调用)\n messages.append(assistant_message)\n\n # 如果没有请求工具调用,模型已经生成了最终响应\n if not assistant_message.tool_calls:\n print(assistant_message.content)\n break\n\n # 执行模型请求的每个工具,将结果附加到消息列表中\n for tool_call in assistant_message.tool_calls:\n result = execute_tool(tool_call.function.name, tool_call.function.arguments)\n messages.append({\n \"role\": \"tool\",\n \"tool_call_id\": tool_call.id,\n \"content\": result,\n })\n # 返回循环顶部,使用更新后的消息列表再次调用模型\n```\n\n循环有一个主要分支:**如果模型返回`tool_calls`,则执行工具并继续;否则,输出结果并退出。** 在此过程中,`messages`列表随着每一轮附加模型的回复和任何工具执行结果而不断增长。\n\n`messages`列表在各轮之间的变化如下:\n\n**初始状态(第一次调用之前):**\n```\nmessages = [\n { role: \"system\", content: \"You are a helpful assistant...\" }, # 由开发者编写\n { role: \"user\", content: \"What's the current time and weather in Vancouver?\" }, # 用户输入\n]\n```\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第2/17部分]\n```javascript\n// ═══ 代理框架构造的请求(第一次调用) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 开发者编写\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← 用户输入\n \"content\": \"What's the current time and weather in Vancouver?\"\n }\n ],\n \"tools\": [ // ← 开发者定义的工具\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": { \"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\" }\n }\n }\n }\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": { \"type\": \"string\", \"description\": \"City name\" },\n \"unit\": { \"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"] }\n }\n }\n }\n }\n ]\n}\n```\n\n**模型返回工具调用请求(不是最终回复):**\n\n```javascript\n// ═══ API返回的响应(模型决定调用工具) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 模型生成\n \"content\": null, // 无文本响应\n \"tool_calls\": [ // 模型请求两个工具调用\n {\n \"id\": \"call_abc123\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\"\n }\n },\n {\n \"id\": \"call_def456\",\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\"\n }\n }\n ]\n }\n }]\n}\n```\n\n模型还没有回答用户的问题。相反,它返回两个**工具调用请求**:一个用于当前时间,一个用于天气。因为这些请求是独立的,代理框架可以并行执行它们。**模型发出调用请求;代理框架执行实际的调用。** 这种职责分工是代理架构的核心:模型决定调用哪个工具以及传递什么参数,而框架调用API、运行代码并返回结果。\n\n**代理框架执行工具,然后发起第二次API调用:**\n\n在收到模型的工具调用请求后,代理框架执行这两个工具(例如,调用时间API和天气API),然后将**完整的对话历史以及工具执行结果**发送回模型:\n\n```javascript\n// ═══ 代理框架构造的请求(第二次调用) ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 与第一次调用相同\n \"content\": \"You are a helpful assistant. Use the provided tools to get real-time information when needed.\"\n },\n {\n \"role\": \"user\", // ← 与第一次调用相同\n \"content\": \"What's the current time and weather in Vancouver?\"\n },\n {\n \"role\": \"assistant\", // ← 第一次调用的模型输出,逐字包含\n \"content\": null,\n \"tool_calls\": [\n { \"id\": \"call_abc123\", \"function\": { \"name\": \"get_current_time\", \"arguments\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\"}\" } },\n { \"id\": \"call_def456\", \"function\": { \"name\": \"get_weather\", \"arguments\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"unit\\\": \\\"celsius\\\"}\" } }\n ]\n },\n {\n \"role\": \"tool\", // ← 由代理框架生成(工具执行结果)\n \"tool_call_id\": \"call_abc123\",\n \"content\": \"{\\\"timezone\\\": \\\"America/Vancouver\\\", \\\"datetime\\\": \\\"2025-09-13T05:18:47\\\", \\\"day_of_week\\\": \\\"Saturday\\\"}\"\n },\n {\n \"role\": \"tool\", // ← 由代理框架生成(工具执行结果)\n \"tool_call_id\": \"call_def456\",\n \"content\": \"{\\\"city\\\": \\\"Vancouver\\\", \\\"temperature\\\": 13.2, \\\"unit\\\": \\\"celsius\\\", \\\"conditions\\\": \\\"clear\\\", \\\"humidity\\\": 93}\"\n }\n ],\n \"tools\": [ ... ] // ← 与上述相同的工具定义,省略\n}\n```\n\n这里有三个关键细节:\n\n1. **第二次请求包含第一次请求的完整对话历史** — 系统消息、用户消息、包含工具调用的助手消息以及新添加的工具结果。这说明了API的无状态性质:代理框架必须在每个请求中包含相关历史。\n2. **第一次助手消息逐字插入消息列表** — 这让下一次模型调用能够访问前一次调用中做出的工具调用决策。\n3. **工具消息通过`tool_call_id`与相应的工具调用关联** — 这告诉模型哪个结果属于哪个请求的调用。\n\n**模型根据工具结果生成最终响应:**\n\n```javascript\n// ═══ API返回的响应(最终回复) ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 模型生成\n \"content\": \"It's currently 5:18 AM on Saturday, September 13, 2025 in Vancouver.\\n\\nWeather: 13.2°C with clear skies and 93% humidity. It's quite cool this morning - you might want to grab a jacket.\"\n }\n }]\n}\n```\n\n这次,模型没有返回`tool_calls`;它返回文本响应,因为工具结果提供了足够的信息来回答用户的问题。如果需要更多信息(例如,用户问“东京呢?”),模型可以再次返回`tool_calls`,代理框架重复相同的循环:执行工具、发送结果并再次调用模型。**这个“请求→工具调用→执行→返回结果→下一次请求”的循环是第1章介绍的ReAct循环的API级实现。**\n\n### 在代码中实现代理的核心循环\n现在JSON结构清晰了,我们可以在Python中连接上述步骤。以下是围绕单个循环构建的最小代理实现:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI()\n\n# ── 工具定义 ──\ntools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_current_time\",\n \"description\": \"Get the current date and time in a specific timezone\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"timezone\": {\"type\": \"string\", \"description\": \"Timezone name, e.g. America/Vancouver\"}\n },\n },\n },\n },\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"get_weather\",\n \"description\": \"Get the current weather for a specific city\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"city\": {\"type\": \"string\", \"description\": \"City name\"},\n \"unit\": {\"type\": \"string\", \"enum\": [\"celsius\", \"fahrenheit\"]},\n },\n },\n },\n },\n]\n\n# ── 工具执行函数(带有固定结果的存根;实际实现必须解析JSON `arguments`并调用实际API ──\ndef execute_tool(name, arguments):\n if name == \"get_current_time\":\n return '{\"datetime\": \"2025-09-13T05:18:47\", \"day_of_week\": \"Saturday\"}'\n elif name == \"get_weather\":\n return '{\"temperature\": 13.2, \"unit\": \"celsius\", \"conditions\": \"clear\", \"humidity\": 93}'\n\n# ── 初始消息列表 ──\nmessages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant. Use tools to get real-time information when needed.\"},\n {\"role\": \"user\", \"content\": \"What's the current time and weather in Vancouver?\"},\n]\n\n# ── 代理核心循环 ──\n# 生产代码需要在此处设置max_iterations限制:如本章后面所述,代理可能永远重复相同的工具调用而陷入困境\nwhile True:\n response = client.chat.completions.create(\n model=\"Qwen3-0.6B\", messages=messages, tools=tools\n )\n assistant_message = response.choices[0].message\n\n # 将模型的响应追加到消息列表(无论是文本还是工具调用)\n messages.append(assistant_message)\n\n # 如果没有请求工具调用,模型已生成最终响应\n if not assistant_message.tool_calls:\n print(assistant_message.content)\n break\n\n # 执行模型请求的每个工具,将结果追加到消息列表\n for tool_call in assistant_message.tool_calls:\n result = execute_tool(tool_call.function.name, tool_call.function.arguments)\n messages.append({\n \"role\": \"tool\",\n \"tool_call_id\": tool_call.id,\n \"content\": result,\n })\n # 返回循环顶部,使用更新后的消息列表再次调用模型\n```\n\n循环有一个主要分支:**如果模型返回`tool_calls`,执行工具并继续;否则,输出结果并退出。** 在此过程中,`messages`列表随着每一轮追加模型的回复和任何工具执行结果而不断增长。\n\n`messages`列表在各轮之间的变化如下:\n\n**初始状态(第一次调用前):**\n```\nmessages = [\n { role: \"system\", content: \"You are a helpful assistant...\" }, # 开发者编写\n { role: \"user\", content: \"What's the current time and weather in Vancouver?\" }, # 用户输入\n]\n```"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853906994433ea4fa82759f0cae4f2ffa723d4f9dd6af3b57",
"model": "doubao-seed-1-6-250615",
"created": 1785390718,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims, such as 'the Agent framework can execute them in parallel' and 'the model decides which tool to call', are accurately translated without omissions or inventions. The key concept of 'ReAct loop' is correctly rendered as 'ReAct循环'.\"},\"fluency\":{\"score\":4,\"evidence\":\"Most sentences are fluent, e.g., '这种责任划分是代理架构的核心' (This division of responsibility is central to Agent architecture). However, '代理框架执行实际的执行' (the Agent framework performs the actual execution) is slightly redundant with repeated '执行' (execution).\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of technical terms: 'tool call requests' → '工具调用请求', 'API-level implementation' → 'API级实现', 'stub' → '存根', 'max_iterations' → 'max_iterations' (retained as is, standard practice).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"All Markdown elements are preserved: JavaScript/Python code blocks with correct syntax highlighting, headings (###), comments in code (e.g., // ← 由开发者编写), and code section labels like '// ═══ 由代理框架构造的请求(第一次调用) ═══'.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical content is accurately translated, including the description of the Agent loop: '请求→工具调用→执行→返回结果→下一次请求' (request→tool call→execution→return results→next request) matches the source exactly.\"},\"fluency\":{\"score\":5,\"evidence\":\"More natural phrasing than X in key sentences, e.g., '代理框架执行实际的调用' (the Agent framework performs the actual calls) avoids redundancy compared to X's '执行实际的执行' (performs the actual execution). Sentences like '循环有一个主要分支' (The loop has one main branch) flow smoothly.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent terminology identical to X: 'tool_calls' → 'tool_calls' (retained), 'stateless nature' → '无状态性质', 'ReAct loop' → 'ReAct循环'. No inconsistent term usage.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown structure is fully preserved, including code blocks, headings, comments, and special characters like '// ═══' and '// ──'. Code syntax and formatting (e.g., Python indentation) are unchanged.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y demonstrates slightly better fluency by resolving redundant phrasing in X (e.g., '执行实际的执行' vs. '执行实际的调用') while maintaining identical accuracy, terminology consistency, and markdown fidelity. This makes Y more readable without sacrificing technical precision.\"} }",
"usage": {
"prompt_tokens": 7855,
"completion_tokens": 617,
"total_tokens": 8472
}
},
"latency_milliseconds": 19063,
"validation": {
"valid": true
}
}
]
}