{ "schema_version": 1, "credential_free": true, "attempts": [ { "attempt": 1, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Current Environment State**: Includes dynamic environment information (system time, working directory, etc.), abnormal operation alerts (\"This tool has been called N times repeatedly\"), and the transformation from implicit state to explicit state. This design principle also applies to human interfaces—both Command Line Interfaces (CLI) and Graphical User Interfaces (GUI) aim to let users clearly perceive the current state of the system.\n\n**Available Capability List**: When the Agent framework supports plugin-based capability extensions (like the Skills system from the previous section), the metadata list of all installed Skills also goes through this same end-of-context injection channel. It tells the model which specialized capabilities are currently available. It changes infrequently (only when the user installs or uninstalls a Skill), and its incremental sending mechanism was detailed in the previous Skills section, so it will not be repeated here.\n\nSide-channel information and the available capability list usually do not change after being added, making them cache-friendly because they do not invalidate the cached prefix. Task planning and environment state are dynamic and must be appended to the end of the context as special user messages, then updated as the task progresses. The update method directly affects KV Cache cost, as discussed below.\n\n### Specific Position of the Agent Status Bar in the Context\n\n![Figure 2-15: Insertion Position of the Agent Status Bar in the API Message List](images/fig2-15.svg)\n\nAn important implementation detail is that the Agent Status Bar is inserted at the end of the context as **a message with the `user` role** at the API level, rather than by modifying the initial `system` message. The reason is the KV Cache constraint discussed earlier: modifying the `system` message would invalidate the cache for the entire prefix. One point requires clarification: the `user` role here is a technical choice at the API protocol level and is not equivalent to \"input from the end-user\" as defined in Chapter 1. The Harness borrows the `user` role message slot to inject system state information generated by the Agent framework. The content does not come from a real user; it simply uses the `user` message format to attach state information to the end of the context.\n\nBelow is the actual message list constructed by the Agent framework during the Nth API call:\n\n```\nmessages: [\n { role: \"system\", content: \"You are a customer service assistant...\" } ← Fixed (KV Cache cached)\n { role: \"user\", content: \"Help me cancel my Xfinity plan\" } ← Original user request\n { role: \"assistant\", content: null, tool_calls: [...] } ← Round 1: model decides to call\n { role: \"tool\", content: \"Call log...\" } ← Round 1: call result\n { role: \"assistant\", content: null, tool_calls: [...] } ← Round 2: model decides to call again\n { role: \"tool\", content: \"Call log...\" } ← Round 2: call result\n ...(more rounds)\n { role: \"user\", content: \"Can you call them again to follow up?\" } ← User follow-up\n { role: \"user\", content: \" ← Status bar injected by Agent framework\n Current State: (as a user message)\n - phone_call invoked 3 times (Xfinity: 3/3 max)\n - Current time: 2025-09-14 10:30:45\n - TODO: [1] Cancel plan (in_progress)\n \" }\n]\n```\n\nNote the last message: its `role` is `user`, but the content is meta-information automatically generated by the Agent framework, wrapped in `` tags so the model can recognize its special nature. This message sits at the very end of the context, immediately adjacent to the new tokens the model is about to generate, thus receiving the highest attention weight. At the same time, because it is appended rather than modified, all previously cached content remains unaffected.\n\nThis design applies the core principle from the KV Cache section to the status bar: append dynamic information at the end, and keep static information unchanged.\n\n### Two Implementations of Status Updates and Their Cache Costs\n\n\"Appending does not break the cache\" only holds for a single injection. Status naturally changes over time: TODO items are completed, tool counts increase, and previous status messages become outdated. There are two ways to update the status bar, each with different cache costs:\n\n**Implementation 1: Replace each round.** Before each API call, remove the previous round's status message from the message list and append the latest status at the end. This keeps only one current status in the context. The cost is that removing the old status invalidates all cached content after its position, which is the same invalidation mechanism discussed in the \"dynamic timestamp\" section of this chapter. The difference is that because the status message is near the end of the context, the invalidation range is limited to the most recent few rounds of messages rather than the entire prefix.\n\n**Implementation 2: Persistent appending.** Once injected, the status message remains permanently in the trajectory, and a new status is appended at the end each round. Claude Code's `` uses this approach: historical status messages remain in the transcript and are never deleted or modified. This method is fully cache-friendly because messages are only appended, never changed, so the prefix remains stable. The cost is that outdated statuses accumulate in the context, consuming tokens and requiring the model to rely on the latest status while ignoring obsolete ones.\n\nThe rule of thumb is: **when status updates are frequent and the trajectory is long, choose Implementation 2**. Replacing the status each round repeatedly invalidates cache entries over a long trajectory, which can cost more than carrying outdated status messages. **When the trajectory is short or a single status message is large** (e.g., a complete TODO list plus environment snapshot), **choose Implementation 1**. Cache invalidation over the last few rounds is cheap, and the context remains clean and unambiguous.\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第13/17部分]\n**当前环境状态**:包括动态环境信息(系统时间、工作目录等)、异常操作警报(“此工具已重复调用N次”)以及从隐含状态到显式状态的转换。这种设计原则也适用于人机界面——命令行界面(CLI)和图形用户界面(GUI)都旨在让用户清晰感知系统的当前状态。\n\n**可用能力列表**:当代理框架支持基于插件的能力扩展(如前一节的技能系统)时,所有已安装技能的元数据列表也通过此相同的上下文末尾注入通道。它告诉模型当前可用的专业能力。它很少变化(仅在用户安装或卸载技能时),其增量发送机制在前一节的技能部分已详细介绍,此处不再重复。\n\n旁道信息和可用能力列表通常在添加后不会改变,使其对缓存友好,因为它们不会使缓存的前缀失效。任务规划和环境状态是动态的,必须作为特殊用户消息附加到上下文末尾,然后随着任务进展进行更新。更新方法直接影响KV缓存成本,如下所述。\n\n### 代理状态栏在上下文中的具体位置\n![图2-15:代理状态栏在API消息列表中的插入位置](images/fig2-15.svg)\n\n一个重要的实现细节是,代理状态栏在API级别作为**具有`user`角色的消息**插入到上下文末尾,而不是通过修改初始的`system`消息。原因是前面讨论的KV缓存约束:修改`system`消息会使整个前缀的缓存失效。有一点需要澄清:这里的`user`角色是API协议层的技术选择,不等同于第1章定义的“最终用户输入”。框架借用`user`消息槽来注入框架生成的系统状态信息。内容不是来自真实用户;它只是使用`user`消息格式将状态信息附加到上下文末尾。\n\n以下是代理框架在第N次API调用期间构造的实际消息列表:\n\n```\nmessages: [\n { role: \"system\", content: \"You are a customer service assistant...\" } ← 固定(KV缓存已缓存)\n { role: \"user\", content: \"Help me cancel my Xfinity plan\" } ← 原始用户请求\n { role: \"assistant\", content: null, tool_calls: [...] } ← 第一轮:模型决定调用\n { role: \"tool\", content: \"Call log...\" } ← 第一轮:调用结果\n { role: \"assistant\", content: null, tool_calls: [...] } ← 第二轮:模型决定再次调用\n { role: \"tool\", content: \"Call log...\" } ← 第二轮:调用结果\n ...(更多轮次)\n { role: \"user\", content: \"Can you call them again to follow up?\" } ← 用户跟进\n { role: \"user\", content: \" ← 框架注入的状态栏\n Current State: (作为用户消息)\n - phone_call invoked 3 times (Xfinity: 3/3 max)\n - Current time: 2025-09-14 10:30:45\n - TODO: [1] Cancel plan (in_progress)\n \" }\n]\n```\n\n注意最后一条消息:它的`role`是`user`,但内容是框架自动生成的元信息,包裹在``标签中,以便模型识别其特殊性质。此消息位于上下文的最末尾,紧邻模型即将生成的新标记,因此获得最高的注意力权重。同时,由于它是附加而不是修改,之前缓存的所有内容保持不受影响。\n\n这种设计将KV缓存部分的核心原则应用到状态栏:在末尾附加动态信息,保持静态信息不变。\n\n### 状态更新的两种实现及其缓存成本\n“附加不会破坏缓存”仅适用于单次注入。状态自然随时间变化:待办事项完成、工具计数增加、之前的状态消息过时。有两种更新状态栏的方法,每种都有不同的缓存成本:\n\n**实现1:每轮替换**。在每次API调用前,从消息列表中删除前一轮的状态消息,并在末尾附加最新状态。这使上下文中仅保留一个当前状态。成本是删除旧状态会使其位置之后的所有缓存内容失效,这与本章“动态时间戳”部分讨论的相同失效机制。不同之处在于,由于状态消息位于上下文末尾附近,失效范围仅限于最近的几轮消息,而不是整个前缀。\n\n**实现2:持久附加**。一旦注入,状态消息永久保留在轨迹中,每轮在末尾附加新状态。Claude Code的``使用这种方法:历史状态消息保留在记录中,永不删除或修改。这种方法完全对缓存友好,因为消息仅附加,从不更改,因此前缀保持稳定。成本是过时的状态累积在上下文中,消耗标记,并要求模型依赖最新状态而忽略过时状态。\n\n经验法则是:**当状态更新频繁且轨迹较长时,选择实现2**。每轮替换状态在长轨迹中反复使缓存条目失效,可能比携带过时状态消息成本更高。**当轨迹较短或单个状态消息较大**(例如完整的待办事项列表加上环境快照),**选择实现1**。最后几轮的缓存失效成本较低,上下文保持干净明确。\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第13/17部分]\n\n**当前环境状态**:包括动态环境信息(系统时间、工作目录等)、异常操作警报(“该工具已被重复调用N次”)以及从隐式状态到显式状态的转换。这一设计原则也适用于人机界面——命令行界面(CLI)和图形用户界面(GUI)都旨在让用户清晰感知系统的当前状态。\n\n**可用能力列表**:当Agent框架支持基于插件的能力扩展(如前一节的技能系统)时,所有已安装技能的元数据列表也通过相同的上下文末尾注入通道。它告知模型当前可用的专门能力。它很少变化(仅在用户安装或卸载技能时变化),其增量发送机制在前一节的技能部分已详细说明,此处不再重复。\n\n侧信道信息和可用能力列表在添加后通常不会改变,这使得它们对缓存友好,因为它们不会使缓存前缀失效。任务规划和环境状态是动态的,必须作为特殊用户消息附加到上下文末尾,并随着任务进展进行更新。更新方法直接影响键值缓存成本,如下文所述。\n\n### Agent状态栏在上下文中的具体位置\n\n![图2-15:API消息列表中Agent状态栏的插入位置](images/fig2-15.svg)\n\n一个重要的实现细节是,Agent状态栏在API层面作为**角色为`user`的消息**插入到上下文末尾,而不是通过修改初始的`system`消息。原因是前面讨论的键值缓存约束:修改`system`消息会使整个前缀的缓存失效。有一点需要澄清:这里的`user`角色是API协议层面的技术选择,不等同于第1章定义的“最终用户输入”。框架借用`user`角色消息槽来注入由Agent框架生成的系统状态信息。内容并非来自真实用户;它只是使用`user`消息格式将状态信息附加到上下文末尾。\n\n以下是Agent框架在第N次API调用期间构建的实际消息列表:\n\n```\nmessages: [\n { role: \"system\", content: \"You are a customer service assistant...\" } ← 固定(KV Cache缓存)\n { role: \"user\", content: \"Help me cancel my Xfinity plan\" } ← 原始用户请求\n { role: \"assistant\", content: null, tool_calls: [...] } ← 第1轮:模型决定调用\n { role: \"tool\", content: \"Call log...\" } ← 第1轮:调用结果\n { role: \"assistant\", content: null, tool_calls: [...] } ← 第2轮:模型决定再次调用\n { role: \"tool\", content: \"Call log...\" } ← 第2轮:调用结果\n ...(更多轮次)\n { role: \"user\", content: \"Can you call them again to follow up?\" } ← 用户跟进\n { role: \"user\", content: \" ← Agent框架注入的状态栏\n Current State: (作为用户消息)\n - phone_call invoked 3 times (Xfinity: 3/3 max)\n - Current time: 2025-09-14 10:30:45\n - TODO: [1] Cancel plan (in_progress)\n \" }\n]\n```\n\n注意最后一条消息:它的`role`是`user`,但内容是Agent框架自动生成的元信息,用``标签包裹,以便模型识别其特殊性。这条消息位于上下文的最末尾,紧挨着模型即将生成的新词元,因此获得最高的注意力权重。同时,因为是附加而不是修改,之前缓存的所有内容不受影响。\n\n这一设计将键值缓存部分的核心原则应用到状态栏:将动态信息附加在末尾,保持静态信息不变。\n\n### 状态更新的两种实现方式及其缓存成本\n\n“附加不破坏缓存”仅适用于单次注入。状态会随时间自然变化:待办事项完成、工具调用次数增加,之前的状态消息会过时。有两种更新状态栏的方式,每种方式的缓存成本不同:\n\n**实现方式1:每轮替换**。在每次API调用前,从消息列表中移除上一轮的状态消息,并在末尾附加最新状态。这样上下文中仅保留一个当前状态。成本是移除旧状态会使其位置之后的所有缓存内容失效,这与本章“动态时间戳”部分讨论的失效机制相同。不同之处在于,由于状态消息靠近上下文末尾,失效范围仅限于最近的几轮消息,而不是整个前缀。\n\n**实现方式2:持久附加**。一旦注入,状态消息永久保留在轨迹中,每轮在末尾附加新状态。Claude Code的``采用这种方式:历史状态消息保留在记录中,从不删除或修改。这种方法完全对缓存友好,因为消息仅被附加,从不改变,所以前缀保持稳定。成本是过时的状态会累积在上下文中,消耗词元,并要求模型依赖最新状态而忽略过时状态。\n\n经验法则是:**当状态更新频繁且轨迹较长时,选择实现方式2**。每轮重复替换状态会在长轨迹上使缓存条目频繁失效,这可能比携带过时状态消息成本更高。**当轨迹较短或单个状态消息较大**(例如完整的待办事项列表加上环境快照),**选择实现方式1**。最后几轮的缓存失效成本较低,上下文保持干净明确。" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853909232803ea4fa82759f0cae4f2ffa723d4f9dd6cef875", "model": "doubao-seed-1-6-250615", "created": 1785390937, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of the section title prefix '上下文工程[第13/17部分]' in the original English source. All technical content is accurately translated without additions or distortions.\"},\"fluency\":{\"score\":4,\"evidence\":\"Minor awkwardness in phrases like '框架借用' (X) vs. 'Agent框架借用' (Y), where Y adds clarity by retaining 'Agent'. Overall, sentences flow naturally with correct grammar.\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'token' as '标记' (X) instead of '词元' (Y), which is the more technically precise term in NLP contexts. Other terms like 'KV Cache' and 'CLI/GUI' are consistently translated.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"All Markdown elements (headings, image links, code blocks, bold text) are preserved exactly. The code block formatting and tags are identical to the source.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, additions, or changes to technical claims. All content, including the section title prefix, is fully translated.\"},\"fluency\":{\"score\":5,\"evidence\":\"Smoother phrasing with explicit references like 'Agent框架' and natural expressions such as '新词元' (new tokens) instead of '标记' (markers), enhancing readability.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and technically correct terminology: 'KV Cache' translated as '键值缓存', 'token' as '词元', and 'Agent framework' explicitly retained as 'Agent框架'.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Perfect preservation of Markdown structure, including headings, image syntax, code fences, and bolded terms. Code block content and formatting are identical to the source.\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y demonstrates higher accuracy by including the section title prefix, superior fluency with natural phrasing, and more precise terminology (e.g., '词元' for 'token'). Both maintain full markdown_code_fidelity, but Y's consistent technical terminology and complete content translation make it preferable.\"}", "usage": { "prompt_tokens": 4012, "completion_tokens": 470, "total_tokens": 4482 } }, "latency_milliseconds": 13819, "validation": { "valid": true } } ] }