{ "schema_version": 1, "credential_free": true, "attempts": [ { "attempt": 1, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n> **Experiment 2-3 ★★: Common but Harmful Context Management Patterns**\n>\n> In the `kv-cache` experiment, we systematically tested several common but harmful context management patterns. These patterns undermine KV Cache effectiveness, and some also impair the Agent's core capabilities.\n>\n> **Dynamic System Prompt** is one of the most common mistakes. Some developers embed timestamps in the system prompt (e.g., \"Current time: 2025-09-14 10:30:45.123456\") to let the Agent \"know\" the current time. While this seems to provide useful context, the timestamp changes with every request, making the entire system prompt different and completely invalidating the KV Cache. The correct approach is to append time information as part of a user message at the end of the conversation, or only obtain it through a tool call when truly needed.\n>\n> **Dynamic User Configuration** attempts to update user status information (such as remaining API calls or account balance) with each request. Embedding this information in the context destroys the cache. A better solution is to handle it through a dedicated state management mechanism when needed.\n>\n> **Dynamic Sorting of Tool Definitions** is another subtle trap. Some systems dynamically reorder tools based on usage frequency, but tool definitions often occupy a large portion of the context (each tool may contain hundreds of tokens of descriptions and parameter specifications). Changing the order invalidates the entire cache. Experiments show that a fixed order has almost no effect on tool-selection accuracy but substantially improves performance.\n>\n> **Sliding Window Conversation History** controls context length by retaining only the most recent messages. For example, if the window size is set to 10 messages, the earliest message is discarded when the 11th message arrives. This approach has two serious problems. First, it breaks prefix consistency and invalidates the KV Cache. Second, it may discard critical tool results. For example, with a sliding window of 10 rounds, if the Agent reads an important file in round 2, it may need that result again by round 15 — but the original result has already fallen out of the window. The model then has to infer from an incomplete conversation, which increases the error rate. In experiments, Agents using sliding windows often fell into loops, repeatedly executing the same tool calls because earlier results had been removed.\n>\n> **Text Formatting Method** is one of the most harmful patterns. It converts structured role-content messages into a plain text stream such as \"USER: ... ASSISTANT: ...\". The key issue is not caching: caching operates on the byte sequence of tokens, so a byte-stable concatenated prefix can still hit the cache. The cache is only broken when the concatenation method itself is unstable, such as when dynamic content is injected into the prefix each time. The real damage is that text formatting deviates from the standard message format used during model training. The model has seen large amounts of role-based dialogue data and has learned to parse that structure. When messages are flattened into plain text, the model must infer role boundaries and dialogue structure from weaker signals, leading to problems such as repeated operations, ignored tool results, text responses when a tool call is required, and parsing errors.\n>\n> **Summary**: The remedies for these harmful patterns all return to the three principles stated at the beginning of this section. One additional point: model providers have optimized heavily for their standard interfaces, and deviating from the standard format is likely to cause problems. As noted above, this is primarily a model capability issue rather than a caching issue.\n\n### KV Cache and Prompt Cache: Two Levels of Caching\n\nBefore proceeding, it is useful to distinguish two easily confused concepts. **KV Cache** is an optimization inside model inference: during a single inference pass, it caches the key-value states of already processed tokens to avoid redundant computation. **Prompt Cache** is an API service-layer optimization: it reuses cached computation for identical prefixes across multiple API requests. Both rely on prefix stability, but they operate at different levels. KV Cache accelerates token generation within a request; Prompt Cache reduces redundant prefix computation across requests. In practice, the API provider matches the request prefix. If multiple requests share the same prefix (for example, the system prompt and tool definitions remain unchanged), the provider can reuse cached prefix computation instead of recomputing those tokens. Reading from the cache costs far less than computing fresh—about one-tenth the price at Anthropic and DeepSeek, and likewise about one-tenth for OpenAI's GPT-5 family (the earlier GPT-4o generation was half price; starting with GPT-5.6, cache writes additionally carry a 1.25× surcharge). How caching is enabled and billed differs by provider: Anthropic requires explicit `cache_control` breakpoints, charges a markup for cache writes, enforces a minimum cacheable length (e.g., 1024 tokens), and applies a TTL limit (about 5 minutes by default); OpenAI uses automatic prefix caching without explicit declaration.\n\nWhen designing context, both levels of caching require a stable prefix—but Prompt Cache has a greater economic impact because it directly affects API billing.\n\n### Caching as an Architectural Constraint\n\nThe following section covers architectural details of production-grade Agents. First-time readers may skip it and return when building an Agent.\n\nIn production-grade Agent systems, caching is not merely a performance optimization—it is an **architectural constraint** that dictates many seemingly unrelated design decisions throughout the system.\n\nClaude Code illustrates a broader pattern: when Prompt Cache has significant economic value, cache consistency can shape architectural choices across the system. Several design decisions reflect this constraint:\n\n**Prompt structure is shaped by cache boundaries.** The system prompt is split by a cache boundary marker: content before the marker can be globally cached across users and sessions, while content after the marker contains user- and session-specific information. This means prompt ordering is driven primarily by caching economics and only secondarily by semantic logic. Each runtime condition placed before the cache boundary (OS type, current mode, user preferences, etc.) increases the number of cache-key variants. If each condition is binary, N conditions produce 2^N combinations. For example, 3 binary conditions (macOS/Linux, normal/debug mode, Chinese/English) produce 2×2×2 = 8 cache keys. Prompt fragments are therefore typed as either \"cacheable\" or \"cache-breaking,\" with explicit warning markers for the latter.\n\n**Sub-agents must be byte-aligned with the parent Agent.** When the main Agent spawns a sub-agent or performs a side query, the sub-agent's prompt, tool definitions, model configuration, message prefix, and reasoning configuration must match the parent Agent's cache key byte-for-byte. The reason is that if the API request initiated by the sub-agent has a prefix identical to the parent Agent's request, it can hit the API provider's Prompt Cache, thereby reducing billing and latency. This constraint propagates upward from the caching layer, influencing how Agents are generated and how parameters are passed.\n\n**Replacement strings for tool results are frozen upon first occurrence.** When large tool outputs are replaced with summary previews, the replacement string is persisted. Even after a session restarts, the system reuses exactly the same replacement string so that the restored message sequence remains byte-identical to the cached stream.\n\nThe core insight is that **caching economics is not a post-hoc optimization but an upfront architectural constraint.** If your Agent system uses Prompt Caching, the requirement for cache key consistency will permeate prompt design, multi-agent coordination, session restoration, and other layers. The earlier this constraint is incorporated into the architecture, the lower the subsequent engineering cost.\n\n### KV Cache Is Not Necessarily One-Shot: Editable, Composable \"Notes\"\n\n(The following is optional advanced material from current research. It can be skipped on first reading without affecting the rest of this chapter; the three practical conclusions above are the foundation.)\n\nSo far, this section has assumed a strict rule: change one byte in the prefix, and the subsequent cache is invalidated. This rule holds in today's inference engines, but it may not be inevitable. A recent line of research starts from a counterintuitive observation[^ch2-2]: during the prefill phase, the model behaves as if it is \"taking notes.\" When it reads a field in the context (e.g., \"User's city: Beijing\"), it does not simply cache that field verbatim. Instead, it writes downstream representations of the **conclusion**—what this field means—into later KV states. Measurements show that the KV states of the field's **own** tokens often contribute less than 1% to the final decision; what influences the output more are the downstream \"notes\" left by that field.\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第6/17部分]\n\n#### 实验2-3 ★★:常见但有害的上下文管理模式\n\n在`kv-cache`实验中,我们系统地测试了几种常见但有害的上下文管理模式。这些模式削弱了KV缓存的有效性,有些还损害了代理的核心能力。\n\n**动态系统提示**是最常见的错误之一。一些开发人员在系统提示中嵌入时间戳(例如,“当前时间:2025-09-14 10:30:45.123456”)以让代理“知道”当前时间。虽然这似乎提供了有用的上下文,但时间戳随每个请求而变化,使整个系统提示不同,完全使KV缓存失效。正确的方法是在对话末尾将时间信息作为用户消息的一部分附加,或者仅在真正需要时通过工具调用获取它。\n\n**动态用户配置**试图在每个请求时更新用户状态信息(如剩余API调用次数或账户余额)。将此信息嵌入上下文中会破坏缓存。更好的解决方案是在需要时通过专用状态管理机制来处理。\n\n**工具定义的动态排序**是另一个微妙的陷阱。一些系统根据使用频率动态重新排序工具,但工具定义通常占据上下文的很大一部分(每个工具可能包含数百个词元的描述和参数规范)。更改顺序会使整个缓存失效。实验表明,固定顺序对工具选择准确性几乎没有影响,但大大提高了性能。\n\n**滑动窗口对话历史**通过仅保留最近的消息来控制上下文长度。例如,如果窗口大小设置为10条消息,当第11条消息到达时,最早的消息将被丢弃。这种方法有两个严重问题。首先,它破坏了前缀一致性,使KV缓存失效。其次,它可能丢弃关键的工具结果。例如,使用10轮的滑动窗口,如果代理在第2轮读取了一个重要文件,它可能在第15轮时再次需要该结果——但原始结果已经超出了窗口。然后模型必须从不完整的对话中推断,这会增加错误率。在实验中,使用滑动窗口的代理经常陷入循环,反复执行相同的工具调用,因为早期的结果已被删除。\n\n**文本格式化方法**是最有害的模式之一。它将结构化的角色-内容消息转换为纯文本流,如“用户:…… 助手:……”。关键问题不是缓存:缓存是在词元的字节序列上操作的,因此字节稳定的连接前缀仍然可以命中缓存。只有当连接方法本身不稳定时,缓存才会被破坏,例如每次向前缀中注入动态内容时。真正的损害是文本格式化偏离了模型训练期间使用的标准消息格式。模型已经看到了大量基于角色的对话数据,并学会了解析该结构。当消息被展平为纯文本时,模型必须从较弱的信号中推断角色边界和对话结构,导致重复操作、忽略工具结果、需要工具调用时的文本响应和解析错误等问题。\n\n**总结**:这些有害模式的补救措施都回到本节开头所述的三个原则。另外一点:模型提供商针对其标准接口进行了大量优化,偏离标准格式可能会导致问题。如上所述,这主要是模型能力问题,而不是缓存问题。\n\n### KV缓存和提示缓存:两级缓存\n\n在继续之前,区分两个容易混淆的概念很有用。**KV缓存**是模型推理内的优化:在单个推理过程中,它缓存已处理词元的键值状态以避免冗余计算。**提示缓存**是API服务层的优化:它在多个API请求中对相同前缀重用缓存的计算。两者都依赖于前缀稳定性,但在不同层次上操作。KV缓存加速请求内的词元生成;提示缓存减少请求间的冗余前缀计算。在实践中,API提供商匹配请求前缀。如果多个请求共享相同的前缀(例如,系统提示和工具定义保持不变),提供商可以重用缓存的前缀计算而不是重新计算那些词元。从缓存读取的成本远低于重新计算的成本——在Anthropic和DeepSeek约为十分之一,OpenAI的GPT-5系列也约为十分之一(早期的GPT-4o一代是一半价格;从GPT-5.6开始,缓存写入额外收取1.25倍的附加费)。缓存如何启用和计费因提供商而异:Anthropic要求显式的`cache_control`断点,对缓存写入收取加价,强制执行最小可缓存长度(例如1024词元),并应用TTL限制(默认约5分钟);OpenAI使用自动前缀缓存,无需显式声明。\n\n在设计上下文时,两级缓存都需要稳定的前缀——但提示缓存具有更大的经济影响,因为它直接影响API计费。\n\n### 缓存作为架构约束\n\n以下部分涵盖生产级代理的架构细节。首次阅读的读者可以跳过,在构建代理时再回来。\n\n在生产级代理系统中,缓存不仅仅是性能优化——它是一个**架构约束**,规定了整个系统中许多看似无关的设计决策。\n\nClaude Code说明了更广泛的模式:当提示缓存具有重大经济价值时,缓存一致性可以塑造整个系统的架构选择。几个设计决策反映了这一约束:\n\n**提示结构由缓存边界塑造**。系统提示由缓存边界标记分割:标记之前的内容可以在用户和会话之间全局缓存,而标记之后的内容包含用户和会话特定的信息。这意味着提示顺序主要由缓存经济性驱动,其次才是语义逻辑。放置在缓存边界之前的每个运行时条件(操作系统类型、当前模式、用户偏好等)都会增加缓存键变体的数量。如果每个条件是二进制的,N个条件产生2^N种组合。例如,3个二进制条件(macOS/Linux、正常/调试模式、中文/英文)产生2×2×2=8个缓存键。因此,提示片段被类型化为“可缓存”或“破坏缓存”,后者有明确的警告标记。\n\n**子代理必须与父代理字节对齐**。当主代理生成子代理或执行辅助查询时,子代理的提示、工具定义、模型配置、消息前缀和推理配置必须与父代理的缓存键逐字节匹配。原因是如果子代理发起的API请求的前缀与父代理的请求相同,它可以命中API提供商的提示缓存,从而减少计费和时延。这一约束从缓存层向上传播,影响代理的生成方式和参数传递方式。\n\n**工具结果的替换字符串在首次出现时冻结**。当大型工具输出被替换为摘要预览时,替换字符串被持久化。即使会话重新启动,系统也会精确重用相同的替换字符串,以便恢复的消息序列与缓存流逐字节相同。\n\n核心见解是**缓存经济性不是事后优化,而是预先的架构约束**。如果你的代理系统使用提示缓存,缓存键一致性的要求将渗透到提示设计、多代理协调、会话恢复等层。越早将这一约束纳入架构,后续的工程成本越低。\n\n### KV缓存不一定是一次性的:可编辑、可组合的“笔记”\n\n(以下是当前研究的可选高级材料。首次阅读时可以跳过,不影响本章其余部分;上述三个实际结论是基础。)\n\n到目前为止,本节假设了一个严格的规则:前缀中改变一个字节,后续的缓存就会失效。这个规则在当今的推理引擎中成立,但不一定是不可避免的。最近的一系列研究从一个反直觉的观察[^ch2-2]开始:在预填充阶段,模型的行为就像在“做笔记”。当它读取上下文中的一个字段(例如,“用户的城市:北京”)时,它不会简单地逐字缓存该字段。相反,它将该字段的**结论**——这个字段意味着什么——写入后续的KV状态中。测量表明,该字段**自身**词元的KV状态通常对最终决策的贡献不到1%;对输出影响更大的是该字段留下的后续“笔记”。\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第6/17部分]\n> **实验2-3 ★★:常见但有害的上下文管理模式**\n>\n> 在`kv-cache`实验中,我们系统地测试了几种常见但有害的上下文管理模式。这些模式破坏了KV缓存的有效性,有些还损害了代理的核心能力。\n>\n> **动态系统提示**是最常见的错误之一。一些开发者在系统提示中嵌入时间戳(例如“当前时间:2025-09-14 10:30:45.123456”),让代理“知道”当前时间。虽然这似乎提供了有用的上下文,但时间戳随每次请求而变化,使整个系统提示不同,完全使KV缓存失效。正确的做法是将时间信息作为用户消息的一部分追加到对话末尾,或仅在真正需要时通过工具调用获取。\n>\n> **动态用户配置**试图在每次请求时更新用户状态信息(例如剩余API调用次数或账户余额)。将此信息嵌入上下文中会破坏缓存。更好的解决方案是在需要时通过专门的状态管理机制处理。\n>\n> **工具定义的动态排序**是另一个微妙的陷阱。一些系统根据使用频率动态重新排序工具,但工具定义通常占据上下文的很大一部分(每个工具可能包含数百个标记的描述和参数规范)。更改顺序会使整个缓存失效。实验表明,固定顺序对工具选择准确性几乎没有影响,但显著提高性能。\n>\n> **滑动窗口对话历史**通过仅保留最近的消息来控制上下文长度。例如,如果窗口大小设置为10条消息,第11条消息到达时最早的消息被丢弃。这种方法有两个严重问题。首先,它破坏了前缀一致性,使KV缓存失效。其次,可能丢弃关键工具结果。例如,使用10轮的滑动窗口,如果代理在第2轮读取了一个重要文件,它可能在第15轮再次需要该结果——但原始结果已经超出窗口。然后模型必须从不完整的对话中推断,这会增加错误率。实验中,使用滑动窗口的代理经常陷入循环,反复执行相同的工具调用,因为早期结果已被删除。\n>\n> **文本格式化方法**是最有害的模式之一。它将结构化的角色-内容消息转换为纯文本流,例如“用户:... 助手:...”。关键问题不是缓存:缓存基于标记的字节序列操作,因此字节稳定的连接前缀仍然可以命中缓存。缓存仅在连接方法本身不稳定时被破坏,例如每次将动态内容注入前缀时。真正的损害是文本格式化偏离了模型训练期间使用的标准消息格式。模型看到了大量基于角色的对话数据,并学会了解析该结构。当消息被展平为纯文本时,模型必须从较弱的信号中推断角色边界和对话结构,导致重复操作、忽略工具结果、需要工具调用时返回文本响应、解析错误等问题。\n>\n> **总结**:这些有害模式的补救措施都回到本节开头所述的三个原则。还有一点:模型提供商针对其标准接口进行了大量优化,偏离标准格式可能会导致问题。如上所述,这主要是模型能力问题,而不是缓存问题。\n\n### KV缓存和提示缓存:两级缓存\n在继续之前,区分两个容易混淆的概念很有用。**KV缓存**是模型推理内部的优化:在单次推理过程中,它缓存已处理标记的键值状态以避免冗余计算。**提示缓存**是API服务层的优化:它在多个API请求之间重用相同前缀的缓存计算。两者都依赖前缀稳定性,但在不同层次上操作。KV缓存加速请求内的标记生成;提示缓存减少请求之间的冗余前缀计算。实际上,API提供商匹配请求前缀。如果多个请求共享相同的前缀(例如,系统提示和工具定义保持不变),提供商可以重用缓存的前缀计算而不是重新计算这些标记。从缓存读取的成本远低于重新计算——在Anthropic和DeepSeek约为十分之一,OpenAI的GPT-5系列也是如此(早期的GPT-4o一代是半价;从GPT-5.6开始,缓存写入额外收取1.25倍的附加费)。缓存如何启用和计费因提供商而异:Anthropic需要显式的`cache_control`断点,对缓存写入收取加价,强制执行最小可缓存长度(例如1024个标记),并应用TTL限制(默认约5分钟);OpenAI使用自动前缀缓存,无需显式声明。\n\n设计上下文时,两级缓存都需要稳定的前缀——但提示缓存对经济影响更大,因为它直接影响API计费。\n\n### 缓存作为架构约束\n以下部分涵盖生产级代理的架构细节。首次阅读的读者可以跳过,在构建代理时再回来。\n\n在生产级代理系统中,缓存不仅仅是性能优化——它是一个**架构约束**,决定了系统中许多看似无关的设计决策。\n\nClaude Code展示了更广泛的模式:当提示缓存具有显著的经济价值时,缓存一致性可以塑造系统中的架构选择。几个设计决策反映了这一约束:\n\n**提示结构由缓存边界塑造**。系统提示通过缓存边界标记分割:标记前的内容可以在用户和会话之间全局缓存,而标记后的内容包含用户和会话特定的信息。这意味着提示顺序主要由缓存经济性驱动,其次才是语义逻辑。放置在缓存边界之前的每个运行时条件(操作系统类型、当前模式、用户偏好等)都会增加缓存键变体的数量。如果每个条件是二进制的,N个条件产生2^N种组合。例如,3个二进制条件(macOS/Linux、正常/调试模式、中文/英文)产生2×2×2=8个缓存键。因此,提示片段被类型化为“可缓存”或“破坏缓存”,后者带有明确的警告标记。\n\n**子代理必须与父代理字节对齐**。当主代理生成子代理或执行旁查询时,子代理的提示、工具定义、模型配置、消息前缀和推理配置必须与父代理的缓存键字节完全匹配。原因是如果子代理发起的API请求具有与父代理请求相同的前缀,它可以命中API提供商的提示缓存,从而降低计费和延迟。此约束从缓存层向上传播,影响代理的生成方式和参数传递方式。\n\n**工具结果的替换字符串在首次出现时冻结**。当大型工具输出被替换为摘要预览时,替换字符串被持久化。即使会话重新启动,系统也会精确重用相同的替换字符串,以便恢复的消息序列与缓存流字节完全相同。\n\n核心洞察是,**缓存经济性不是事后优化,而是前期架构约束**。如果你的代理系统使用提示缓存,缓存键一致性的要求将渗透到提示设计、多代理协调、会话恢复等层。越早将此约束纳入架构,后续工程成本越低。\n\n### KV缓存不一定是单次的:可编辑、可组合的“笔记”\n(以下是当前研究的可选高级材料。首次阅读时可以跳过,不影响本章其余内容;上述三个实践结论是基础。)\n\n到目前为止,本节假设了一个严格规则:前缀中更改一个字节,后续缓存就会失效。这个规则在当今的推理引擎中成立,但不一定是不可避免的。最近的一系列研究从一个反直觉的观察开始[^ch2-2]:在预填充阶段,模型的行为就像在“做笔记”。当它读取上下文中的一个字段(例如“用户的城市:北京”)时,它不会简单地逐字缓存该字段。相反,它将该字段的**结论**的下游表示写入后续的KV状态。测量表明,该字段**自身**标记的KV状态通常对最终决策的贡献不到1%;对输出影响更大的是该字段留下的下游“笔记”。" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853907800793ea4fa82759f0cae4f2ffa723d4f9dd6640636", "model": "doubao-seed-1-6-250615", "created": 1785390796, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, 'the correct approach is to append time information as part of a user message at the end of the conversation, or only obtain it through a tool call when truly needed' is accurately translated as '正确的方法是在对话末尾将时间信息作为用户消息的一部分附加,或者仅在真正需要时通过工具调用获取它'.\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural and smooth Chinese expression. For instance, 'This means prompt ordering is driven primarily by caching economics and only secondarily by semantic logic' is translated as '这意味着提示顺序主要由缓存经济性驱动,其次才是语义逻辑' with proper word order and collocation.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and technically correct terminology. 'KV Cache' is consistently 'KV缓存', 'Prompt Cache' as '提示缓存', 'token' as '词元' throughout.\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Headings, code blocks (e.g., `kv-cache`), and equations are preserved. However, the blockquote formatting for the experiment section in the original English is not retained in X; X uses '####' heading instead of blockquote for 'Experiment 2-3 ★★: Common but Harmful Context Management Patterns'.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, 'the model must infer role boundaries and dialogue structure from weaker signals, leading to problems such as repeated operations...' is accurately translated as '模型必须从较弱的信号中推断角色边界和对话结构,导致重复操作...等问题'.\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent but has minor awkwardness. For example, 'sub-agents must be byte-aligned with the parent Agent' is translated as '子代理必须与父代理字节对齐' which is correct but slightly less natural than X's '子代理必须与父代理字节对齐' (same wording here, but other instances like '旁查询' for 'side query' in Y is less standard than X's '辅助查询').\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent terminology. 'KV Cache' as 'KV缓存', 'Prompt Cache' as '提示缓存', 'token' as '标记' (consistent choice, though different from X's '词元', both are technically acceptable).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Blockquote formatting for the experiment section is preserved with '> ' prefixes, headings (###), code blocks (e.g., `kv-cache`), and footnotes ([^ch2-2]) are all correctly maintained.\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher markdown_code_fidelity (5 vs. 4) by preserving the original blockquote formatting for the experiment section, while both have equal accuracy and terminology scores. Y's fluency is slightly lower than X's but still acceptable, and the markdown fidelity advantage makes Y preferred.\"}", "usage": { "prompt_tokens": 5594, "completion_tokens": 656, "total_tokens": 6250 } }, "latency_milliseconds": 16581, "validation": { "valid": true } } ] }