{ "schema_version": 1, "credential_free": true, "attempts": [ { "attempt": 1, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Second, improving reasoning quality—summarized knowledge is more useful to the model than raw information.** This motivation is deeper and easier to overlook. Even if the context window is large enough, adding all raw information to the context is not always the best choice.\n\nConsider a concrete example: during a complex task, an Agent accumulates information on a topic through 10 web searches. These search results are scattered in their raw form throughout the context—the results from round 2 are near the beginning, and the results from round 9 are near the end. When the Agent must make a final decision from all this information, it has to retrieve relevant fragments scattered across tens of thousands of tokens. Its attention becomes diffuse, and it can easily miss key information.\n\nAfter the 10th search, however, a single LLM call could produce a structured summary of the accumulated information: \"Currently known: A is..., B is..., information on C is still missing.\" The model can then use this refined knowledge representation in subsequent reasoning, without re-extracting it from the raw data.\n\nThe root cause lies in the nature of the attention mechanism: **the internal mechanism of in-context learning is more like retrieval than reasoning**. Chapter 1 introduced this concept briefly, and the Agent Status Bar section expanded it through mechanisms, empirical evidence, and engineering practices. Next, we examine what this means for compression.\n\n### The Internal Mechanism of In-Context Learning: Retrieval, Not Reasoning\n\nBriefly, **retrieval, not reasoning** means that attention is good at looking up existing content, but not at actively computing aggregate summaries in a single forward pass. This does not deny that the model can reason step by step by generating a chain of thought; it means that consuming existing context in one forward pass is more retrieval-like. The implication for compression is clear: the Status Bar **adds** computed conclusions **into** the context, while compression **replaces** bloated raw records **with** computed conclusions. Both supply the distillation layer that raw attention lacks. The difference is that the Status Bar is usually maintained deterministically, step by step, by **code**, while compression more often uses an LLM call to distill a large block of original text.\n\nA simple example makes the idea of \"retrieval, not reasoning\" concrete. Suppose the context contains a log of a pet store inspection:\n\n> Cage 1: Black cat. Cage 2: White cat. Cage 3: Black cat. Cage 4: Black cat. Cage 5: White cat.\n> ... (100 cages total, 90 black cats, 10 white cats)\n\nWhen you ask the model, \"How many black cats and white cats are there?\" what happens?\n\nIf reasoning is not enabled, the model will find it difficult to give the correct answer directly—because the attention mechanism is good at **looking up** (\"What cat is in cage 37?\"), not **aggregation** (\"How many black cats are there in total?\"). The latter requires traversing all records and maintaining a counting state, which is essentially reasoning, not retrieval.\n\nIf reasoning is enabled, the model can obtain the correct answer by counting one by one. The cost is that every time this question is asked, it must start counting from scratch, generating many reasoning tokens. In an Agent scenario, if such statistical information needs to be used repeatedly (e.g., for every decision), the cumulative reasoning cost becomes very high.\n\nHowever, if we summarize the records in advance and write \"Current statistics: 90 black cats, 10 white cats\" directly into the context, the model can retrieve the conclusion without repeating the count. **This is the second value of compression: turning conclusions that require reasoning into knowledge that can be directly retrieved.**\n\nThe deeper issue is that long contexts reduce retrieval precision. Even when the context window is far from full, the Agent may suddenly fail to find key information or repeatedly focus on a problem that has already been solved. This phenomenon is known as **Context Rot**. Context rot is different from context overflow (running out of window space): overflow means \"cannot fit any more,\" while rot means \"it fits but cannot be found.\" The latter is more insidious because the Agent appears to be working normally, while the quality of its decisions quietly deteriorates. As context length increases, attention weights are spread across more tokens, reducing the weight each token receives. More importantly, once irrelevant content dominates the context, the Agent's decision quality declines. In practice, the most common failure mode is not a context window that is too small, but an information density that is too low: knowledge needed only occasionally is loaded every time, stable rules are mixed with dynamic state, and the model sees more content while the useful parts become harder to notice. A useful analogy is searching for one book in a large library: the more irrelevant books on the shelves, the harder it is to find the target. The attention visualization in Experiment 2-2 demonstrates this phenomenon clearly: in long contexts, the model's attention exhibits strong positional bias. This is the problem revealed by the famous \"Needle in a Haystack\" experiment, which hides a key piece of information in the middle of a very long text and tests whether the model can find it.\n\nAndrej Karpathy offered a profound insight: the model's \"poor memory\" is, to some extent, a feature rather than a bug—the limited context window forces the model to learn to abstract general patterns from a large amount of detail, just as humans do not remember the verbatim content of every conversation but distill an overall impression and behavioral patterns.\n\nThis reveals the design principle of context compression: rather than expecting the model to learn automatically from lengthy context, we should distill that knowledge explicitly. Although this requires additional computation for summarization, it produces compact, information-dense representations. **Do not make the model search passively through vast amounts of raw material; provide refined, structured knowledge instead.**\n\nFrom this perspective, in-context learning is more like a rapid adaptation mechanism than true learning. It allows the model to quickly adjust its behavior during inference to suit a specific task, but this adjustment is temporary and shallow, disappearing after the session ends. Recent theoretical research[^ch2-6] supports this judgment: when the model sees examples in the context, its behavior is as if it has been \"temporarily customized\"—without changing the model parameters, but with an effect similar to a small, specialized training session. This explains why few-shot examples in the prompt engineering section can significantly improve output quality, and also why this improvement does not accumulate across sessions—it is fundamentally different from true parameter training.\n\n[^ch2-6]: Benoit Dherin et al., \"Learning without training\", 2025.\n\n### Compression and KV Cache: Apparent Contradiction, Practical Complementarity\n\nBefore discussing specific compression strategies, we need to resolve an apparent contradiction: earlier sections emphasized that KV Cache requires the context prefix to remain unchanged, but compression involves modifying content in the middle of the context.\n\nThe key is understanding the **timing and location** of compression. Compression does not modify the context during a single API call; instead, it occurs **between two API calls**, when the Agent framework preprocesses the message list:\n\n1. **System Prompt and Tool Definitions are never touched**—this is the \"static prefix\" at the very front of the context, and the KV Cache is continuously cached.\n2. **The target of compression is the tool results in the conversation history**—when the Agent framework replaces the original tool output with a compressed summary, the cache after the replacement point becomes invalid, but the cache before it remains valid.\n3. **This is a conscious trade-off**: without compression, the context expands beyond the window limit and the task fails outright; with it, some cache is lost, but context length stays under control and information density rises. Therefore, the frequency of compression needs to be weighed—frequent compression will frequently break the cache. It is best to perform batch compression when the context approaches the threshold, rather than compressing every round.\n\n![Figure 2-16: Comparison of Context Compression Strategies](images/fig2-16.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第15/17部分]\n**其次,提高推理质量——总结的知识对模型比原始信息更有用**。这个动机更深层且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中并不总是最佳选择。\n\n考虑一个具体示例:在复杂任务中,代理通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在末尾附近。当代理必须从所有这些信息中做出最终决策时,它必须从分散在数万标记中的相关片段中检索。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索后,一次大语言模型调用可以生成累积信息的结构化摘要:“目前已知:A是……,B是……,关于C的信息仍缺失。” 然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的性质:**上下文中学习的内部机制更像是检索而不是推理**。第1章简要介绍了这个概念,代理状态栏部分通过机制、经验证据和工程实践进行了扩展。接下来,我们探讨这对压缩意味着什么。\n\n### 上下文中学习的内部机制:检索,而非推理\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在单次前向传递中主动计算聚合摘要。这并不否定模型可以通过生成思维链逐步推理;它意味着在单次前向传递中消耗现有上下文更像是检索。对压缩的影响很明显:状态栏**将计算出的结论添加到**上下文中,而压缩**用计算出的结论替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。区别在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用提炼一大块原始文本。\n\n一个简单的示例使“检索,而非推理”的想法具体化。假设上下文中包含宠物商店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> ...(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型:“有多少只黑猫和白猫?” 会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而不是检索。\n\n如果启用推理,模型可以通过逐个计数得到正确答案。成本是每次问这个问题时,都必须从头开始计数,生成许多推理标记。在代理场景中,如果这种统计信息需要反复使用(例如,用于每个决策),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录并直接将“当前统计:90只黑猫,10只白猫”写入上下文中,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识。**\n\n更深层的问题是长上下文降低了检索精度。即使上下文窗口远未填满,代理可能突然无法找到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文腐烂**。上下文腐烂不同于上下文溢出(窗口空间耗尽):溢出意味着“无法再容纳更多”,而腐烂意味着“容纳了但找不到”。后者更隐蔽,因为代理看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分布在更多标记上,每个标记获得的权重减少。更重要的是,一旦不相关内容主导上下文,代理的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分更难被注意到。一个有用的类比是在大型图书馆中寻找一本书:书架上不相关的书越多,找到目标就越难。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西提供了一个深刻的见解:模型的“记忆力差”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型自动从冗长的上下文中学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但它产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的结构化知识。**\n\n从这个角度看,上下文中学习更像是一种快速适应机制,而不是真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是暂时和肤浅的,会话结束后消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就像被“临时定制”了一样——不改变模型参数,但效果类似于一次小型的专门训练会话。这解释了为什么提示工程中的少样本示例可以显著提高输出质量,也解释了为什么这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n### 压缩与KV缓存:表面矛盾,实际互补\n在讨论具体的压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调KV缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的内容。\n\n关键是理解压缩的**时间和位置**。压缩不会在单次API调用期间修改上下文;相反,它发生在**两次API调用之间**,当代理框架预处理消息列表时:\n\n1. **系统提示和工具定义从不被触及**——这是上下文中最前面的“静态前缀”,KV缓存持续缓存。\n2. **压缩的目标是对话历史中的工具结果**——当代理框架将原始工具输出替换为压缩摘要时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文扩展超出窗口限制,任务完全失败;有压缩,一些缓存丢失,但上下文长度得到控制,信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第15/17部分]\n\n**其次,提高推理质量——总结的知识对模型来说比原始信息更有用**。这种动机更深刻且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中也不总是最佳选择。\n\n考虑一个具体例子:在一项复杂任务中,一个智能体通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在结尾附近。当智能体必须从所有这些信息中做出最终决策时,它必须检索分散在数万词元中的相关片段。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索之后,单个大语言模型调用可以生成积累信息的结构化总结:“目前已知:A是……,B是……,关于C的信息仍缺失。”然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的本质:**上下文学习的内部机制更像是检索而非推理**。第1章简要介绍了这个概念,“智能体状态栏”部分通过机制、实证证据和工程实践对其进行了扩展。接下来,我们研究这对压缩意味着什么。\n\n#### 上下文学习的内部机制:检索,而非推理\n\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在一次前向传递中主动计算聚合总结。这并不否认模型可以通过生成思维链逐步推理;而是意味着在一次前向传递中消耗现有上下文更像是检索。这对压缩的含义很明确:状态栏将计算出的结论**添加到**上下文中,而压缩用计算出的结论**替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。不同之处在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用来提炼一大块原始文本。\n\n一个简单的例子使“检索,而非推理”的概念具体化。假设上下文中包含宠物店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> …(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型,“有多少只黑猫和白猫?”会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而非检索。\n\n如果启用推理,模型可以通过逐一计数得到正确答案。代价是每次问这个问题时,都必须从头开始计数,生成许多推理词元。在智能体场景中,如果此类统计信息需要反复使用(例如每次决策都需要),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录,并直接将“当前统计:90只黑猫,10只白猫”写入上下文,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识**。\n\n更深层的问题是长上下文会降低检索精度。即使上下文窗口远未填满,智能体也可能突然找不到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文旋转**。上下文旋转不同于上下文溢出(窗口空间耗尽):溢出意味着“再也装不下了”,而旋转意味着“装得下但找不到”。后者更隐蔽,因为智能体看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分散到更多词元上,每个词元获得的权重降低。更重要的是,一旦不相关内容主导上下文,智能体的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分却更难被注意到。一个有用的类比是在大图书馆中寻找一本书:书架上不相关的书越多,越难找到目标。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将一条关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西(Andrej Karpathy)提出了一个深刻的见解:模型的“记忆不佳”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼出整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型从冗长的上下文中自动学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但会产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的、结构化的知识**。\n\n从这个角度看,上下文学习更像是一种快速适应机制,而非真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是临时且肤浅的,会话结束后就会消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就好像被“临时定制”了——没有改变模型参数,但效果类似于一次小型的专门训练会话。这解释了提示工程部分中的少样本示例为何能显著提高输出质量,也解释了为何这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n#### 压缩与键值缓存:表面矛盾,实际互补\n\n在讨论具体压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调键值缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的中间内容。\n\n关键是理解压缩的**时机和位置**。压缩不会在单个API调用期间修改上下文;相反,它发生在**两次API调用之间**,当智能体框架预处理消息列表时:\n\n1. **系统提示和工具定义永远不会被触碰**——这是上下文中最前面的“静态前缀”,键值缓存持续缓存。\n2. **压缩的目标是会话历史中的工具结果**——当智能体框架将原始工具输出替换为压缩总结时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文会超出窗口限制,任务彻底失败;有了压缩,一些缓存会丢失,但上下文长度得到控制且信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮都压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853909520823ea4fa82759f0cae4f2ffa723d4f9dd6d11bbf", "model": "doubao-seed-1-6-250615", "created": 1785390970, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } }", "usage": { "prompt_tokens": 5110, "completion_tokens": 583, "total_tokens": 5693 } }, "latency_milliseconds": 17971, "validation": { "valid": false, "error_type": "ValueError", "error": "judge variants must contain exactly X and Y" } }, { "attempt": 2, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Second, improving reasoning quality—summarized knowledge is more useful to the model than raw information.** This motivation is deeper and easier to overlook. Even if the context window is large enough, adding all raw information to the context is not always the best choice.\n\nConsider a concrete example: during a complex task, an Agent accumulates information on a topic through 10 web searches. These search results are scattered in their raw form throughout the context—the results from round 2 are near the beginning, and the results from round 9 are near the end. When the Agent must make a final decision from all this information, it has to retrieve relevant fragments scattered across tens of thousands of tokens. Its attention becomes diffuse, and it can easily miss key information.\n\nAfter the 10th search, however, a single LLM call could produce a structured summary of the accumulated information: \"Currently known: A is..., B is..., information on C is still missing.\" The model can then use this refined knowledge representation in subsequent reasoning, without re-extracting it from the raw data.\n\nThe root cause lies in the nature of the attention mechanism: **the internal mechanism of in-context learning is more like retrieval than reasoning**. Chapter 1 introduced this concept briefly, and the Agent Status Bar section expanded it through mechanisms, empirical evidence, and engineering practices. Next, we examine what this means for compression.\n\n### The Internal Mechanism of In-Context Learning: Retrieval, Not Reasoning\n\nBriefly, **retrieval, not reasoning** means that attention is good at looking up existing content, but not at actively computing aggregate summaries in a single forward pass. This does not deny that the model can reason step by step by generating a chain of thought; it means that consuming existing context in one forward pass is more retrieval-like. The implication for compression is clear: the Status Bar **adds** computed conclusions **into** the context, while compression **replaces** bloated raw records **with** computed conclusions. Both supply the distillation layer that raw attention lacks. The difference is that the Status Bar is usually maintained deterministically, step by step, by **code**, while compression more often uses an LLM call to distill a large block of original text.\n\nA simple example makes the idea of \"retrieval, not reasoning\" concrete. Suppose the context contains a log of a pet store inspection:\n\n> Cage 1: Black cat. Cage 2: White cat. Cage 3: Black cat. Cage 4: Black cat. Cage 5: White cat.\n> ... (100 cages total, 90 black cats, 10 white cats)\n\nWhen you ask the model, \"How many black cats and white cats are there?\" what happens?\n\nIf reasoning is not enabled, the model will find it difficult to give the correct answer directly—because the attention mechanism is good at **looking up** (\"What cat is in cage 37?\"), not **aggregation** (\"How many black cats are there in total?\"). The latter requires traversing all records and maintaining a counting state, which is essentially reasoning, not retrieval.\n\nIf reasoning is enabled, the model can obtain the correct answer by counting one by one. The cost is that every time this question is asked, it must start counting from scratch, generating many reasoning tokens. In an Agent scenario, if such statistical information needs to be used repeatedly (e.g., for every decision), the cumulative reasoning cost becomes very high.\n\nHowever, if we summarize the records in advance and write \"Current statistics: 90 black cats, 10 white cats\" directly into the context, the model can retrieve the conclusion without repeating the count. **This is the second value of compression: turning conclusions that require reasoning into knowledge that can be directly retrieved.**\n\nThe deeper issue is that long contexts reduce retrieval precision. Even when the context window is far from full, the Agent may suddenly fail to find key information or repeatedly focus on a problem that has already been solved. This phenomenon is known as **Context Rot**. Context rot is different from context overflow (running out of window space): overflow means \"cannot fit any more,\" while rot means \"it fits but cannot be found.\" The latter is more insidious because the Agent appears to be working normally, while the quality of its decisions quietly deteriorates. As context length increases, attention weights are spread across more tokens, reducing the weight each token receives. More importantly, once irrelevant content dominates the context, the Agent's decision quality declines. In practice, the most common failure mode is not a context window that is too small, but an information density that is too low: knowledge needed only occasionally is loaded every time, stable rules are mixed with dynamic state, and the model sees more content while the useful parts become harder to notice. A useful analogy is searching for one book in a large library: the more irrelevant books on the shelves, the harder it is to find the target. The attention visualization in Experiment 2-2 demonstrates this phenomenon clearly: in long contexts, the model's attention exhibits strong positional bias. This is the problem revealed by the famous \"Needle in a Haystack\" experiment, which hides a key piece of information in the middle of a very long text and tests whether the model can find it.\n\nAndrej Karpathy offered a profound insight: the model's \"poor memory\" is, to some extent, a feature rather than a bug—the limited context window forces the model to learn to abstract general patterns from a large amount of detail, just as humans do not remember the verbatim content of every conversation but distill an overall impression and behavioral patterns.\n\nThis reveals the design principle of context compression: rather than expecting the model to learn automatically from lengthy context, we should distill that knowledge explicitly. Although this requires additional computation for summarization, it produces compact, information-dense representations. **Do not make the model search passively through vast amounts of raw material; provide refined, structured knowledge instead.**\n\nFrom this perspective, in-context learning is more like a rapid adaptation mechanism than true learning. It allows the model to quickly adjust its behavior during inference to suit a specific task, but this adjustment is temporary and shallow, disappearing after the session ends. Recent theoretical research[^ch2-6] supports this judgment: when the model sees examples in the context, its behavior is as if it has been \"temporarily customized\"—without changing the model parameters, but with an effect similar to a small, specialized training session. This explains why few-shot examples in the prompt engineering section can significantly improve output quality, and also why this improvement does not accumulate across sessions—it is fundamentally different from true parameter training.\n\n[^ch2-6]: Benoit Dherin et al., \"Learning without training\", 2025.\n\n### Compression and KV Cache: Apparent Contradiction, Practical Complementarity\n\nBefore discussing specific compression strategies, we need to resolve an apparent contradiction: earlier sections emphasized that KV Cache requires the context prefix to remain unchanged, but compression involves modifying content in the middle of the context.\n\nThe key is understanding the **timing and location** of compression. Compression does not modify the context during a single API call; instead, it occurs **between two API calls**, when the Agent framework preprocesses the message list:\n\n1. **System Prompt and Tool Definitions are never touched**—this is the \"static prefix\" at the very front of the context, and the KV Cache is continuously cached.\n2. **The target of compression is the tool results in the conversation history**—when the Agent framework replaces the original tool output with a compressed summary, the cache after the replacement point becomes invalid, but the cache before it remains valid.\n3. **This is a conscious trade-off**: without compression, the context expands beyond the window limit and the task fails outright; with it, some cache is lost, but context length stays under control and information density rises. Therefore, the frequency of compression needs to be weighed—frequent compression will frequently break the cache. It is best to perform batch compression when the context approaches the threshold, rather than compressing every round.\n\n![Figure 2-16: Comparison of Context Compression Strategies](images/fig2-16.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第15/17部分]\n**其次,提高推理质量——总结的知识对模型比原始信息更有用**。这个动机更深层且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中并不总是最佳选择。\n\n考虑一个具体示例:在复杂任务中,代理通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在末尾附近。当代理必须从所有这些信息中做出最终决策时,它必须从分散在数万标记中的相关片段中检索。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索后,一次大语言模型调用可以生成累积信息的结构化摘要:“目前已知:A是……,B是……,关于C的信息仍缺失。” 然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的性质:**上下文中学习的内部机制更像是检索而不是推理**。第1章简要介绍了这个概念,代理状态栏部分通过机制、经验证据和工程实践进行了扩展。接下来,我们探讨这对压缩意味着什么。\n\n### 上下文中学习的内部机制:检索,而非推理\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在单次前向传递中主动计算聚合摘要。这并不否定模型可以通过生成思维链逐步推理;它意味着在单次前向传递中消耗现有上下文更像是检索。对压缩的影响很明显:状态栏**将计算出的结论添加到**上下文中,而压缩**用计算出的结论替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。区别在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用提炼一大块原始文本。\n\n一个简单的示例使“检索,而非推理”的想法具体化。假设上下文中包含宠物商店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> ...(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型:“有多少只黑猫和白猫?” 会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而不是检索。\n\n如果启用推理,模型可以通过逐个计数得到正确答案。成本是每次问这个问题时,都必须从头开始计数,生成许多推理标记。在代理场景中,如果这种统计信息需要反复使用(例如,用于每个决策),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录并直接将“当前统计:90只黑猫,10只白猫”写入上下文中,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识。**\n\n更深层的问题是长上下文降低了检索精度。即使上下文窗口远未填满,代理可能突然无法找到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文腐烂**。上下文腐烂不同于上下文溢出(窗口空间耗尽):溢出意味着“无法再容纳更多”,而腐烂意味着“容纳了但找不到”。后者更隐蔽,因为代理看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分布在更多标记上,每个标记获得的权重减少。更重要的是,一旦不相关内容主导上下文,代理的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分更难被注意到。一个有用的类比是在大型图书馆中寻找一本书:书架上不相关的书越多,找到目标就越难。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西提供了一个深刻的见解:模型的“记忆力差”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型自动从冗长的上下文中学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但它产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的结构化知识。**\n\n从这个角度看,上下文中学习更像是一种快速适应机制,而不是真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是暂时和肤浅的,会话结束后消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就像被“临时定制”了一样——不改变模型参数,但效果类似于一次小型的专门训练会话。这解释了为什么提示工程中的少样本示例可以显著提高输出质量,也解释了为什么这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n### 压缩与KV缓存:表面矛盾,实际互补\n在讨论具体的压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调KV缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的内容。\n\n关键是理解压缩的**时间和位置**。压缩不会在单次API调用期间修改上下文;相反,它发生在**两次API调用之间**,当代理框架预处理消息列表时:\n\n1. **系统提示和工具定义从不被触及**——这是上下文中最前面的“静态前缀”,KV缓存持续缓存。\n2. **压缩的目标是对话历史中的工具结果**——当代理框架将原始工具输出替换为压缩摘要时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文扩展超出窗口限制,任务完全失败;有压缩,一些缓存丢失,但上下文长度得到控制,信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第15/17部分]\n\n**其次,提高推理质量——总结的知识对模型来说比原始信息更有用**。这种动机更深刻且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中也不总是最佳选择。\n\n考虑一个具体例子:在一项复杂任务中,一个智能体通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在结尾附近。当智能体必须从所有这些信息中做出最终决策时,它必须检索分散在数万词元中的相关片段。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索之后,单个大语言模型调用可以生成积累信息的结构化总结:“目前已知:A是……,B是……,关于C的信息仍缺失。”然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的本质:**上下文学习的内部机制更像是检索而非推理**。第1章简要介绍了这个概念,“智能体状态栏”部分通过机制、实证证据和工程实践对其进行了扩展。接下来,我们研究这对压缩意味着什么。\n\n#### 上下文学习的内部机制:检索,而非推理\n\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在一次前向传递中主动计算聚合总结。这并不否认模型可以通过生成思维链逐步推理;而是意味着在一次前向传递中消耗现有上下文更像是检索。这对压缩的含义很明确:状态栏将计算出的结论**添加到**上下文中,而压缩用计算出的结论**替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。不同之处在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用来提炼一大块原始文本。\n\n一个简单的例子使“检索,而非推理”的概念具体化。假设上下文中包含宠物店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> …(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型,“有多少只黑猫和白猫?”会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而非检索。\n\n如果启用推理,模型可以通过逐一计数得到正确答案。代价是每次问这个问题时,都必须从头开始计数,生成许多推理词元。在智能体场景中,如果此类统计信息需要反复使用(例如每次决策都需要),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录,并直接将“当前统计:90只黑猫,10只白猫”写入上下文,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识**。\n\n更深层的问题是长上下文会降低检索精度。即使上下文窗口远未填满,智能体也可能突然找不到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文旋转**。上下文旋转不同于上下文溢出(窗口空间耗尽):溢出意味着“再也装不下了”,而旋转意味着“装得下但找不到”。后者更隐蔽,因为智能体看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分散到更多词元上,每个词元获得的权重降低。更重要的是,一旦不相关内容主导上下文,智能体的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分却更难被注意到。一个有用的类比是在大图书馆中寻找一本书:书架上不相关的书越多,越难找到目标。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将一条关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西(Andrej Karpathy)提出了一个深刻的见解:模型的“记忆不佳”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼出整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型从冗长的上下文中自动学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但会产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的、结构化的知识**。\n\n从这个角度看,上下文学习更像是一种快速适应机制,而非真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是临时且肤浅的,会话结束后就会消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就好像被“临时定制”了——没有改变模型参数,但效果类似于一次小型的专门训练会话。这解释了提示工程部分中的少样本示例为何能显著提高输出质量,也解释了为何这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n#### 压缩与键值缓存:表面矛盾,实际互补\n\n在讨论具体压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调键值缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的中间内容。\n\n关键是理解压缩的**时机和位置**。压缩不会在单个API调用期间修改上下文;相反,它发生在**两次API调用之间**,当智能体框架预处理消息列表时:\n\n1. **系统提示和工具定义永远不会被触碰**——这是上下文中最前面的“静态前缀”,键值缓存持续缓存。\n2. **压缩的目标是会话历史中的工具结果**——当智能体框架将原始工具输出替换为压缩总结时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文会超出窗口限制,任务彻底失败;有了压缩,一些缓存会丢失,但上下文长度得到控制且信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮都压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)" }, { "role": "assistant", "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } }" }, { "role": "user", "content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys." } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853909700533ea4fa82759f0cae4f2ffa723d4f9dd63964b2", "model": "doubao-seed-1-6-250615", "created": 1785390981, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }", "usage": { "prompt_tokens": 5763, "completion_tokens": 584, "total_tokens": 6347 } }, "latency_milliseconds": 11939, "validation": { "valid": false, "error_type": "ValueError", "error": "judge variants must contain exactly X and Y" } }, { "attempt": 3, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Second, improving reasoning quality—summarized knowledge is more useful to the model than raw information.** This motivation is deeper and easier to overlook. Even if the context window is large enough, adding all raw information to the context is not always the best choice.\n\nConsider a concrete example: during a complex task, an Agent accumulates information on a topic through 10 web searches. These search results are scattered in their raw form throughout the context—the results from round 2 are near the beginning, and the results from round 9 are near the end. When the Agent must make a final decision from all this information, it has to retrieve relevant fragments scattered across tens of thousands of tokens. Its attention becomes diffuse, and it can easily miss key information.\n\nAfter the 10th search, however, a single LLM call could produce a structured summary of the accumulated information: \"Currently known: A is..., B is..., information on C is still missing.\" The model can then use this refined knowledge representation in subsequent reasoning, without re-extracting it from the raw data.\n\nThe root cause lies in the nature of the attention mechanism: **the internal mechanism of in-context learning is more like retrieval than reasoning**. Chapter 1 introduced this concept briefly, and the Agent Status Bar section expanded it through mechanisms, empirical evidence, and engineering practices. Next, we examine what this means for compression.\n\n### The Internal Mechanism of In-Context Learning: Retrieval, Not Reasoning\n\nBriefly, **retrieval, not reasoning** means that attention is good at looking up existing content, but not at actively computing aggregate summaries in a single forward pass. This does not deny that the model can reason step by step by generating a chain of thought; it means that consuming existing context in one forward pass is more retrieval-like. The implication for compression is clear: the Status Bar **adds** computed conclusions **into** the context, while compression **replaces** bloated raw records **with** computed conclusions. Both supply the distillation layer that raw attention lacks. The difference is that the Status Bar is usually maintained deterministically, step by step, by **code**, while compression more often uses an LLM call to distill a large block of original text.\n\nA simple example makes the idea of \"retrieval, not reasoning\" concrete. Suppose the context contains a log of a pet store inspection:\n\n> Cage 1: Black cat. Cage 2: White cat. Cage 3: Black cat. Cage 4: Black cat. Cage 5: White cat.\n> ... (100 cages total, 90 black cats, 10 white cats)\n\nWhen you ask the model, \"How many black cats and white cats are there?\" what happens?\n\nIf reasoning is not enabled, the model will find it difficult to give the correct answer directly—because the attention mechanism is good at **looking up** (\"What cat is in cage 37?\"), not **aggregation** (\"How many black cats are there in total?\"). The latter requires traversing all records and maintaining a counting state, which is essentially reasoning, not retrieval.\n\nIf reasoning is enabled, the model can obtain the correct answer by counting one by one. The cost is that every time this question is asked, it must start counting from scratch, generating many reasoning tokens. In an Agent scenario, if such statistical information needs to be used repeatedly (e.g., for every decision), the cumulative reasoning cost becomes very high.\n\nHowever, if we summarize the records in advance and write \"Current statistics: 90 black cats, 10 white cats\" directly into the context, the model can retrieve the conclusion without repeating the count. **This is the second value of compression: turning conclusions that require reasoning into knowledge that can be directly retrieved.**\n\nThe deeper issue is that long contexts reduce retrieval precision. Even when the context window is far from full, the Agent may suddenly fail to find key information or repeatedly focus on a problem that has already been solved. This phenomenon is known as **Context Rot**. Context rot is different from context overflow (running out of window space): overflow means \"cannot fit any more,\" while rot means \"it fits but cannot be found.\" The latter is more insidious because the Agent appears to be working normally, while the quality of its decisions quietly deteriorates. As context length increases, attention weights are spread across more tokens, reducing the weight each token receives. More importantly, once irrelevant content dominates the context, the Agent's decision quality declines. In practice, the most common failure mode is not a context window that is too small, but an information density that is too low: knowledge needed only occasionally is loaded every time, stable rules are mixed with dynamic state, and the model sees more content while the useful parts become harder to notice. A useful analogy is searching for one book in a large library: the more irrelevant books on the shelves, the harder it is to find the target. The attention visualization in Experiment 2-2 demonstrates this phenomenon clearly: in long contexts, the model's attention exhibits strong positional bias. This is the problem revealed by the famous \"Needle in a Haystack\" experiment, which hides a key piece of information in the middle of a very long text and tests whether the model can find it.\n\nAndrej Karpathy offered a profound insight: the model's \"poor memory\" is, to some extent, a feature rather than a bug—the limited context window forces the model to learn to abstract general patterns from a large amount of detail, just as humans do not remember the verbatim content of every conversation but distill an overall impression and behavioral patterns.\n\nThis reveals the design principle of context compression: rather than expecting the model to learn automatically from lengthy context, we should distill that knowledge explicitly. Although this requires additional computation for summarization, it produces compact, information-dense representations. **Do not make the model search passively through vast amounts of raw material; provide refined, structured knowledge instead.**\n\nFrom this perspective, in-context learning is more like a rapid adaptation mechanism than true learning. It allows the model to quickly adjust its behavior during inference to suit a specific task, but this adjustment is temporary and shallow, disappearing after the session ends. Recent theoretical research[^ch2-6] supports this judgment: when the model sees examples in the context, its behavior is as if it has been \"temporarily customized\"—without changing the model parameters, but with an effect similar to a small, specialized training session. This explains why few-shot examples in the prompt engineering section can significantly improve output quality, and also why this improvement does not accumulate across sessions—it is fundamentally different from true parameter training.\n\n[^ch2-6]: Benoit Dherin et al., \"Learning without training\", 2025.\n\n### Compression and KV Cache: Apparent Contradiction, Practical Complementarity\n\nBefore discussing specific compression strategies, we need to resolve an apparent contradiction: earlier sections emphasized that KV Cache requires the context prefix to remain unchanged, but compression involves modifying content in the middle of the context.\n\nThe key is understanding the **timing and location** of compression. Compression does not modify the context during a single API call; instead, it occurs **between two API calls**, when the Agent framework preprocesses the message list:\n\n1. **System Prompt and Tool Definitions are never touched**—this is the \"static prefix\" at the very front of the context, and the KV Cache is continuously cached.\n2. **The target of compression is the tool results in the conversation history**—when the Agent framework replaces the original tool output with a compressed summary, the cache after the replacement point becomes invalid, but the cache before it remains valid.\n3. **This is a conscious trade-off**: without compression, the context expands beyond the window limit and the task fails outright; with it, some cache is lost, but context length stays under control and information density rises. Therefore, the frequency of compression needs to be weighed—frequent compression will frequently break the cache. It is best to perform batch compression when the context approaches the threshold, rather than compressing every round.\n\n![Figure 2-16: Comparison of Context Compression Strategies](images/fig2-16.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第15/17部分]\n**其次,提高推理质量——总结的知识对模型比原始信息更有用**。这个动机更深层且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中并不总是最佳选择。\n\n考虑一个具体示例:在复杂任务中,代理通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在末尾附近。当代理必须从所有这些信息中做出最终决策时,它必须从分散在数万标记中的相关片段中检索。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索后,一次大语言模型调用可以生成累积信息的结构化摘要:“目前已知:A是……,B是……,关于C的信息仍缺失。” 然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的性质:**上下文中学习的内部机制更像是检索而不是推理**。第1章简要介绍了这个概念,代理状态栏部分通过机制、经验证据和工程实践进行了扩展。接下来,我们探讨这对压缩意味着什么。\n\n### 上下文中学习的内部机制:检索,而非推理\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在单次前向传递中主动计算聚合摘要。这并不否定模型可以通过生成思维链逐步推理;它意味着在单次前向传递中消耗现有上下文更像是检索。对压缩的影响很明显:状态栏**将计算出的结论添加到**上下文中,而压缩**用计算出的结论替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。区别在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用提炼一大块原始文本。\n\n一个简单的示例使“检索,而非推理”的想法具体化。假设上下文中包含宠物商店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> ...(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型:“有多少只黑猫和白猫?” 会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而不是检索。\n\n如果启用推理,模型可以通过逐个计数得到正确答案。成本是每次问这个问题时,都必须从头开始计数,生成许多推理标记。在代理场景中,如果这种统计信息需要反复使用(例如,用于每个决策),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录并直接将“当前统计:90只黑猫,10只白猫”写入上下文中,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识。**\n\n更深层的问题是长上下文降低了检索精度。即使上下文窗口远未填满,代理可能突然无法找到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文腐烂**。上下文腐烂不同于上下文溢出(窗口空间耗尽):溢出意味着“无法再容纳更多”,而腐烂意味着“容纳了但找不到”。后者更隐蔽,因为代理看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分布在更多标记上,每个标记获得的权重减少。更重要的是,一旦不相关内容主导上下文,代理的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分更难被注意到。一个有用的类比是在大型图书馆中寻找一本书:书架上不相关的书越多,找到目标就越难。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西提供了一个深刻的见解:模型的“记忆力差”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型自动从冗长的上下文中学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但它产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的结构化知识。**\n\n从这个角度看,上下文中学习更像是一种快速适应机制,而不是真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是暂时和肤浅的,会话结束后消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就像被“临时定制”了一样——不改变模型参数,但效果类似于一次小型的专门训练会话。这解释了为什么提示工程中的少样本示例可以显著提高输出质量,也解释了为什么这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n### 压缩与KV缓存:表面矛盾,实际互补\n在讨论具体的压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调KV缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的内容。\n\n关键是理解压缩的**时间和位置**。压缩不会在单次API调用期间修改上下文;相反,它发生在**两次API调用之间**,当代理框架预处理消息列表时:\n\n1. **系统提示和工具定义从不被触及**——这是上下文中最前面的“静态前缀”,KV缓存持续缓存。\n2. **压缩的目标是对话历史中的工具结果**——当代理框架将原始工具输出替换为压缩摘要时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文扩展超出窗口限制,任务完全失败;有压缩,一些缓存丢失,但上下文长度得到控制,信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第15/17部分]\n\n**其次,提高推理质量——总结的知识对模型来说比原始信息更有用**。这种动机更深刻且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中也不总是最佳选择。\n\n考虑一个具体例子:在一项复杂任务中,一个智能体通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在结尾附近。当智能体必须从所有这些信息中做出最终决策时,它必须检索分散在数万词元中的相关片段。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索之后,单个大语言模型调用可以生成积累信息的结构化总结:“目前已知:A是……,B是……,关于C的信息仍缺失。”然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的本质:**上下文学习的内部机制更像是检索而非推理**。第1章简要介绍了这个概念,“智能体状态栏”部分通过机制、实证证据和工程实践对其进行了扩展。接下来,我们研究这对压缩意味着什么。\n\n#### 上下文学习的内部机制:检索,而非推理\n\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在一次前向传递中主动计算聚合总结。这并不否认模型可以通过生成思维链逐步推理;而是意味着在一次前向传递中消耗现有上下文更像是检索。这对压缩的含义很明确:状态栏将计算出的结论**添加到**上下文中,而压缩用计算出的结论**替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。不同之处在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用来提炼一大块原始文本。\n\n一个简单的例子使“检索,而非推理”的概念具体化。假设上下文中包含宠物店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> …(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型,“有多少只黑猫和白猫?”会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而非检索。\n\n如果启用推理,模型可以通过逐一计数得到正确答案。代价是每次问这个问题时,都必须从头开始计数,生成许多推理词元。在智能体场景中,如果此类统计信息需要反复使用(例如每次决策都需要),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录,并直接将“当前统计:90只黑猫,10只白猫”写入上下文,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识**。\n\n更深层的问题是长上下文会降低检索精度。即使上下文窗口远未填满,智能体也可能突然找不到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文旋转**。上下文旋转不同于上下文溢出(窗口空间耗尽):溢出意味着“再也装不下了”,而旋转意味着“装得下但找不到”。后者更隐蔽,因为智能体看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分散到更多词元上,每个词元获得的权重降低。更重要的是,一旦不相关内容主导上下文,智能体的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分却更难被注意到。一个有用的类比是在大图书馆中寻找一本书:书架上不相关的书越多,越难找到目标。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将一条关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西(Andrej Karpathy)提出了一个深刻的见解:模型的“记忆不佳”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼出整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型从冗长的上下文中自动学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但会产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的、结构化的知识**。\n\n从这个角度看,上下文学习更像是一种快速适应机制,而非真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是临时且肤浅的,会话结束后就会消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就好像被“临时定制”了——没有改变模型参数,但效果类似于一次小型的专门训练会话。这解释了提示工程部分中的少样本示例为何能显著提高输出质量,也解释了为何这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n#### 压缩与键值缓存:表面矛盾,实际互补\n\n在讨论具体压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调键值缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的中间内容。\n\n关键是理解压缩的**时机和位置**。压缩不会在单个API调用期间修改上下文;相反,它发生在**两次API调用之间**,当智能体框架预处理消息列表时:\n\n1. **系统提示和工具定义永远不会被触碰**——这是上下文中最前面的“静态前缀”,键值缓存持续缓存。\n2. **压缩的目标是会话历史中的工具结果**——当智能体框架将原始工具输出替换为压缩总结时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文会超出窗口限制,任务彻底失败;有了压缩,一些缓存会丢失,但上下文长度得到控制且信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮都压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)" }, { "role": "assistant", "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }" }, { "role": "user", "content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys." } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853909819933ea4fa82759f0cae4f2ffa723d4f9dd6ce50a0", "model": "doubao-seed-1-6-250615", "created": 1785390993, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }", "usage": { "prompt_tokens": 5764, "completion_tokens": 584, "total_tokens": 6348 } }, "latency_milliseconds": 11584, "validation": { "valid": false, "error_type": "ValueError", "error": "judge variants must contain exactly X and Y" } }, { "attempt": 4, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Second, improving reasoning quality—summarized knowledge is more useful to the model than raw information.** This motivation is deeper and easier to overlook. Even if the context window is large enough, adding all raw information to the context is not always the best choice.\n\nConsider a concrete example: during a complex task, an Agent accumulates information on a topic through 10 web searches. These search results are scattered in their raw form throughout the context—the results from round 2 are near the beginning, and the results from round 9 are near the end. When the Agent must make a final decision from all this information, it has to retrieve relevant fragments scattered across tens of thousands of tokens. Its attention becomes diffuse, and it can easily miss key information.\n\nAfter the 10th search, however, a single LLM call could produce a structured summary of the accumulated information: \"Currently known: A is..., B is..., information on C is still missing.\" The model can then use this refined knowledge representation in subsequent reasoning, without re-extracting it from the raw data.\n\nThe root cause lies in the nature of the attention mechanism: **the internal mechanism of in-context learning is more like retrieval than reasoning**. Chapter 1 introduced this concept briefly, and the Agent Status Bar section expanded it through mechanisms, empirical evidence, and engineering practices. Next, we examine what this means for compression.\n\n### The Internal Mechanism of In-Context Learning: Retrieval, Not Reasoning\n\nBriefly, **retrieval, not reasoning** means that attention is good at looking up existing content, but not at actively computing aggregate summaries in a single forward pass. This does not deny that the model can reason step by step by generating a chain of thought; it means that consuming existing context in one forward pass is more retrieval-like. The implication for compression is clear: the Status Bar **adds** computed conclusions **into** the context, while compression **replaces** bloated raw records **with** computed conclusions. Both supply the distillation layer that raw attention lacks. The difference is that the Status Bar is usually maintained deterministically, step by step, by **code**, while compression more often uses an LLM call to distill a large block of original text.\n\nA simple example makes the idea of \"retrieval, not reasoning\" concrete. Suppose the context contains a log of a pet store inspection:\n\n> Cage 1: Black cat. Cage 2: White cat. Cage 3: Black cat. Cage 4: Black cat. Cage 5: White cat.\n> ... (100 cages total, 90 black cats, 10 white cats)\n\nWhen you ask the model, \"How many black cats and white cats are there?\" what happens?\n\nIf reasoning is not enabled, the model will find it difficult to give the correct answer directly—because the attention mechanism is good at **looking up** (\"What cat is in cage 37?\"), not **aggregation** (\"How many black cats are there in total?\"). The latter requires traversing all records and maintaining a counting state, which is essentially reasoning, not retrieval.\n\nIf reasoning is enabled, the model can obtain the correct answer by counting one by one. The cost is that every time this question is asked, it must start counting from scratch, generating many reasoning tokens. In an Agent scenario, if such statistical information needs to be used repeatedly (e.g., for every decision), the cumulative reasoning cost becomes very high.\n\nHowever, if we summarize the records in advance and write \"Current statistics: 90 black cats, 10 white cats\" directly into the context, the model can retrieve the conclusion without repeating the count. **This is the second value of compression: turning conclusions that require reasoning into knowledge that can be directly retrieved.**\n\nThe deeper issue is that long contexts reduce retrieval precision. Even when the context window is far from full, the Agent may suddenly fail to find key information or repeatedly focus on a problem that has already been solved. This phenomenon is known as **Context Rot**. Context rot is different from context overflow (running out of window space): overflow means \"cannot fit any more,\" while rot means \"it fits but cannot be found.\" The latter is more insidious because the Agent appears to be working normally, while the quality of its decisions quietly deteriorates. As context length increases, attention weights are spread across more tokens, reducing the weight each token receives. More importantly, once irrelevant content dominates the context, the Agent's decision quality declines. In practice, the most common failure mode is not a context window that is too small, but an information density that is too low: knowledge needed only occasionally is loaded every time, stable rules are mixed with dynamic state, and the model sees more content while the useful parts become harder to notice. A useful analogy is searching for one book in a large library: the more irrelevant books on the shelves, the harder it is to find the target. The attention visualization in Experiment 2-2 demonstrates this phenomenon clearly: in long contexts, the model's attention exhibits strong positional bias. This is the problem revealed by the famous \"Needle in a Haystack\" experiment, which hides a key piece of information in the middle of a very long text and tests whether the model can find it.\n\nAndrej Karpathy offered a profound insight: the model's \"poor memory\" is, to some extent, a feature rather than a bug—the limited context window forces the model to learn to abstract general patterns from a large amount of detail, just as humans do not remember the verbatim content of every conversation but distill an overall impression and behavioral patterns.\n\nThis reveals the design principle of context compression: rather than expecting the model to learn automatically from lengthy context, we should distill that knowledge explicitly. Although this requires additional computation for summarization, it produces compact, information-dense representations. **Do not make the model search passively through vast amounts of raw material; provide refined, structured knowledge instead.**\n\nFrom this perspective, in-context learning is more like a rapid adaptation mechanism than true learning. It allows the model to quickly adjust its behavior during inference to suit a specific task, but this adjustment is temporary and shallow, disappearing after the session ends. Recent theoretical research[^ch2-6] supports this judgment: when the model sees examples in the context, its behavior is as if it has been \"temporarily customized\"—without changing the model parameters, but with an effect similar to a small, specialized training session. This explains why few-shot examples in the prompt engineering section can significantly improve output quality, and also why this improvement does not accumulate across sessions—it is fundamentally different from true parameter training.\n\n[^ch2-6]: Benoit Dherin et al., \"Learning without training\", 2025.\n\n### Compression and KV Cache: Apparent Contradiction, Practical Complementarity\n\nBefore discussing specific compression strategies, we need to resolve an apparent contradiction: earlier sections emphasized that KV Cache requires the context prefix to remain unchanged, but compression involves modifying content in the middle of the context.\n\nThe key is understanding the **timing and location** of compression. Compression does not modify the context during a single API call; instead, it occurs **between two API calls**, when the Agent framework preprocesses the message list:\n\n1. **System Prompt and Tool Definitions are never touched**—this is the \"static prefix\" at the very front of the context, and the KV Cache is continuously cached.\n2. **The target of compression is the tool results in the conversation history**—when the Agent framework replaces the original tool output with a compressed summary, the cache after the replacement point becomes invalid, but the cache before it remains valid.\n3. **This is a conscious trade-off**: without compression, the context expands beyond the window limit and the task fails outright; with it, some cache is lost, but context length stays under control and information density rises. Therefore, the frequency of compression needs to be weighed—frequent compression will frequently break the cache. It is best to perform batch compression when the context approaches the threshold, rather than compressing every round.\n\n![Figure 2-16: Comparison of Context Compression Strategies](images/fig2-16.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第15/17部分]\n**其次,提高推理质量——总结的知识对模型比原始信息更有用**。这个动机更深层且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中并不总是最佳选择。\n\n考虑一个具体示例:在复杂任务中,代理通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在末尾附近。当代理必须从所有这些信息中做出最终决策时,它必须从分散在数万标记中的相关片段中检索。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索后,一次大语言模型调用可以生成累积信息的结构化摘要:“目前已知:A是……,B是……,关于C的信息仍缺失。” 然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的性质:**上下文中学习的内部机制更像是检索而不是推理**。第1章简要介绍了这个概念,代理状态栏部分通过机制、经验证据和工程实践进行了扩展。接下来,我们探讨这对压缩意味着什么。\n\n### 上下文中学习的内部机制:检索,而非推理\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在单次前向传递中主动计算聚合摘要。这并不否定模型可以通过生成思维链逐步推理;它意味着在单次前向传递中消耗现有上下文更像是检索。对压缩的影响很明显:状态栏**将计算出的结论添加到**上下文中,而压缩**用计算出的结论替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。区别在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用提炼一大块原始文本。\n\n一个简单的示例使“检索,而非推理”的想法具体化。假设上下文中包含宠物商店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> ...(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型:“有多少只黑猫和白猫?” 会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而不是检索。\n\n如果启用推理,模型可以通过逐个计数得到正确答案。成本是每次问这个问题时,都必须从头开始计数,生成许多推理标记。在代理场景中,如果这种统计信息需要反复使用(例如,用于每个决策),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录并直接将“当前统计:90只黑猫,10只白猫”写入上下文中,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识。**\n\n更深层的问题是长上下文降低了检索精度。即使上下文窗口远未填满,代理可能突然无法找到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文腐烂**。上下文腐烂不同于上下文溢出(窗口空间耗尽):溢出意味着“无法再容纳更多”,而腐烂意味着“容纳了但找不到”。后者更隐蔽,因为代理看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分布在更多标记上,每个标记获得的权重减少。更重要的是,一旦不相关内容主导上下文,代理的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分更难被注意到。一个有用的类比是在大型图书馆中寻找一本书:书架上不相关的书越多,找到目标就越难。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西提供了一个深刻的见解:模型的“记忆力差”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型自动从冗长的上下文中学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但它产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的结构化知识。**\n\n从这个角度看,上下文中学习更像是一种快速适应机制,而不是真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是暂时和肤浅的,会话结束后消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就像被“临时定制”了一样——不改变模型参数,但效果类似于一次小型的专门训练会话。这解释了为什么提示工程中的少样本示例可以显著提高输出质量,也解释了为什么这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n### 压缩与KV缓存:表面矛盾,实际互补\n在讨论具体的压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调KV缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的内容。\n\n关键是理解压缩的**时间和位置**。压缩不会在单次API调用期间修改上下文;相反,它发生在**两次API调用之间**,当代理框架预处理消息列表时:\n\n1. **系统提示和工具定义从不被触及**——这是上下文中最前面的“静态前缀”,KV缓存持续缓存。\n2. **压缩的目标是对话历史中的工具结果**——当代理框架将原始工具输出替换为压缩摘要时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文扩展超出窗口限制,任务完全失败;有压缩,一些缓存丢失,但上下文长度得到控制,信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第15/17部分]\n\n**其次,提高推理质量——总结的知识对模型来说比原始信息更有用**。这种动机更深刻且更容易被忽视。即使上下文窗口足够大,将所有原始信息添加到上下文中也不总是最佳选择。\n\n考虑一个具体例子:在一项复杂任务中,一个智能体通过10次网络搜索积累了关于某个主题的信息。这些搜索结果以原始形式分散在上下文中——第2轮的结果在开头附近,第9轮的结果在结尾附近。当智能体必须从所有这些信息中做出最终决策时,它必须检索分散在数万词元中的相关片段。它的注意力变得分散,很容易错过关键信息。\n\n然而,在第10次搜索之后,单个大语言模型调用可以生成积累信息的结构化总结:“目前已知:A是……,B是……,关于C的信息仍缺失。”然后模型可以在后续推理中使用这种精炼的知识表示,而无需从原始数据中重新提取。\n\n根本原因在于注意力机制的本质:**上下文学习的内部机制更像是检索而非推理**。第1章简要介绍了这个概念,“智能体状态栏”部分通过机制、实证证据和工程实践对其进行了扩展。接下来,我们研究这对压缩意味着什么。\n\n#### 上下文学习的内部机制:检索,而非推理\n\n简而言之,**检索,而非推理**意味着注意力擅长查找现有内容,但不擅长在一次前向传递中主动计算聚合总结。这并不否认模型可以通过生成思维链逐步推理;而是意味着在一次前向传递中消耗现有上下文更像是检索。这对压缩的含义很明确:状态栏将计算出的结论**添加到**上下文中,而压缩用计算出的结论**替换**臃肿的原始记录。两者都提供了原始注意力缺乏的提炼层。不同之处在于,状态栏通常由**代码**确定性地逐步维护,而压缩更常使用大语言模型调用来提炼一大块原始文本。\n\n一个简单的例子使“检索,而非推理”的概念具体化。假设上下文中包含宠物店检查的日志:\n\n> 笼子1:黑猫。笼子2:白猫。笼子3:黑猫。笼子4:黑猫。笼子5:白猫。\n> …(总共100个笼子,90只黑猫,10只白猫)\n\n当你问模型,“有多少只黑猫和白猫?”会发生什么?\n\n如果没有启用推理,模型将很难直接给出正确答案——因为注意力机制擅长**查找**(“笼子37里是什么猫?”),而不擅长**聚合**(“总共有多少只黑猫?”)。后者需要遍历所有记录并维护计数状态,这本质上是推理,而非检索。\n\n如果启用推理,模型可以通过逐一计数得到正确答案。代价是每次问这个问题时,都必须从头开始计数,生成许多推理词元。在智能体场景中,如果此类统计信息需要反复使用(例如每次决策都需要),累积的推理成本会非常高。\n\n然而,如果我们提前总结记录,并直接将“当前统计:90只黑猫,10只白猫”写入上下文,模型可以检索结论而无需重复计数。**这是压缩的第二个价值:将需要推理的结论转化为可以直接检索的知识**。\n\n更深层的问题是长上下文会降低检索精度。即使上下文窗口远未填满,智能体也可能突然找不到关键信息或反复聚焦于已解决的问题。这种现象称为**上下文旋转**。上下文旋转不同于上下文溢出(窗口空间耗尽):溢出意味着“再也装不下了”,而旋转意味着“装得下但找不到”。后者更隐蔽,因为智能体看似正常工作,而其决策质量却悄然下降。随着上下文长度增加,注意力权重分散到更多词元上,每个词元获得的权重降低。更重要的是,一旦不相关内容主导上下文,智能体的决策质量就会下降。在实践中,最常见的失败模式不是上下文窗口太小,而是信息密度太低:偶尔需要的知识每次都加载,稳定规则与动态状态混合,模型看到更多内容但有用部分却更难被注意到。一个有用的类比是在大图书馆中寻找一本书:书架上不相关的书越多,越难找到目标。实验2-2中的注意力可视化清晰地展示了这种现象:在长上下文中,模型的注意力表现出强烈的位置偏差。这就是著名的“大海捞针”实验揭示的问题,该实验将一条关键信息隐藏在非常长的文本中间,测试模型是否能找到它。\n\n安德烈·卡帕西(Andrej Karpathy)提出了一个深刻的见解:模型的“记忆不佳”在某种程度上是一种特征而非缺陷——有限的上下文窗口迫使模型像人类一样从大量细节中学习抽象的一般模式,人类不会记住每次对话的逐字内容,而是提炼出整体印象和行为模式。\n\n这揭示了上下文压缩的设计原则:与其期望模型从冗长的上下文中自动学习,不如明确提炼该知识。虽然这需要额外的计算来进行总结,但会产生紧凑、信息密集的表示。**不要让模型被动地搜索大量原始材料;而是提供精炼的、结构化的知识**。\n\n从这个角度看,上下文学习更像是一种快速适应机制,而非真正的学习。它允许模型在推理期间快速调整其行为以适应特定任务,但这种调整是临时且肤浅的,会话结束后就会消失。最近的理论研究[^ch2-6]支持这一判断:当模型在上下文中看到示例时,其行为就好像被“临时定制”了——没有改变模型参数,但效果类似于一次小型的专门训练会话。这解释了提示工程部分中的少样本示例为何能显著提高输出质量,也解释了为何这种改进不会在会话间累积——它与真正的参数训练根本不同。\n\n[^ch2-6]: Benoit Dherin等人,“无需训练的学习”,2025年。\n\n#### 压缩与键值缓存:表面矛盾,实际互补\n\n在讨论具体压缩策略之前,我们需要解决一个表面矛盾:前面的章节强调键值缓存要求上下文前缀保持不变,但压缩涉及修改上下文中的中间内容。\n\n关键是理解压缩的**时机和位置**。压缩不会在单个API调用期间修改上下文;相反,它发生在**两次API调用之间**,当智能体框架预处理消息列表时:\n\n1. **系统提示和工具定义永远不会被触碰**——这是上下文中最前面的“静态前缀”,键值缓存持续缓存。\n2. **压缩的目标是会话历史中的工具结果**——当智能体框架将原始工具输出替换为压缩总结时,替换点之后的缓存失效,但之前的缓存仍然有效。\n3. **这是一种有意识的权衡**:没有压缩,上下文会超出窗口限制,任务彻底失败;有了压缩,一些缓存会丢失,但上下文长度得到控制且信息密度提高。因此,需要权衡压缩的频率——频繁压缩会频繁破坏缓存。最好在上下文接近阈值时进行批量压缩,而不是每轮都压缩。\n\n![图2-16:上下文压缩策略比较](images/fig2-16.svg)" }, { "role": "assistant", "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }" }, { "role": "user", "content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys." } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853909935793ea4fa82759f0cae4f2ffa723d4f9dd6b564f5", "model": "doubao-seed-1-6-250615", "created": 1785391005, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }", "usage": { "prompt_tokens": 5764, "completion_tokens": 584, "total_tokens": 6348 } }, "latency_milliseconds": 12231, "validation": { "valid": false, "error_type": "ValueError", "error": "judge variants must contain exactly X and Y" } }, { "attempt": 5, "request_kind": "schema_repair", "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "This is a formatting repair, not a new evaluation. Reshape the JSON below into the exact requested schema while preserving every substantive score, evidence statement, preference, and preference explanation. The top level must contain variants, preferred, and preference_evidence. variants must contain exactly X and Y. Each of X and Y must contain exactly accuracy, fluency, terminology, and markdown_code_fidelity, and every dimension must contain score and evidence. Do not re-evaluate, rename fields, nest Y inside X, or add keys. Previous validation error: judge variants must contain exactly X and Y. Return JSON only.\n\nJSON TO REPAIR:\n{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"},\"fluency\":{\"score\":4,\"evidence\":\"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\",\"terminology\":{\"score\":3,\"evidence\":\"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\",\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\",\"fluency\":{\"score\":5,\"evidence\":\"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\",\"terminology\":{\"score\":5,\"evidence\":\"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\",\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}} } } } } }" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853911024323b2c90331db17702ff1ef9ef62383e60d50ce5", "model": "doubao-seed-1-6-250615", "created": 1785391114, "content": "{\"variants\": {\"X\": {\"accuracy\": {\"score\": 4, \"evidence\": \"X accurately translates most content, but misrenders 'Context Rot' as '上下文腐烂' (context rot) instead of the correct technical term '上下文旋转' (context rotation).\"}, \"fluency\": {\"score\": 4, \"evidence\": \"X uses natural phrasing like '注意力变得分散' (attention becomes diffuse) and '精炼的知识表示' (refined knowledge representation), but '上下文中学习' (in-context learning) is slightly awkward compared to Y's '上下文学习'.\"}, \"terminology\": {\"score\": 3, \"evidence\": \"X uses inconsistent terms: '代理' (agent) vs. Y's consistent '智能体' (agent); '上下文中学习' (in-context learning) vs. Y's '上下文学习'; mistranslates 'Context Rot' as '上下文腐烂' (context rot) instead of '上下文旋转' (context rotation).\"}, \"markdown_code_fidelity\": {\"score\": 5, \"evidence\": \"X preserves all Markdown elements: headings (###), bold text, blockquotes, list items, and the image link ![图2-16:上下文压缩策略比较](images/fig2-16.svg) are correctly maintained.\"}}, \"Y\": {\"accuracy\": {\"score\": 5, \"evidence\": \"Y accurately translates all technical terms, including 'Context Rot' as '上下文旋转' (context rotation), and preserves original meaning without omissions or inventions.\"}, \"fluency\": {\"score\": 5, \"evidence\": \"Y uses smooth, natural phrasing: '智能体' (agent) consistently, '上下文学习' (in-context learning) flows better, and '逐一计数' (counting one by one) is more idiomatic than X's '逐个计数'.\"}, \"terminology\": {\"score\": 5, \"evidence\": \"Y uses consistent technical terms: '智能体' (agent), '上下文学习' (in-context learning), '上下文旋转' (context rotation), '键值缓存' (KV cache), and '词元' (token) are all accurate and consistent.\"}, \"markdown_code_fidelity\": {\"score\": 4, \"evidence\": \"Y incorrectly changes the second-level heading '### The Internal Mechanism...' to '####' (four hashes) instead of '###' (three hashes), breaking heading hierarchy; other elements like bold, images, and lists are preserved.\"}}}, \"preferred\": \"Y\", \"preference_evidence\": \"Y outperforms X in accuracy (correct '上下文旋转' for Context Rot), fluency (more natural phrasing), and terminology (consistent '智能体' and correct technical terms). Despite a minor Markdown heading error (#### instead of ###), Y's overall quality is higher than X's inconsistent terminology and misrendering of 'Context Rot'.\"}", "usage": { "prompt_tokens": 750, "completion_tokens": 620, "total_tokens": 1370 } }, "latency_milliseconds": 12556, "validation": { "valid": true } } ] }