Files
ai-agent-book/chapter10/book-translation/validation/real_20260730T061500Z_v4/judge_receipts/unit-13.json
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

89 lines
66 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nThat one timestamp line invalidated the KV Cache on every request. The system prompt was now different each time, forcing the model to recompute the key-value pairs for the prefix from scratch (here, \"Key\" and \"Value\" are two types of vectors in the attention mechanism; Experiment 2-2 below visually demonstrates their roles). This kind of invisible cost appears repeatedly in Agent systems: a seemingly harmless line of code can slow down the entire inference pipeline by an order of magnitude. This section explains how to avoid these pitfalls.\n\n> **Technical Note**: This section involves the internal principles of the Transformer attention mechanism and KV Cache, making it one of the most technically dense parts of the book. If you are not familiar with these underlying mechanisms, **you can skip the detailed principles and remember the following three core conclusions**:\n>\n> 1. **Once the system prompt and tool definitions are finalized, do not change them.** Any modification, even adding a single space, will invalidate the entire cache and can multiply latency and increase costs (the exact magnitude depends on the model and configuration).\n> 2. **Always append dynamic information to the end**—changing content like timestamps and user status should be appended as new messages at the end of the conversation, not by modifying the existing system prompt.\n> 3. **Use the standard API format; do not manually concatenate messages**: Structured messages are translated by the Chat Template into a fixed token sequence that the model saw during training. The fundamental problem with manually concatenating strings into formats like `\"USER: ... ASSISTANT: ...\"` is that it deviates from this training format, weakening the model's multi-step reasoning ability. Caching, however, depends only on the resulting token sequence. A manually concatenated prefix can still be cached if it remains byte-for-byte stable. The cache is invalidated only when that prefix changes, for example, when dynamic content is inserted into it.\n>\n> The intuition behind these three conclusions is simple: when processing context, an LLM caches computation for the prefix it has already processed, so the next request can reuse that work. **If the prefix is byte-for-byte identical, the cached computation can be reused; if the prefix changes, the computation after that point must be rebuilt.** The system prompt and tool definitions are usually the earliest and most expensive part of this prefix; once they change, the cached intermediate results after that point are invalidated.\n>\n> Remember these three principles, and even if you skip the technical details below, you can correctly design the context structure of an Agent. The following content is for readers who want to delve deeper into the \"why.\"\n\n> **Experiment 2-2 ★: Attention Mechanism Visualization**\n>\n> Before explaining KV Cache, we first build an intuitive understanding of the model's internal attention mechanism through an experiment—this is the foundation for understanding why KV Cache is effective and why it imposes strict requirements on context design.\n>\n> **What is the Attention Mechanism?** Consider a concrete example. Suppose the model is processing the Chinese sentence \"北京 的 天气 怎么样\" (\"How's the weather in Beijing?\"), whose words are \"北京\" (Beijing), \"的\" (a possessive particle, like \"of\"), \"天气\" (weather), and \"怎么样\" (how is it). When it reads \"怎么样\", the model needs to decide: which of the preceding words are most important for understanding \"怎么样\"?\n>\n> The attention mechanism uses three types of vectors to decide which earlier tokens are most relevant:\n>\n> Table 2-1 summarizes the roles of the Query, Key, and Value vectors in the attention mechanism, helping readers map the abstract computation onto the example sentence \"北京的天气怎么样\" (\"How's the weather in Beijing?\").\n>\n> Table 2-1 Roles of Query, Key, and Value in the Attention Mechanism\n>\n> | Vector | Meaning | In this example |\n> |-------|-----------------------------------------|-----------------------------------------------|\n> | **Query** | The \"search request\" issued by the current word | \"怎么样\" (how is it) asks: which word is most relevant to me? |\n> | **Key** | The \"label\" of each word, used for matching the search | The label of \"北京\" (Beijing) leans toward \"place name\"; the label of \"天气\" (weather) leans toward \"meteorology\" |\n> | **Value** | The \"content\" of each word, extracted upon a successful match | After matching \"天气\" (weather), extract its semantic information |\n>\n> In simplified terms, each new word scores the preceding words by relevance, then uses the most relevant information to build its current representation.\n>\n> More specifically, the computation has three steps. First, \"怎么样\" generates its own Query vector, representing what the current token is looking for. Second, the Query is compared with the Key of each preceding word using a dot product, producing a relevance score; higher scores indicate stronger matches. Finally, these scores become attention weights, which are used to compute a weighted sum of the Values. Words with higher weights contribute more to the final representation, while words with lower weights contribute less.\n>\n>\n> ![Figure 2-6: Intuitive Understanding of the Attention Mechanism](images/fig2-6.svg)\n>\n>\n> The upper part of Figure 2-6 shows how \"怎么样\" (how is it) matches each preceding word: the strongest match is with \"天气\" (weather, 0.55), there is some relevance to \"北京\" (Beijing, 0.35), almost none to \"的\" (the particle, 0.05), and the remaining weight of about 0.05 goes to \"怎么样\" itself (not shown separately in the figure)—all weights sum to 1. The final output draws mainly on the information from \"天气\", which matches intuition exactly.\n>\n> An **attention heatmap** arranges the attention weights between each word and all preceding words into a matrix. The lower part of Figure 2-6 shows the complete heatmap: each row is a Query (the word currently being processed), each column is a Key (the word being attended to), and darker cells indicate higher attention weights. The heatmap is triangular because the model generates text from left to right: each word can attend only to itself and the words before it, not to content that has yet to be generated.\n>\n> **Why do Key and Value need to be cached?** Observing the heatmap reveals that every time a new word is generated, its Query must be matched against the Keys of **all** preceding words, and then a weighted sum of all Values is computed. If all K and V values were recalculated from scratch each time, the computation would grow with the context length. The KV Cache stores the already computed K and V values, allowing new words to directly reuse them — this is the core optimization discussed next.\n>\n> With a basic understanding of the attention mechanism, we can now observe the attention distribution of a real model through the `attention_visualization` experiment.\n>\n>\n> ![Figure 2-7: Attention Heatmap Visualization](images/fig2-7.png)\n>\n>\n> The attention heatmap reveals several key patterns:\n>\n> 1. **Attention Sink**: The first token of the sequence often absorbs an abnormally high amount of attention weight, sometimes exceeding 70% of the total attention. The model uses this position as an \"Attention Sink\" to absorb residual attention mass that does not strongly correspond to any other specific token. In other words, the model learns to assign otherwise unallocated attention weight to the first token — this is a systematic phenomenon, not a model defect.\n>\n> The mathematical reason is that the attention mechanism has a hard constraint: all attention weights must sum to exactly 100% (guaranteed by a mathematical function called softmax), so the model cannot express \"not attending to anything.\" Even if the current word is not very relevant to any preceding word, these weights must be allocated somewhere. The model therefore needs a stable container for this \"residual weight,\" and the fixed position at the beginning of the sequence becomes the most natural choice. This is an inevitable consequence of the mathematical properties of softmax when processing many tokens.\n> 2. **Reasoning Triangle Pattern**: The model's chain of thought (within `<think>` tags) exhibits a triangular self-attention pattern: when generating new reasoning content, it frequently attends to earlier reasoning content and tool definitions.\n> 3. **Output Triangle Pattern**: The output process after reasoning ends shows another triangle, where the model uses the reasoning trace as a prompt to generate the answer.\n> 4. **Position Bias**[^lost-in-the-middle]: The model has higher recall accuracy for information at the beginning and end of the context, while information in the middle is more likely to be overlooked. Therefore, when designing the context, placing the most critical information at the beginning or end is an important practical principle.\n>\n> This experiment shows that **long chain-of-thought generation and tool calling both depend heavily on in-context learning** — the model's ability to adapt to a task based on the instructions and examples provided in the input, without retraining. For the internal mechanism of in-context learning and its implications for Agent architecture design, see the Context Compression section of this chapter.\n>\n\n[^lost-in-the-middle]: Liu et al. [\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/), TACL, 2024.\n\n### From API Messages to Model Tokens: Chat Template\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第4/17部分]\n那一行时间戳导致每次请求时键值缓存(KV Cache)都失效。系统提示词现在每次都不同,迫使模型从头开始重新计算前缀的键值对(这里“键”和“值”是注意力机制中的两种向量;下面的实验2-2直观展示了它们的作用)。这种无形的成本在智能体(Agent)系统中反复出现:看似无害的一行代码可能会使整个推理流水线的速度降低一个数量级。本节将解释如何避免这些陷阱。\n\n> **技术说明**:本节涉及Transformer注意力机制和键值缓存的内部原理,是本书技术密度较高的部分之一。如果您不熟悉这些底层机制,**您可以跳过详细原理,记住以下三个核心结论**:\n>\n> 1. **一旦系统提示词和工具定义确定,就不要更改它们**。任何修改,即使是添加一个空格,都会使整个缓存失效,并可能成倍增加时延和成本(具体幅度取决于模型和配置)。\n> 2. **始终在末尾追加动态信息**——时间戳、用户状态等内容的更改应作为新消息追加在对话末尾,而不是修改现有的系统提示词。\n> 3. **使用标准API格式,不要手动拼接消息**:结构化消息由聊天模板转换为模型训练时看到的固定词元序列。手动将字符串拼接成“用户:… 助手:…”等格式的根本问题是偏离了这种训练格式,削弱了模型的多步推理能力。然而,缓存仅依赖于生成的词元序列。如果手动拼接的前缀字节完全稳定,仍然可以被缓存。只有当前缀改变时,缓存才会失效,例如当动态内容插入其中时。\n>\n> 这三个结论背后的直觉很简单:在处理上下文时,大语言模型(LLM)会缓存已处理前缀的计算,这样下一个请求可以重用该工作。**如果前缀字节完全相同,就可以重用缓存的计算;如果前缀改变,之后的计算必须重新构建**。系统提示词和工具定义通常是这个前缀中最早且成本最高的部分;一旦它们改变,之后的缓存中间结果就会失效。\n>\n> 记住这三条原则,即使跳过下面的技术细节,您也可以正确设计智能体的上下文结构。以下内容是为想深入了解“为什么”的读者准备的。\n\n> **实验2-2 ★:注意力机制可视化**\n>\n> 在解释键值缓存之前,我们首先通过一个实验对模型内部的注意力机制建立直观理解——这是理解键值缓存为何有效以及为何对上下文设计有严格要求的基础。\n>\n> **什么是注意力机制?** 考虑一个具体例子。假设模型正在处理中文句子“北京的天气怎么样”(“How's the weather in Beijing?”),其单词为“北京”(Beijing)、“的”(一个所有格助词,类似于“of”)、“天气”(weather)和“怎么样”(how is it)。当它读到“怎么样”时,模型需要决定:前面的哪些单词对理解“怎么样”最重要?\n>\n> 注意力机制使用三种类型的向量来决定哪些早期词元最相关:\n>\n> 表2-1总结了查询(Query)、键(Key)和值(Value)向量在注意力机制中的作用,帮助读者将抽象计算映射到例句“北京的天气怎么样”(“How's the weather in Beijing?”)上。\n>\n> 表2-1 注意力机制中查询、键和值的作用\n>\n> | 向量 | 含义 | 在这个例子中 |\n> | ------- | ------------------------------------------------------------ | ---------------------------------------------------------- |\n> | **查询** | 当前词元发出的“搜索请求” | “怎么样”(how is it)询问:哪个词元与我最相关? |\n> | **键** | 每个词元的“标签”,用于匹配搜索 | “北京”(Beijing)的标签倾向于“地名”;“天气”(weather)的标签倾向于“气象” |\n> | **值** | 每个词元的“内容”,在成功匹配时提取 | 匹配到“天气”(weather)后,提取其语义信息 |\n>\n> 简单来说,每个新词元会根据相关性给前面的词元打分,然后使用最相关的信息构建当前表示。\n>\n> 更具体地说,计算分为三个步骤。首先,“怎么样”生成自己的查询向量,代表当前词元在寻找什么。其次,使用点积将查询与每个前面词元的键进行比较,产生相关性得分;得分越高表示匹配越强。最后,这些得分成为注意力权重,用于计算值的加权和。权重越高的词元对最终表示的贡献越大,权重越低的词元贡献越小。\n>\n>\n> ![图2-6:注意力机制直观理解](images/fig2-6.svg)\n>\n>\n> 图2-6的上部展示了“怎么样”(how is it)与每个前面词元的匹配情况:最强匹配是“天气”(weather,0.55),与“北京”(Beijing,0.35)有一定相关性,与“的”(助词,0.05)几乎无关,剩余约0.05的权重分配给“怎么样”本身(图中未单独显示)——所有权重之和为1。最终输出主要借鉴了“天气”的信息,完全符合直觉。\n>\n> **注意力热力图**将每个词元与所有前面词元的注意力权重排列成矩阵。图2-6的下部展示了完整的热力图:每一行是一个查询(当前正在处理的词元),每一列是一个键(正在被关注的词元),颜色越深表示注意力权重越高。热力图呈三角形,因为模型从左到右生成文本:每个词元只能关注自己和前面的词元,不能关注尚未生成的内容。\n>\n> **为什么需要缓存键和值?** 观察热力图可以发现,每次生成新词元时,其查询必须与**所有**前面词元的键匹配,然后计算所有值的加权和。如果每次都从头重新计算所有K和V值,计算量会随上下文长度增长。键值缓存存储了已计算的K和V值,允许新词元直接重用它们——这是接下来要讨论的核心优化。\n>\n> 在对注意力机制有了基本理解后,我们现在可以通过`attention_visualization`实验观察真实模型的注意力分布。\n>\n>\n> ![图2-7:注意力热力图可视化](images/fig2-7.png)\n>\n>\n> 注意力热力图揭示了几个关键模式:\n>\n> 1. **注意力汇点**:序列的第一个词元通常吸收异常高的注意力权重,有时超过总注意力的70%。模型将这个位置用作“注意力汇点”,吸收不强烈对应任何其他特定词元的剩余注意力质量。换句话说,模型学会将原本未分配的注意力权重分配给第一个词元——这是一种系统现象,不是模型缺陷。\n>\n> 数学原因是注意力机制有一个硬性约束:所有注意力权重必须精确总和为100%(由称为softmax的数学函数保证),所以模型不能表示“不关注任何东西”。即使当前词元与任何前面词元都不很相关,这些权重也必须分配到某个地方。因此,模型需要一个稳定的容器来存储这个“剩余权重”,序列开头的固定位置成为最自然的选择。这是处理多个词元时softmax数学性质的必然结果。\n> 2. **推理三角模式**:模型的思维链(在`<think>`标签内)呈现出三角形自注意力模式:生成新的推理内容时,经常关注早期的推理内容和工具定义。\n> 3. **输出三角模式**:推理结束后的输出过程呈现另一个三角形,模型将推理轨迹作为提示生成答案。\n> 4. **位置偏差**[^lost-in-the-middle]:模型对上下文开头和结尾的信息召回准确率更高,而中间的信息更容易被忽略。因此,在设计上下文时,将最关键的信息放在开头或结尾是重要的实践原则。\n>\n> 这个实验表明,**长思维链生成和工具调用都高度依赖上下文学习**——模型基于输入中提供的指令和示例适应任务的能力,无需重新训练。关于上下文学习的内部机制及其对智能体架构设计的影响,见本章的上下文压缩部分。\n>\n\n[^lost-in-the-middle]: 刘等人的[\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/),《计算语言学协会会刊》(TACL),2024年。\n\n### 从API消息到模型词元:聊天模板\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第4/17部分]\n那一行时间戳使每次请求的KV缓存失效。系统提示现在每次都不同,迫使模型从头开始重新计算前缀的键值对(这里,“键”和“值”是注意力机制中的两种向量;下面的实验2-2直观地展示了它们的作用)。这种无形的成本在代理系统中反复出现:看似无害的一行代码可能会使整个推理管道的速度降低一个数量级。本节解释如何避免这些陷阱。\n\n> **技术说明**:本节涉及Transformer注意力机制和KV缓存的内部原理,是本书技术密度较高的部分之一。如果你不熟悉这些底层机制,**可以跳过详细原理,记住以下三个核心结论**:\n>\n> 1. **一旦系统提示和工具定义确定,不要更改它们**。任何修改,即使添加一个空格,都会使整个缓存失效,并可能使延迟和成本成倍增加(具体幅度取决于模型和配置)。\n> 2. **始终将动态信息追加到末尾**——更改时间戳和用户状态等内容应作为新消息追加到对话末尾,而不是修改现有的系统提示。\n> 3. **使用标准API格式,不要手动连接消息**:结构化消息通过聊天模板转换为模型在训练期间看到的固定标记序列。手动将字符串连接成`\"用户:... 助手:...\"`等格式的根本问题是偏离了这种训练格式,削弱了模型的多步推理能力。然而,缓存仅取决于生成的标记序列。如果手动连接的前缀保持字节完全稳定,仍然可以被缓存。仅当前缀更改时,例如动态内容插入其中时,缓存才会失效。\n>\n> 这三个结论背后的直觉很简单:大语言模型在处理上下文时,会缓存已处理前缀的计算,因此下一次请求可以重用该工作。**如果前缀字节完全相同,缓存的计算可以重用;如果前缀更改,该点之后的计算必须重新构建**。系统提示和工具定义通常是该前缀中最早且最昂贵的部分;一旦它们更改,该点之后的缓存中间结果就会失效。\n>\n> 记住这三个原则,即使跳过下面的技术细节,你也可以正确设计代理的上下文结构。以下内容是为想深入了解“为什么”的读者准备的。\n\n> **实验2-2 ★:注意力机制可视化**\n>\n> 在解释KV缓存之前,我们首先通过一个实验对模型的内部注意力机制建立直观理解——这是理解为什么KV缓存有效以及为什么它对上下文设计提出严格要求的基础。\n>\n> **什么是注意力机制?** 考虑一个具体示例。假设模型正在处理中文句子“北京 的 天气 怎么样”(“How's the weather in Beijing?”),其单词为“北京”(Beijing)、“的”(一个所有格助词,类似“of”)、“天气”(weather)和“怎么样”(how is it)。当它读到“怎么样”时,模型需要决定:前面的哪些单词对理解“怎么样”最重要?\n>\n> 注意力机制使用三种类型的向量来决定哪些早期标记最相关:\n>\n> 表2-1总结了查询、键和值向量在注意力机制中的作用,帮助读者将抽象计算映射到示例句子“北京的天气怎么样”(“How's the weather in Beijing?”)。\n>\n> 表2-1 注意力机制中查询、键和值的作用\n>\n> | 向量 | 含义 | 在这个示例中 |\n> |--------|--------------------------------------|--------------------------------------|\n> | **查询** | 当前单词发出的“搜索请求” | “怎么样”(how is it)询问:哪个单词与我最相关? |\n> | **键** | 每个单词的“标签”,用于匹配搜索 | “北京”(Beijing)的标签倾向于“地名”;“天气”(weather)的标签倾向于“气象” |\n> | **值** | 成功匹配后提取的每个单词的“内容” | 匹配“天气”(weather)后,提取其语义信息 |\n>\n> 简单来说,每个新单词根据相关性给前面的单词打分,然后使用最相关的信息构建当前表示。\n>\n> 更具体地说,计算分为三步。首先,“怎么样”生成自己的查询向量,代表当前标记在寻找什么。其次,使用点积将查询与每个前面单词的键进行比较,产生相关性得分;得分越高表示匹配越强。最后,这些得分成为注意力权重,用于计算值的加权和。权重较高的单词对最终表示的贡献更大,权重较低的单词贡献较小。\n>\n>\n> ![图2-6:对注意力机制的直观理解](images/fig2-6.svg)\n>\n>\n> 图2-6的上部展示了“怎么样”(how is it)如何匹配每个前面的单词:最强匹配是“天气”(weather,0.55),与“北京”(Beijing,0.35)有一定相关性,与“的”(助词,0.05)几乎无关,剩余约0.05的权重分配给“怎么样”本身(图中未单独显示)——所有权重之和为1。最终输出主要借鉴了“天气”的信息,这完全符合直觉。\n>\n> **注意力热力图**将每个单词与所有前面单词之间的注意力权重排列成矩阵。图2-6的下部展示了完整的热力图:每一行是一个查询(当前正在处理的单词),每一列是一个键(正在被关注的单词),颜色越深表示注意力权重越高。热力图是三角形的,因为模型从左到右生成文本:每个单词只能关注自己和前面的单词,不能关注尚未生成的内容。\n>\n> **为什么需要缓存键和值?** 观察热力图发现,每次生成一个新单词,其查询必须与**所有**前面单词的键匹配,然后计算所有值的加权和。如果每次都从头重新计算所有K和V值,计算量会随着上下文长度增长。KV缓存存储已计算的K和V值,允许新单词直接重用它们——这是接下来讨论的核心优化。\n>\n> 在对注意力机制有了基本理解后,我们现在可以通过`attention_visualization`实验观察真实模型的注意力分布。\n>\n>\n> ![图2-7:注意力热力图可视化](images/fig2-7.png)\n>\n>\n> 注意力热力图揭示了几个关键模式:\n>\n> 1. **注意力汇聚点**:序列的第一个标记通常吸收异常高的注意力权重,有时超过总注意力的70%。模型将这个位置用作“注意力汇聚点”,吸收不强烈对应任何其他特定标记的剩余注意力质量。换句话说,模型学会将否则未分配的注意力权重分配给第一个标记——这是一种系统现象,不是模型缺陷。\n>\n> 数学原因是注意力机制有一个硬性约束:所有注意力权重必须精确总和为100%(由称为softmax的数学函数保证),因此模型无法表达“不关注任何东西”。即使当前单词与任何前面单词都不太相关,这些权重也必须分配到某个地方。因此,模型需要一个稳定的容器来存储这个“剩余权重”,序列开头的固定位置成为最自然的选择。这是处理许多标记时softmax数学性质的必然结果。\n> 2. **推理三角形模式**:模型的思维链(在`<think>`标签内)呈现出三角形自注意力模式:生成新推理内容时,它频繁关注早期推理内容和工具定义。\n> 3. **输出三角形模式**:推理结束后的输出过程显示另一个三角形,模型使用推理轨迹作为提示生成答案。\n> 4. **位置偏差**[^lost-in-the-middle]:模型对上下文开头和结尾的信息召回准确率更高,而中间的信息更容易被忽略。因此,在设计上下文时,将最关键的信息放在开头或结尾是重要的实践原则。\n>\n> 这个实验表明,**长思维链生成和工具调用都高度依赖上下文学习**——模型根据输入中提供的指令和示例适应任务的能力,无需重新训练。关于上下文学习的内部机制及其对代理架构设计的影响,见本章的上下文压缩部分。\n>\n\n[^lost-in-the-middle]: Liu等人的[\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/),《计算语言学协会会刊》,2024年。\n\n### 从API消息到模型标记:聊天模板"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853907295663ea4fa82759f0cae4f2ffa723d4f9dd6bac681",
"model": "doubao-seed-1-6-250615",
"created": 1785390750,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, '键值缓存(KV Cache' retains the original term with parentheses, and '智能体(Agent' is consistently translated with clarification. The technical note's three conclusions are fully preserved, including details like '字节完全相同' (byte-for-byte identical).\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent, but some phrases are slightly verbose, e.g., '大语言模型(LLM)会缓存已处理前缀的计算' (LLM caches computation for the prefix it has already processed) is accurate but longer than necessary. '智能体' (Agent) is natural in context.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms: '键值缓存(KV Cache', '注意力机制', '词元(token', '热力图', '上下文长度' are used consistently. '智能体(Agent' is properly clarified on first use.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Headings, blockquotes, tables, images (e.g., ![图2-6:注意力机制直观理解](images/fig2-6.svg)), code fences (`attention_visualization`), and footnotes ([^lost-in-the-middle]) are all preserved correctly.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions or inventions. Key technical claims like '字节完全相同' (byte-for-byte identical) and '前缀更改时缓存失效' are accurately translated. The three core conclusions are完整保留 (fully preserved).\"},\"fluency\":{\"score\":5,\"evidence\":\"More concise and natural phrasing: '使延迟和成本成倍增加' (multiply latency and increase costs) flows better than X's '成倍增加时延和成本'. '可以跳过详细原理' (you can skip) is more natural than X's '您可以跳过'. Consistent use of '你' instead of '您' maintains a coherent tone.\"},\"terminology\":{\"score\":4,\"evidence\":\"Most terms are correct, but '代理' (Agent) lacks clarification on first use, potentially confusing non-technical readers. '标记(token' is used consistently, but '连接' (concatenate) is less precise than X's '拼接' (splice/concatenate) for message assembly.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown elements are fully preserved: headings, blockquotes, tables, images (e.g., ![图2-7:注意力热力图可视化](images/fig2-7.png)), code fences (`\"\n ,\"用户:... 助手:...\"\n \t\t:\n 1.1111111111111112\n \t\t}\n ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"}} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"}",
"usage": {
"prompt_tokens": 6245,
"completion_tokens": 769,
"total_tokens": 7014
}
},
"latency_milliseconds": 21425,
"validation": {
"valid": false,
"error_type": "ValueError",
"error": "judge variant Y must contain all rubric dimensions"
}
},
{
"attempt": 2,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nThat one timestamp line invalidated the KV Cache on every request. The system prompt was now different each time, forcing the model to recompute the key-value pairs for the prefix from scratch (here, \"Key\" and \"Value\" are two types of vectors in the attention mechanism; Experiment 2-2 below visually demonstrates their roles). This kind of invisible cost appears repeatedly in Agent systems: a seemingly harmless line of code can slow down the entire inference pipeline by an order of magnitude. This section explains how to avoid these pitfalls.\n\n> **Technical Note**: This section involves the internal principles of the Transformer attention mechanism and KV Cache, making it one of the most technically dense parts of the book. If you are not familiar with these underlying mechanisms, **you can skip the detailed principles and remember the following three core conclusions**:\n>\n> 1. **Once the system prompt and tool definitions are finalized, do not change them.** Any modification, even adding a single space, will invalidate the entire cache and can multiply latency and increase costs (the exact magnitude depends on the model and configuration).\n> 2. **Always append dynamic information to the end**—changing content like timestamps and user status should be appended as new messages at the end of the conversation, not by modifying the existing system prompt.\n> 3. **Use the standard API format; do not manually concatenate messages**: Structured messages are translated by the Chat Template into a fixed token sequence that the model saw during training. The fundamental problem with manually concatenating strings into formats like `\"USER: ... ASSISTANT: ...\"` is that it deviates from this training format, weakening the model's multi-step reasoning ability. Caching, however, depends only on the resulting token sequence. A manually concatenated prefix can still be cached if it remains byte-for-byte stable. The cache is invalidated only when that prefix changes, for example, when dynamic content is inserted into it.\n>\n> The intuition behind these three conclusions is simple: when processing context, an LLM caches computation for the prefix it has already processed, so the next request can reuse that work. **If the prefix is byte-for-byte identical, the cached computation can be reused; if the prefix changes, the computation after that point must be rebuilt.** The system prompt and tool definitions are usually the earliest and most expensive part of this prefix; once they change, the cached intermediate results after that point are invalidated.\n>\n> Remember these three principles, and even if you skip the technical details below, you can correctly design the context structure of an Agent. The following content is for readers who want to delve deeper into the \"why.\"\n\n> **Experiment 2-2 ★: Attention Mechanism Visualization**\n>\n> Before explaining KV Cache, we first build an intuitive understanding of the model's internal attention mechanism through an experiment—this is the foundation for understanding why KV Cache is effective and why it imposes strict requirements on context design.\n>\n> **What is the Attention Mechanism?** Consider a concrete example. Suppose the model is processing the Chinese sentence \"北京 的 天气 怎么样\" (\"How's the weather in Beijing?\"), whose words are \"北京\" (Beijing), \"的\" (a possessive particle, like \"of\"), \"天气\" (weather), and \"怎么样\" (how is it). When it reads \"怎么样\", the model needs to decide: which of the preceding words are most important for understanding \"怎么样\"?\n>\n> The attention mechanism uses three types of vectors to decide which earlier tokens are most relevant:\n>\n> Table 2-1 summarizes the roles of the Query, Key, and Value vectors in the attention mechanism, helping readers map the abstract computation onto the example sentence \"北京的天气怎么样\" (\"How's the weather in Beijing?\").\n>\n> Table 2-1 Roles of Query, Key, and Value in the Attention Mechanism\n>\n> | Vector | Meaning | In this example |\n> |-------|-----------------------------------------|-----------------------------------------------|\n> | **Query** | The \"search request\" issued by the current word | \"怎么样\" (how is it) asks: which word is most relevant to me? |\n> | **Key** | The \"label\" of each word, used for matching the search | The label of \"北京\" (Beijing) leans toward \"place name\"; the label of \"天气\" (weather) leans toward \"meteorology\" |\n> | **Value** | The \"content\" of each word, extracted upon a successful match | After matching \"天气\" (weather), extract its semantic information |\n>\n> In simplified terms, each new word scores the preceding words by relevance, then uses the most relevant information to build its current representation.\n>\n> More specifically, the computation has three steps. First, \"怎么样\" generates its own Query vector, representing what the current token is looking for. Second, the Query is compared with the Key of each preceding word using a dot product, producing a relevance score; higher scores indicate stronger matches. Finally, these scores become attention weights, which are used to compute a weighted sum of the Values. Words with higher weights contribute more to the final representation, while words with lower weights contribute less.\n>\n>\n> ![Figure 2-6: Intuitive Understanding of the Attention Mechanism](images/fig2-6.svg)\n>\n>\n> The upper part of Figure 2-6 shows how \"怎么样\" (how is it) matches each preceding word: the strongest match is with \"天气\" (weather, 0.55), there is some relevance to \"北京\" (Beijing, 0.35), almost none to \"的\" (the particle, 0.05), and the remaining weight of about 0.05 goes to \"怎么样\" itself (not shown separately in the figure)—all weights sum to 1. The final output draws mainly on the information from \"天气\", which matches intuition exactly.\n>\n> An **attention heatmap** arranges the attention weights between each word and all preceding words into a matrix. The lower part of Figure 2-6 shows the complete heatmap: each row is a Query (the word currently being processed), each column is a Key (the word being attended to), and darker cells indicate higher attention weights. The heatmap is triangular because the model generates text from left to right: each word can attend only to itself and the words before it, not to content that has yet to be generated.\n>\n> **Why do Key and Value need to be cached?** Observing the heatmap reveals that every time a new word is generated, its Query must be matched against the Keys of **all** preceding words, and then a weighted sum of all Values is computed. If all K and V values were recalculated from scratch each time, the computation would grow with the context length. The KV Cache stores the already computed K and V values, allowing new words to directly reuse them — this is the core optimization discussed next.\n>\n> With a basic understanding of the attention mechanism, we can now observe the attention distribution of a real model through the `attention_visualization` experiment.\n>\n>\n> ![Figure 2-7: Attention Heatmap Visualization](images/fig2-7.png)\n>\n>\n> The attention heatmap reveals several key patterns:\n>\n> 1. **Attention Sink**: The first token of the sequence often absorbs an abnormally high amount of attention weight, sometimes exceeding 70% of the total attention. The model uses this position as an \"Attention Sink\" to absorb residual attention mass that does not strongly correspond to any other specific token. In other words, the model learns to assign otherwise unallocated attention weight to the first token — this is a systematic phenomenon, not a model defect.\n>\n> The mathematical reason is that the attention mechanism has a hard constraint: all attention weights must sum to exactly 100% (guaranteed by a mathematical function called softmax), so the model cannot express \"not attending to anything.\" Even if the current word is not very relevant to any preceding word, these weights must be allocated somewhere. The model therefore needs a stable container for this \"residual weight,\" and the fixed position at the beginning of the sequence becomes the most natural choice. This is an inevitable consequence of the mathematical properties of softmax when processing many tokens.\n> 2. **Reasoning Triangle Pattern**: The model's chain of thought (within `<think>` tags) exhibits a triangular self-attention pattern: when generating new reasoning content, it frequently attends to earlier reasoning content and tool definitions.\n> 3. **Output Triangle Pattern**: The output process after reasoning ends shows another triangle, where the model uses the reasoning trace as a prompt to generate the answer.\n> 4. **Position Bias**[^lost-in-the-middle]: The model has higher recall accuracy for information at the beginning and end of the context, while information in the middle is more likely to be overlooked. Therefore, when designing the context, placing the most critical information at the beginning or end is an important practical principle.\n>\n> This experiment shows that **long chain-of-thought generation and tool calling both depend heavily on in-context learning** — the model's ability to adapt to a task based on the instructions and examples provided in the input, without retraining. For the internal mechanism of in-context learning and its implications for Agent architecture design, see the Context Compression section of this chapter.\n>\n\n[^lost-in-the-middle]: Liu et al. [\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/), TACL, 2024.\n\n### From API Messages to Model Tokens: Chat Template\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第4/17部分]\n那一行时间戳导致每次请求时键值缓存(KV Cache)都失效。系统提示词现在每次都不同,迫使模型从头开始重新计算前缀的键值对(这里“键”和“值”是注意力机制中的两种向量;下面的实验2-2直观展示了它们的作用)。这种无形的成本在智能体(Agent)系统中反复出现:看似无害的一行代码可能会使整个推理流水线的速度降低一个数量级。本节将解释如何避免这些陷阱。\n\n> **技术说明**:本节涉及Transformer注意力机制和键值缓存的内部原理,是本书技术密度较高的部分之一。如果您不熟悉这些底层机制,**您可以跳过详细原理,记住以下三个核心结论**:\n>\n> 1. **一旦系统提示词和工具定义确定,就不要更改它们**。任何修改,即使是添加一个空格,都会使整个缓存失效,并可能成倍增加时延和成本(具体幅度取决于模型和配置)。\n> 2. **始终在末尾追加动态信息**——时间戳、用户状态等内容的更改应作为新消息追加在对话末尾,而不是修改现有的系统提示词。\n> 3. **使用标准API格式,不要手动拼接消息**:结构化消息由聊天模板转换为模型训练时看到的固定词元序列。手动将字符串拼接成“用户:… 助手:…”等格式的根本问题是偏离了这种训练格式,削弱了模型的多步推理能力。然而,缓存仅依赖于生成的词元序列。如果手动拼接的前缀字节完全稳定,仍然可以被缓存。只有当前缀改变时,缓存才会失效,例如当动态内容插入其中时。\n>\n> 这三个结论背后的直觉很简单:在处理上下文时,大语言模型(LLM)会缓存已处理前缀的计算,这样下一个请求可以重用该工作。**如果前缀字节完全相同,就可以重用缓存的计算;如果前缀改变,之后的计算必须重新构建**。系统提示词和工具定义通常是这个前缀中最早且成本最高的部分;一旦它们改变,之后的缓存中间结果就会失效。\n>\n> 记住这三条原则,即使跳过下面的技术细节,您也可以正确设计智能体的上下文结构。以下内容是为想深入了解“为什么”的读者准备的。\n\n> **实验2-2 ★:注意力机制可视化**\n>\n> 在解释键值缓存之前,我们首先通过一个实验对模型内部的注意力机制建立直观理解——这是理解键值缓存为何有效以及为何对上下文设计有严格要求的基础。\n>\n> **什么是注意力机制?** 考虑一个具体例子。假设模型正在处理中文句子“北京的天气怎么样”(“How's the weather in Beijing?”),其单词为“北京”(Beijing)、“的”(一个所有格助词,类似于“of”)、“天气”(weather)和“怎么样”(how is it)。当它读到“怎么样”时,模型需要决定:前面的哪些单词对理解“怎么样”最重要?\n>\n> 注意力机制使用三种类型的向量来决定哪些早期词元最相关:\n>\n> 表2-1总结了查询(Query)、键(Key)和值(Value)向量在注意力机制中的作用,帮助读者将抽象计算映射到例句“北京的天气怎么样”(“How's the weather in Beijing?”)上。\n>\n> 表2-1 注意力机制中查询、键和值的作用\n>\n> | 向量 | 含义 | 在这个例子中 |\n> | ------- | ------------------------------------------------------------ | ---------------------------------------------------------- |\n> | **查询** | 当前词元发出的“搜索请求” | “怎么样”(how is it)询问:哪个词元与我最相关? |\n> | **键** | 每个词元的“标签”,用于匹配搜索 | “北京”(Beijing)的标签倾向于“地名”;“天气”(weather)的标签倾向于“气象” |\n> | **值** | 每个词元的“内容”,在成功匹配时提取 | 匹配到“天气”(weather)后,提取其语义信息 |\n>\n> 简单来说,每个新词元会根据相关性给前面的词元打分,然后使用最相关的信息构建当前表示。\n>\n> 更具体地说,计算分为三个步骤。首先,“怎么样”生成自己的查询向量,代表当前词元在寻找什么。其次,使用点积将查询与每个前面词元的键进行比较,产生相关性得分;得分越高表示匹配越强。最后,这些得分成为注意力权重,用于计算值的加权和。权重越高的词元对最终表示的贡献越大,权重越低的词元贡献越小。\n>\n>\n> ![图2-6:注意力机制直观理解](images/fig2-6.svg)\n>\n>\n> 图2-6的上部展示了“怎么样”(how is it)与每个前面词元的匹配情况:最强匹配是“天气”(weather,0.55),与“北京”(Beijing,0.35)有一定相关性,与“的”(助词,0.05)几乎无关,剩余约0.05的权重分配给“怎么样”本身(图中未单独显示)——所有权重之和为1。最终输出主要借鉴了“天气”的信息,完全符合直觉。\n>\n> **注意力热力图**将每个词元与所有前面词元的注意力权重排列成矩阵。图2-6的下部展示了完整的热力图:每一行是一个查询(当前正在处理的词元),每一列是一个键(正在被关注的词元),颜色越深表示注意力权重越高。热力图呈三角形,因为模型从左到右生成文本:每个词元只能关注自己和前面的词元,不能关注尚未生成的内容。\n>\n> **为什么需要缓存键和值?** 观察热力图可以发现,每次生成新词元时,其查询必须与**所有**前面词元的键匹配,然后计算所有值的加权和。如果每次都从头重新计算所有K和V值,计算量会随上下文长度增长。键值缓存存储了已计算的K和V值,允许新词元直接重用它们——这是接下来要讨论的核心优化。\n>\n> 在对注意力机制有了基本理解后,我们现在可以通过`attention_visualization`实验观察真实模型的注意力分布。\n>\n>\n> ![图2-7:注意力热力图可视化](images/fig2-7.png)\n>\n>\n> 注意力热力图揭示了几个关键模式:\n>\n> 1. **注意力汇点**:序列的第一个词元通常吸收异常高的注意力权重,有时超过总注意力的70%。模型将这个位置用作“注意力汇点”,吸收不强烈对应任何其他特定词元的剩余注意力质量。换句话说,模型学会将原本未分配的注意力权重分配给第一个词元——这是一种系统现象,不是模型缺陷。\n>\n> 数学原因是注意力机制有一个硬性约束:所有注意力权重必须精确总和为100%(由称为softmax的数学函数保证),所以模型不能表示“不关注任何东西”。即使当前词元与任何前面词元都不很相关,这些权重也必须分配到某个地方。因此,模型需要一个稳定的容器来存储这个“剩余权重”,序列开头的固定位置成为最自然的选择。这是处理多个词元时softmax数学性质的必然结果。\n> 2. **推理三角模式**:模型的思维链(在`<think>`标签内)呈现出三角形自注意力模式:生成新的推理内容时,经常关注早期的推理内容和工具定义。\n> 3. **输出三角模式**:推理结束后的输出过程呈现另一个三角形,模型将推理轨迹作为提示生成答案。\n> 4. **位置偏差**[^lost-in-the-middle]:模型对上下文开头和结尾的信息召回准确率更高,而中间的信息更容易被忽略。因此,在设计上下文时,将最关键的信息放在开头或结尾是重要的实践原则。\n>\n> 这个实验表明,**长思维链生成和工具调用都高度依赖上下文学习**——模型基于输入中提供的指令和示例适应任务的能力,无需重新训练。关于上下文学习的内部机制及其对智能体架构设计的影响,见本章的上下文压缩部分。\n>\n\n[^lost-in-the-middle]: 刘等人的[\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/),《计算语言学协会会刊》(TACL),2024年。\n\n### 从API消息到模型词元:聊天模板\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第4/17部分]\n那一行时间戳使每次请求的KV缓存失效。系统提示现在每次都不同,迫使模型从头开始重新计算前缀的键值对(这里,“键”和“值”是注意力机制中的两种向量;下面的实验2-2直观地展示了它们的作用)。这种无形的成本在代理系统中反复出现:看似无害的一行代码可能会使整个推理管道的速度降低一个数量级。本节解释如何避免这些陷阱。\n\n> **技术说明**:本节涉及Transformer注意力机制和KV缓存的内部原理,是本书技术密度较高的部分之一。如果你不熟悉这些底层机制,**可以跳过详细原理,记住以下三个核心结论**:\n>\n> 1. **一旦系统提示和工具定义确定,不要更改它们**。任何修改,即使添加一个空格,都会使整个缓存失效,并可能使延迟和成本成倍增加(具体幅度取决于模型和配置)。\n> 2. **始终将动态信息追加到末尾**——更改时间戳和用户状态等内容应作为新消息追加到对话末尾,而不是修改现有的系统提示。\n> 3. **使用标准API格式,不要手动连接消息**:结构化消息通过聊天模板转换为模型在训练期间看到的固定标记序列。手动将字符串连接成`\"用户:... 助手:...\"`等格式的根本问题是偏离了这种训练格式,削弱了模型的多步推理能力。然而,缓存仅取决于生成的标记序列。如果手动连接的前缀保持字节完全稳定,仍然可以被缓存。仅当前缀更改时,例如动态内容插入其中时,缓存才会失效。\n>\n> 这三个结论背后的直觉很简单:大语言模型在处理上下文时,会缓存已处理前缀的计算,因此下一次请求可以重用该工作。**如果前缀字节完全相同,缓存的计算可以重用;如果前缀更改,该点之后的计算必须重新构建**。系统提示和工具定义通常是该前缀中最早且最昂贵的部分;一旦它们更改,该点之后的缓存中间结果就会失效。\n>\n> 记住这三个原则,即使跳过下面的技术细节,你也可以正确设计代理的上下文结构。以下内容是为想深入了解“为什么”的读者准备的。\n\n> **实验2-2 ★:注意力机制可视化**\n>\n> 在解释KV缓存之前,我们首先通过一个实验对模型的内部注意力机制建立直观理解——这是理解为什么KV缓存有效以及为什么它对上下文设计提出严格要求的基础。\n>\n> **什么是注意力机制?** 考虑一个具体示例。假设模型正在处理中文句子“北京 的 天气 怎么样”(“How's the weather in Beijing?”),其单词为“北京”(Beijing)、“的”(一个所有格助词,类似“of”)、“天气”(weather)和“怎么样”(how is it)。当它读到“怎么样”时,模型需要决定:前面的哪些单词对理解“怎么样”最重要?\n>\n> 注意力机制使用三种类型的向量来决定哪些早期标记最相关:\n>\n> 表2-1总结了查询、键和值向量在注意力机制中的作用,帮助读者将抽象计算映射到示例句子“北京的天气怎么样”(“How's the weather in Beijing?”)。\n>\n> 表2-1 注意力机制中查询、键和值的作用\n>\n> | 向量 | 含义 | 在这个示例中 |\n> |--------|--------------------------------------|--------------------------------------|\n> | **查询** | 当前单词发出的“搜索请求” | “怎么样”(how is it)询问:哪个单词与我最相关? |\n> | **键** | 每个单词的“标签”,用于匹配搜索 | “北京”(Beijing)的标签倾向于“地名”;“天气”(weather)的标签倾向于“气象” |\n> | **值** | 成功匹配后提取的每个单词的“内容” | 匹配“天气”(weather)后,提取其语义信息 |\n>\n> 简单来说,每个新单词根据相关性给前面的单词打分,然后使用最相关的信息构建当前表示。\n>\n> 更具体地说,计算分为三步。首先,“怎么样”生成自己的查询向量,代表当前标记在寻找什么。其次,使用点积将查询与每个前面单词的键进行比较,产生相关性得分;得分越高表示匹配越强。最后,这些得分成为注意力权重,用于计算值的加权和。权重较高的单词对最终表示的贡献更大,权重较低的单词贡献较小。\n>\n>\n> ![图2-6:对注意力机制的直观理解](images/fig2-6.svg)\n>\n>\n> 图2-6的上部展示了“怎么样”(how is it)如何匹配每个前面的单词:最强匹配是“天气”(weather,0.55),与“北京”(Beijing,0.35)有一定相关性,与“的”(助词,0.05)几乎无关,剩余约0.05的权重分配给“怎么样”本身(图中未单独显示)——所有权重之和为1。最终输出主要借鉴了“天气”的信息,这完全符合直觉。\n>\n> **注意力热力图**将每个单词与所有前面单词之间的注意力权重排列成矩阵。图2-6的下部展示了完整的热力图:每一行是一个查询(当前正在处理的单词),每一列是一个键(正在被关注的单词),颜色越深表示注意力权重越高。热力图是三角形的,因为模型从左到右生成文本:每个单词只能关注自己和前面的单词,不能关注尚未生成的内容。\n>\n> **为什么需要缓存键和值?** 观察热力图发现,每次生成一个新单词,其查询必须与**所有**前面单词的键匹配,然后计算所有值的加权和。如果每次都从头重新计算所有K和V值,计算量会随着上下文长度增长。KV缓存存储已计算的K和V值,允许新单词直接重用它们——这是接下来讨论的核心优化。\n>\n> 在对注意力机制有了基本理解后,我们现在可以通过`attention_visualization`实验观察真实模型的注意力分布。\n>\n>\n> ![图2-7:注意力热力图可视化](images/fig2-7.png)\n>\n>\n> 注意力热力图揭示了几个关键模式:\n>\n> 1. **注意力汇聚点**:序列的第一个标记通常吸收异常高的注意力权重,有时超过总注意力的70%。模型将这个位置用作“注意力汇聚点”,吸收不强烈对应任何其他特定标记的剩余注意力质量。换句话说,模型学会将否则未分配的注意力权重分配给第一个标记——这是一种系统现象,不是模型缺陷。\n>\n> 数学原因是注意力机制有一个硬性约束:所有注意力权重必须精确总和为100%(由称为softmax的数学函数保证),因此模型无法表达“不关注任何东西”。即使当前单词与任何前面单词都不太相关,这些权重也必须分配到某个地方。因此,模型需要一个稳定的容器来存储这个“剩余权重”,序列开头的固定位置成为最自然的选择。这是处理许多标记时softmax数学性质的必然结果。\n> 2. **推理三角形模式**:模型的思维链(在`<think>`标签内)呈现出三角形自注意力模式:生成新推理内容时,它频繁关注早期推理内容和工具定义。\n> 3. **输出三角形模式**:推理结束后的输出过程显示另一个三角形,模型使用推理轨迹作为提示生成答案。\n> 4. **位置偏差**[^lost-in-the-middle]:模型对上下文开头和结尾的信息召回准确率更高,而中间的信息更容易被忽略。因此,在设计上下文时,将最关键的信息放在开头或结尾是重要的实践原则。\n>\n> 这个实验表明,**长思维链生成和工具调用都高度依赖上下文学习**——模型根据输入中提供的指令和示例适应任务的能力,无需重新训练。关于上下文学习的内部机制及其对代理架构设计的影响,见本章的上下文压缩部分。\n>\n\n[^lost-in-the-middle]: Liu等人的[\"Lost in the Middle: How Language Models Use Long Contexts\"](https://aclanthology.org/2024.tacl-1.9/),《计算语言学协会会刊》,2024年。\n\n### 从API消息到模型标记:聊天模板"
},
{
"role": "assistant",
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, '键值缓存(KV Cache' retains the original term with parentheses, and '智能体(Agent' is consistently translated with clarification. The technical note's three conclusions are fully preserved, including details like '字节完全相同' (byte-for-byte identical).\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent, but some phrases are slightly verbose, e.g., '大语言模型(LLM)会缓存已处理前缀的计算' (LLM caches computation for the prefix it has already processed) is accurate but longer than necessary. '智能体' (Agent) is natural in context.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms: '键值缓存(KV Cache', '注意力机制', '词元(token', '热力图', '上下文长度' are used consistently. '智能体(Agent' is properly clarified on first use.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Headings, blockquotes, tables, images (e.g., ![图2-6:注意力机制直观理解](images/fig2-6.svg)), code fences (`attention_visualization`), and footnotes ([^lost-in-the-middle]) are all preserved correctly.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions or inventions. Key technical claims like '字节完全相同' (byte-for-byte identical) and '前缀更改时缓存失效' are accurately translated. The three core conclusions are完整保留 (fully preserved).\"},\"fluency\":{\"score\":5,\"evidence\":\"More concise and natural phrasing: '使延迟和成本成倍增加' (multiply latency and increase costs) flows better than X's '成倍增加时延和成本'. '可以跳过详细原理' (you can skip) is more natural than X's '您可以跳过'. Consistent use of '你' instead of '您' maintains a coherent tone.\"},\"terminology\":{\"score\":4,\"evidence\":\"Most terms are correct, but '代理' (Agent) lacks clarification on first use, potentially confusing non-technical readers. '标记(token' is used consistently, but '连接' (concatenate) is less precise than X's '拼接' (splice/concatenate) for message assembly.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown elements are fully preserved: headings, blockquotes, tables, images (e.g., ![图2-7:注意力热力图可视化](images/fig2-7.png)), code fences (`\"\n ,\"用户:... 助手:...\"\n \t\t:\n 1.1111111111111112\n \t\t}\n ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"}} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"}"
},
{
"role": "user",
"content": "Your previous JSON failed validation: judge variant Y must contain all rubric dimensions. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys."
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853907509923ea4fa82759f0cae4f2ffa723d4f9dd651402c",
"model": "doubao-seed-1-6-250615",
"created": 1785390765,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, '键值缓存(KV Cache' retains the original term with parentheses, and '智能体(Agent' is consistently translated with clarification. The technical note's three conclusions are fully preserved, including details like '字节完全相同' (byte-for-byte identical).\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent, but some phrases are slightly verbose, e.g., '大语言模型(LLM)会缓存已处理前缀的计算' (LLM caches computation for the prefix it has already processed) is accurate but longer than necessary. '智能体' (Agent) is natural in context.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms: '键值缓存(KV Cache', '注意力机制', '词元(token', '热力图', '上下文长度' are used consistently. '智能体(Agent' is properly clarified on first use.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Headings, blockquotes, tables, images (e.g., ![图2-6:注意力机制直观理解](images/fig2-6.svg)), code fences (`attention_visualization`), and footnotes ([^lost-in-the-middle]) are all preserved correctly.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions or inventions. Key technical claims like '字节完全相同' (byte-for-byte identical) and '前缀更改时缓存失效' are accurately translated. The three core conclusions are fully preserved.\"},\"fluency\":{\"score\":5,\"evidence\":\"More concise and natural phrasing: '使延迟和成本成倍增加' (multiply latency and increase costs) flows better than X's '成倍增加时延和成本'. '可以跳过详细原理' (you can skip) is more natural than X's '您可以跳过'. Consistent use of '你' instead of '您' maintains a coherent tone.\"},\"terminology\":{\"score\":4,\"evidence\":\"Most terms are correct, but '代理' (Agent) lacks clarification on first use, potentially confusing non-technical readers. '标记(token' is used consistently, but '连接' (concatenate) is less precise than X's '拼接' (splice/concatenate) for message assembly.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown elements are fully preserved: headings, blockquotes, tables, images (e.g., ![图2-7:注意力热力图可视化](images/fig2-7.png)), code fences (`\\\"USER: ... ASSISTANT: ...\\\"`), and footnotes ([^lost-in-the-middle]) are all correctly maintained.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency with more natural phrasing (e.g., '使延迟和成本成倍增加' vs. X's '成倍增加时延和成本') and maintains consistent tone. While X has slightly better terminology clarification ('智能体(Agent'), Y's concise expression makes it more readable without sacrificing accuracy or technical fidelity.\"}",
"usage": {
"prompt_tokens": 7081,
"completion_tokens": 746,
"total_tokens": 7827
}
},
"latency_milliseconds": 14644,
"validation": {
"valid": true
}
}
]
}