Files
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

89 lines
65 KiB
JSON
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n- **Wasted tokens**: Most content is irrelevant to the current task.\n- **Diluted attention**: Too much irrelevant information in the context dilutes the model's attention to key content (the context compression section later in this chapter discusses this in detail under the concept of \"context rot\").\n\nThis is the natural evolution from static prompt engineering to dynamic prompts: **instead of loading all knowledge into the Agent at once, allow it to load knowledge on demand**. The Agent Skills system is the engineering implementation of this idea.\n\n### Skills: Composable Units of Domain Capability\n\nThe core idea of Agent Skills is to modularize the Agent's capabilities into independent, loadable knowledge packages[^ch2-3]. Each Skill is essentially a collection of prompts and files containing specialized domain guidance, like an operating manual for a specific task. Unlike the traditional approach of placing all instructions into a single system prompt, Skills use Progressive Disclosure: first show the Agent a table-of-contents summary, then load the full content only when needed. Instead of loading every domain manual into context at once, the framework provides a directory and lets the Agent retrieve the relevant manual as needed.\n\n[^ch2-3]: Anthropic, \"Equipping Agents for the Real World with Agent Skills\", 2025.\n\n**Layer 1 (Metadata)**: Each Skill must include a `SKILL.md` file that starts with YAML frontmatter (a metadata block at the top of the file delimited by `---`, similar to a book's copyright page), containing `name` and `description` fields. The Agent framework scans all installed Skills at startup and injects their `name` and `description` into the dialogue context. This usually costs only a few hundred tokens, and the trade-offs around injection location are discussed in the next subsection. The goal is to let the Agent discover which specialized capabilities are available without loading all Skill content into context.\n\nRouting depends heavily on the metadata's `description` field. It should be concise enough to keep the always-loaded token count low, but written as a routing rule rather than a feature summary. The clearest pattern is \"Use when / Do not use when,\" supported by **negative examples** that identify situations in which the Skill should not be triggered. Negative examples are not optional; they are essential to accurate Skill routing. Broad descriptions such as \"help with backend\" activate on unrelated tasks, while explicit exclusions make routing substantially more precise. For routing purposes, \"when to use me\" matters far more than \"what I can do.\"\n\n**Layer 2 (Core Workflow)**: When the Agent determines that a specific Skill is needed for a task, it loads the complete `SKILL.md` via a dedicated Skill tool, and the content appears in the conversation history as a tool result. Using the PPTX Skill[^ch2-4] as an example, it contains the core workflow for handling PowerPoint files: how to extract text via markitdown (Microsoft's open-source document-to-Markdown tool), how to unzip the PPTX file to access the raw XML structure, and the path conventions for key files.\n\n[^ch2-4]: Anthropic, \"PPTX Skill\", 2025. https://github.com/anthropics/skills/\n\n**Layer 3 (Details)**: File references allow deeper navigation into more detailed sub-documents. The main file references `html2pptx.md` (detailed workflow for creating PowerPoint from HTML templates), `reference.md` (format technical details), and others. The Agent selectively reads relevant sub-documents based on specific needs.\n\nSkills not only contain instructional documentation but can also bundle executable code tools and template files—turning them from pure knowledge transfer into operational capabilities.\n\nThe value of Skills lies not only in context management but also in providing a sustainable path for accumulating domain knowledge. Each Skill is a self-contained knowledge module that can be independently developed, tested, version-controlled, and shared. This modularity transforms Agent capability expansion from centralized system prompt editing into a distributed Skill ecosystem, similar in spirit to package managers such as Python's pip or Node.js's npm. Each Skill encapsulates best practices for a specific domain. Anthropic's official Skills repository already covers document processing (PPTX, PDF, DOCX), data analysis, code generation, and other domains, allowing developers to use, customize, or create entirely new Skills.\n\nThis reveals an important principle for Agent developers: **when choosing an Agent interaction mode, align with the interaction patterns the model and API are designed to support**. When building Agents with Claude, fully leverage Skills and structured system prompts; when using other models, follow the conventions optimized by that model vendor. The Agent usage patterns promoted by foundation model companies often reflect the modes those models are trained and evaluated to support.\n\n### Skills Implementation Methods and Trade-offs\n\nAfter defining Skills, the next question is a concrete engineering problem: where in the context should Skill content be placed? This design decision directly affects KV Cache efficiency and the model's ability to follow the Skill's instructions. In principle, there are two straightforward approaches, but both have significant costs. Production systems such as Claude Code use a third approach that avoids the main drawbacks of both.\n\n**Approach One: Inject into System Prompt (system message)**. Append Skill content directly to the system prompt. The model's instruction-following ability is strongest for content in the system position (because training heavily uses instructions in this position), so Skill execution is most effective. The problem: each time a new Skill is loaded, the system message content changes, invalidating the KV Cache prefix. If the Agent frequently switches Skills (e.g., a task requires first using a search Skill, then a document Skill), the cache is repeatedly invalidated, significantly increasing latency and cost.\n\n**Approach Two: Read as a regular file, with content appearing in the middle of the context**. The Agent reads the Skill file via a generic file-reading tool, and the file content appears as a tool result in the conversation history—i.e., the middle of the context. This approach does not affect the KV Cache at all (the system prompt remains unchanged), but it places higher demands on the model's **instruction-following** ability: the model needs to accurately identify and follow the instructions within the Skill in the middle of a long context, rather than treating it as ordinary tool output to reference. In practice, different models vary significantly in their support for this mode—Claude performs most reliably because its training heavily uses instruction-following data in the middle position; other models often degrade when following instructions injected in the middle of the context.\n\n**Approach Three (Production Implementation): Metadata provided as dynamic context, full content loaded on demand via a dedicated tool**. Claude Code's core approach is to separate Skill \"routing\" from \"execution\": the model first receives the metadata for the available Skills and uses it to determine whether the current task requires a particular Skill; only after a Skill is selected does it load the complete `SKILL.md`. This design balances context overhead, Prompt Cache reuse, and instruction-following ability.\n\n- **Metadata list**—the `name` + `description` of all installed Skills (usually only a few hundred tokens)—is made available to the model in advance, allowing it to determine which Skills are relevant to the current task. Importantly, **the message role used to inject this metadata into the context is an implementation detail of the Claude Code Agent Harness, not a fixed requirement of the Agent Skills mechanism itself**. In some historical versions of Claude Code, this type of dynamic context appeared as user-role content wrapped in `<system-reminder>`; newer implementation paths that support mid-conversation system messages can instead use an appended system-role context block. Regardless of the representation, the common goal is to make the model aware of the currently available Skills without repeatedly rewriting the stable context prefix.\n\n- **Full content**—once the model determines from the metadata that a Skill is suitable for the current task, it reads the corresponding `SKILL.md` on demand through the Skill tool, and the content then enters the current execution context. This avoids loading the complete instructions for every Skill at the beginning of the session, reducing the amount of irrelevant context.\n\nIt is therefore important to distinguish two levels: **\"Skill metadata must be visible to the model in advance\" is a relatively stable mechanism, while \"user role, system role, or a wrapper such as `<system-reminder>`\" is a version-specific implementation choice.** `<system-reminder>` is not a protocol format exclusive to Agent Skills; it is one representation used by the Claude Code Agent Harness to inject dynamic system context.\n\nNote that **dynamically adding system context during a conversation is not unique to Skills**. In addition to metadata about available Skills, an Agent may need to keep the model informed of the current task state, runtime environment, or other dynamic information. The next section on the **Agent Status Bar** will explore this mechanism further, and the Skill metadata list can be viewed as one concrete example.\n\nThe following two figures show the effect of this design from two perspectives: the position of Skills in the trajectory and the evolution of the KV Cache.\n\n![Figure 2-12: Complete Structure of the Agent Trajectory After Enabling Skills](images/fig2-12.svg){height=55%}\n\n![Figure 2-13: Evolution of KV Cache as the Agent Trajectory Grows](images/fig2-13.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第10/17部分]\n- **浪费的词元**:大多数内容与当前任务无关。\n- **分散的注意力**:上下文中过多不相关的信息会分散模型对关键内容的注意力(本章后面的上下文压缩部分将在“上下文老化”概念下详细讨论这一点)。\n\n这是从静态提示工程到动态提示的自然演进:**不是一次性将所有知识加载到智能体中,而是允许按需加载知识**。智能体技能系统是这一理念的工程实现。\n\n### 技能:领域能力的可组合单元\n智能体技能的核心思想是将智能体的能力模块化,成为独立的、可加载的知识包[^ch2-3]。每个技能本质上是一组提示词和包含专门领域指导的文件,就像特定任务的操作手册。与将所有指令放在单个系统提示中的传统方法不同,技能使用逐步披露:首先向智能体展示目录摘要,然后仅在需要时加载完整内容。框架提供一个目录,而不是一次性将所有领域手册加载到上下文中,让智能体根据需要检索相关手册。\n\n[^ch2-3]Anthropic,“用智能体技能为现实世界装备智能体”,2025年。\n\n**第1层(元数据)**:每个技能必须包含一个`SKILL.md`文件,以YAML前置元数据(文件顶部由`---`界定的元数据块,类似于书籍的版权页)开头,包含`name`和`description`字段。智能体框架在启动时扫描所有已安装的技能,并将它们的`name`和`description`注入对话上下文。这通常只花费几百个词元,关于注入位置的权衡将在下一个小节讨论。目标是让智能体在不将所有技能内容加载到上下文中的情况下发现可用的专门能力。\n\n路由在很大程度上依赖于元数据的`description`字段。它应该足够简洁,以保持始终加载的词元数低,但应写成路由规则而不是功能摘要。最清晰的模式是“何时使用/何时不使用”,由**负面示例**支持,负面示例识别不应触发技能的情况。负面示例不是可选的;它们对于准确的技能路由至关重要。像“帮助后端”这样宽泛的描述会在不相关的任务上激活,而明确的排除能使路由更精确。出于路由目的,“何时使用我”比“我能做什么”重要得多。\n\n**第2层(核心工作流)**:当智能体确定任务需要特定技能时,它通过专用的技能工具加载完整的`SKILL.md`,内容作为工具结果出现在对话历史中。以PPTX技能[^ch2-4]为例,它包含处理PowerPoint文件的核心工作流:如何通过markitdown(微软的开源文档转Markdown工具)提取文本,如何解压缩PPTX文件以访问原始XML结构,以及关键文件的路径约定。\n\n[^ch2-4]Anthropic,“PPTX技能”,2025年。https://github.com/anthropics/skills/\n\n**第3层(细节)**:文件引用允许深入导航到更详细的子文档。主文件引用`html2pptx.md`(从HTML模板创建PowerPoint的详细工作流)、`reference.md`(格式技术细节)等。智能体根据特定需求有选择地读取相关子文档。\n\n技能不仅包含指导文档,还可以捆绑可执行代码工具和模板文件——将它们从纯知识传递转变为操作能力。\n\n技能的价值不仅在于上下文管理,还在于提供积累领域知识的可持续路径。每个技能是一个独立的知识模块,可以独立开发、测试、版本控制和共享。这种模块化将智能体能力扩展从集中式系统提示编辑转变为分布式技能生态系统,与Python的pip或Node.js的npm等包管理器在精神上相似。每个技能封装了特定领域的最佳实践。Anthropic的官方技能仓库已经涵盖文档处理(PPTX、PDF、DOCX)、数据分析、代码生成等领域,允许开发者使用、定制或创建全新的技能。\n\n这为智能体开发者揭示了一个重要原则:**在选择智能体交互模式时,要与模型和API设计支持的交互模式保持一致**。使用Claude构建智能体时,充分利用技能和结构化系统提示;使用其他模型时,遵循该模型供应商优化的约定。基础模型公司推广的智能体使用模式通常反映了这些模型训练和评估所支持的模式。\n\n### 技能实现方法及权衡\n定义技能后,下一个问题是具体的工程问题:技能内容应放置在上下文中的哪个位置?这个设计决策直接影响键值缓存(KV Cache)的效率和模型遵循技能指令的能力。原则上有两种直接方法,但都有显著成本。Claude Code等生产系统使用第三种方法,避免了两种方法的主要缺点。\n\n**方法一:注入系统提示(系统消息)**。将技能内容直接附加到系统提示中。模型在系统位置的内容的指令遵循能力最强(因为训练大量使用该位置的指令),所以技能执行最有效。问题在于:每次加载新技能时,系统消息内容改变,使KV Cache前缀失效。如果智能体频繁切换技能(例如,任务需要先使用搜索技能,然后使用文档技能),缓存会反复失效,显著增加时延和成本。\n\n**方法二:作为普通文件读取,内容出现在上下文中间**。智能体通过通用文件读取工具读取技能文件,文件内容作为工具结果出现在对话历史中——即上下文中间。这种方法完全不影响KV Cache(系统提示保持不变),但对模型的**指令遵循**能力提出了更高要求:模型需要准确识别并遵循上下文中间技能中的指令,而不是将其视为普通工具输出来引用。实际上,不同模型对这种模式的支持差异很大——Claude表现最可靠,因为其训练大量使用中间位置的指令遵循数据;其他模型在遵循上下文中间注入的指令时往往表现下降。\n\n**方法三(生产实现):元数据作为动态上下文,通过专用工具按需加载完整内容**。Claude Code的核心方法是将技能“路由”与“执行”分离:模型首先接收可用技能的元数据,并利用它确定当前任务是否需要特定技能;仅在选择技能后才加载完整的`SKILL.md`。这种设计平衡了上下文开销、提示缓存重用和指令遵循能力。\n\n- **元数据列表**——所有已安装技能的`name` + `description`(通常只有几百个词元)预先提供给模型,使其能够确定哪些技能与当前任务相关。重要的是,**将此元数据注入上下文的消息角色是Claude Code智能体框架的实现细节,而不是智能体技能机制本身的固定要求**。在Claude Code的一些历史版本中,这种动态上下文以包裹在`<system-reminder>`中的用户角色内容形式出现;支持会话中系统消息的较新实现路径可以改为使用附加的系统角色上下文块。无论表示形式如何,共同目标是让模型了解当前可用的技能,而无需反复重写稳定的上下文前缀。\n\n- **完整内容**——一旦模型从元数据中确定某个技能适合当前任务,它就通过技能工具按需读取相应的`SKILL.md`,内容随后进入当前执行上下文。这避免了在会话开始时加载所有技能的完整指令,减少了不相关上下文的数量。\n\n因此,区分两个层次很重要:**“技能元数据必须预先对模型可见”是相对稳定的机制,而“用户角色、系统角色或`<system-reminder>`等包装”是特定版本的实现选择**。`<system-reminder>`不是智能体技能专属的协议格式;它是Claude Code智能体框架注入动态系统上下文的一种表示形式。\n\n请注意,**在会话中动态添加系统上下文并非技能独有**。除了可用技能的元数据外,智能体可能需要让模型了解当前任务状态、运行时环境或其他动态信息。下一节关于**智能体状态栏**将进一步探讨这种机制,技能元数据列表可视为一个具体示例。\n\n以下两个图从两个角度展示了这种设计的效果:技能在轨迹中的位置和KV Cache的演进。\n\n![图2-12:启用技能后智能体轨迹的完整结构](images/fig2-12.svg){height=55%}\n\n![图2-13:智能体轨迹增长时KV Cache的演进](images/fig2-13.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第10/17部分]\n- **标记浪费**:大多数内容与当前任务无关。\n- **注意力稀释**:上下文中过多的无关信息稀释了模型对关键内容的注意力(本章后面的上下文压缩部分将在“上下文腐烂”概念下详细讨论这一点)。\n\n这是从静态提示工程到动态提示的自然演进:**不是一次性将所有知识加载到代理中,而是允许它按需加载知识**。代理技能系统是这一想法的工程实现。\n\n### 技能:领域能力的可组合单元\n代理技能的核心思想是将代理的能力模块化,形成独立的、可加载的知识包[^ch2-3]。每个技能本质上是一组包含专业领域指导的提示和文件,就像特定任务的操作手册。与将所有指令放入单个系统提示的传统方法不同,技能使用渐进披露:首先向代理展示目录摘要,然后仅在需要时加载完整内容。框架提供一个目录,让代理根据需要检索相关手册,而不是一次性将所有领域手册加载到上下文中。\n\n[^ch2-3]: Anthropic,“用代理技能为现实世界装备代理”,2025年。\n\n**第1层(元数据)**:每个技能必须包含一个`SKILL.md`文件,以YAML前matter开头(文件顶部由`---`分隔的元数据块,类似于书籍的版权页),包含`name`和`description`字段。代理框架在启动时扫描所有已安装的技能,并将它们的`name`和`description`注入对话上下文。这通常只花费几百个标记,下一小节将讨论注入位置的权衡。目标是让代理在不将所有技能内容加载到上下文中的情况下,发现可用的专业能力。\n\n路由在很大程度上依赖于元数据的`description`字段。它应该足够简洁,以保持始终加载的标记数低,但应写成路由规则而不是功能摘要。最清晰的模式是“何时使用/何时不使用”,由**否定示例**支持,这些示例识别不应触发技能的情况。否定示例不是可选的;它们对于准确的技能路由至关重要。像“帮助处理后端”这样宽泛的描述会在不相关的任务上激活,而明确的排除使路由大大更精确。出于路由目的,“何时使用我”比“我能做什么”重要得多。\n\n**第2层(核心工作流)**:当代理确定任务需要特定技能时,它通过专用的技能工具加载完整的`SKILL.md`,内容作为工具结果出现在对话历史中。以PPTX技能[^ch2-4]为例,它包含处理PowerPoint文件的核心工作流:如何通过markitdown(微软的开源文档转Markdown工具)提取文本,如何解压缩PPTX文件以访问原始XML结构,以及关键文件的路径约定。\n\n[^ch2-4]: Anthropic,“PPTX技能”,2025年。https://github.com/anthropics/skills/\n\n**第3层(细节)**:文件引用允许更深入地导航到更详细的子文档。主文件引用`html2pptx.md`(从HTML模板创建PowerPoint的详细工作流)、`reference.md`(格式技术细节)等。代理根据特定需求有选择地读取相关子文档。\n\n技能不仅包含说明性文档,还可以捆绑可执行代码工具和模板文件——将它们从纯知识传递转变为操作能力。\n\n技能的价值不仅在于上下文管理,还在于提供积累领域知识的可持续路径。每个技能是一个自包含的知识模块,可以独立开发、测试、版本控制和共享。这种模块化将代理能力扩展从集中式系统提示编辑转变为分布式技能生态系统,与Python的pip或Node.js的npm等包管理器精神相似。每个技能封装了特定领域的最佳实践。Anthropic的官方技能库已经涵盖文档处理(PPTX、PDF、DOCX)、数据分析、代码生成等领域,允许开发者使用、定制或创建全新的技能。\n\n这揭示了代理开发者的一个重要原则:**在选择代理交互模式时,与模型和API设计支持的交互模式保持一致**。使用Claude构建代理时,充分利用技能和结构化系统提示;使用其他模型时,遵循该模型供应商优化的约定。基础模型公司推广的代理使用模式通常反映了那些模型经过训练和评估支持的模式。\n\n### 技能实现方法及权衡\n定义技能后,下一个问题是具体的工程问题:技能内容应放置在上下文中的哪个位置?这个设计决策直接影响KV缓存效率和模型遵循技能指令的能力。原则上有两种直接方法,但都有显著成本。Claude Code等生产系统使用第三种方法,避免了两种方法的主要缺点。\n\n**方法一:注入系统提示(系统消息)**。将技能内容直接附加到系统提示中。模型在系统位置的内容的指令遵循能力最强(因为训练大量使用该位置的指令),因此技能执行最有效。问题:每次加载新技能时,系统消息内容更改,使KV缓存前缀失效。如果代理频繁切换技能(例如,任务需要首先使用搜索技能,然后使用文档技能),缓存会反复失效,显著增加延迟和成本。\n\n**方法二:作为普通文件读取,内容出现在上下文中间**。代理通过通用文件读取工具读取技能文件,文件内容作为工具结果出现在对话历史中——即上下文中间。这种方法完全不影响KV缓存(系统提示保持不变),但对模型的**指令遵循**能力提出了更高要求:模型需要在长上下文中准确识别并遵循技能中的指令,而不是将其视为普通工具输出来参考。实际上,不同模型对此模式的支持差异很大——Claude执行最可靠,因为其训练大量使用中间位置的指令遵循数据;其他模型在遵循注入上下文中间的指令时往往退化。\n\n**方法三(生产实现):元数据作为动态上下文,通过专用工具按需加载完整内容**。Claude Code的核心方法是将技能“路由”与“执行”分离:模型首先接收可用技能的元数据,并使用它来确定当前任务是否需要特定技能;仅在选择技能后才加载完整的`SKILL.md`。这种设计平衡了上下文开销、提示缓存重用和指令遵循能力。\n\n- **元数据列表**——所有已安装技能的`name`+`description`(通常只有几百个标记)——预先提供给模型,使其能够确定当前任务相关的技能。重要的是,**用于将此元数据注入上下文的消息角色是Claude Code代理框架的实现细节,而不是代理技能机制本身的固定要求**。在Claude Code的某些历史版本中,这种类型的动态上下文以包裹在`<system-reminder>`中的用户角色内容形式出现;支持会话中间系统消息的较新实现路径可以改为使用附加的系统角色上下文块。无论表示如何,共同目标是让模型在不重复重写稳定上下文前缀的情况下,了解当前可用的技能。\n\n- **完整内容**——一旦模型从元数据中确定技能适合当前任务,它通过技能工具按需读取相应的`SKILL.md`,内容随后进入当前执行上下文。这避免了在会话开始时加载每个技能的完整指令,减少了不相关上下文的数量。\n\n因此,区分两个层次很重要:**“技能元数据必须提前对模型可见”是相对稳定的机制,而“用户角色、系统角色或`<system-reminder>`等包装器”是特定版本的实现选择**。`<system-reminder>`不是代理技能独有的协议格式;它是Claude Code代理框架注入动态系统上下文的一种表示。\n\n注意,**在会话期间动态添加系统上下文并非技能独有**。除了可用技能的元数据,代理可能需要让模型了解当前任务状态、运行时环境或其他动态信息。下一节关于**代理状态栏**将进一步探讨该机制,技能元数据列表可视为一个具体示例。\n\n以下两个图从两个角度展示了该设计的效果:技能在轨迹中的位置和KV缓存的演进。\n\n![图2-12:启用技能后代理轨迹的完整结构](images/fig2-12.svg){height=55%}\n\n![图2-13:代理轨迹增长时KV缓存的演进](images/fig2-13.svg)"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853908586023ea4fa82759f0cae4f2ffa723d4f9dd6fe2a16",
"model": "doubao-seed-1-6-250615",
"created": 1785390877,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"准确翻译了所有关键概念,如“context rot”译为“上下文老化”(符合技术文档规范),“Progressive Disclosure”译为“逐步披露”,未出现遗漏或误译。例如,“Negative examples are not optional; they are essential to accurate Skill routing.”译为“负面示例不是可选的;它们对于准确的技能路由至关重要。”完全忠实原文。\"},\"fluency\":{\"score\":5,\"evidence\":\"语言流畅自然,符合中文技术文档表达习惯。例如,“the model's instruction-following ability is strongest for content in the system position”译为“模型在系统位置的内容的指令遵循能力最强”,语句通顺,无语法错误。”},\" ,\"terminology\":{\"score\":5,\"evidence\":\"术语一致且准确,如“Agent”统一译为“智能体”,“KV Cache”译为“键值缓存(KV Cache)”并保持一致,“token”译为“词元”,“instruction-following”译为“指令遵循”,符合技术翻译规范。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"所有Markdown元素保留完整,包括标题层级(###)、列表项(-)、引用块(**...**)、图片链接(![Figure...](images/...))、脚注([^ch2-3])及代码格式(`SKILL.md`),未出现格式丢失或错乱。\"},\"Y\":{\"accuracy\":{\"score\":4,\"evidence\":\"存在少量术语误译,如“context rot”译为“上下文腐烂”(非标准技术术语,应为“上下文老化”或“上下文衰退”),“token”译为“标记”(技术文档中更常用“词元”)。“Negative examples”译为“否定示例”(应为“负面示例”更准确)。\"},\"fluency\":{\"score\":4,\"evidence\":\"整体流畅,但部分表达略显生硬。例如,“the model needs to accurately identify and follow the instructions within the Skill in the middle of a long context”译为“模型需要在长上下文中准确识别并遵循技能中的指令”,“长上下文”表述稍显冗余,可优化为“上下文中间”。\"},\"terminology\":{\"score\":4,\"evidence\":\"部分术语不一致或不规范,如“Agent”译为“代理”(技术文档中“智能体”更通用),“KV Cache”译为“KV缓存”(缺少括号标注原文),“instruction-following”译为“指令遵循”(正确),但“token”译为“标记”(不标准)。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown格式保留完整,标题、列表、图片、脚注、代码格式等均未丢失,与原文结构一致。\"}},\"preferred\":\"X\",\"preference_evidence\":\"X在准确性(无关键术语误译)、术语规范性(“词元”“负面示例”“智能体”等标准译法)和流畅度(更自然的中文表达)上优于Y,且Markdown格式保真度相同。Y存在“上下文腐烂”“标记”等不规范术语,影响技术文档的专业性。\"} } } }",
"usage": {
"prompt_tokens": 5937,
"completion_tokens": 688,
"total_tokens": 6625
}
},
"latency_milliseconds": 19352,
"validation": {
"valid": false,
"error_type": "ValueError",
"error": "judge variants must contain exactly X and Y"
}
},
{
"attempt": 2,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n- **Wasted tokens**: Most content is irrelevant to the current task.\n- **Diluted attention**: Too much irrelevant information in the context dilutes the model's attention to key content (the context compression section later in this chapter discusses this in detail under the concept of \"context rot\").\n\nThis is the natural evolution from static prompt engineering to dynamic prompts: **instead of loading all knowledge into the Agent at once, allow it to load knowledge on demand**. The Agent Skills system is the engineering implementation of this idea.\n\n### Skills: Composable Units of Domain Capability\n\nThe core idea of Agent Skills is to modularize the Agent's capabilities into independent, loadable knowledge packages[^ch2-3]. Each Skill is essentially a collection of prompts and files containing specialized domain guidance, like an operating manual for a specific task. Unlike the traditional approach of placing all instructions into a single system prompt, Skills use Progressive Disclosure: first show the Agent a table-of-contents summary, then load the full content only when needed. Instead of loading every domain manual into context at once, the framework provides a directory and lets the Agent retrieve the relevant manual as needed.\n\n[^ch2-3]: Anthropic, \"Equipping Agents for the Real World with Agent Skills\", 2025.\n\n**Layer 1 (Metadata)**: Each Skill must include a `SKILL.md` file that starts with YAML frontmatter (a metadata block at the top of the file delimited by `---`, similar to a book's copyright page), containing `name` and `description` fields. The Agent framework scans all installed Skills at startup and injects their `name` and `description` into the dialogue context. This usually costs only a few hundred tokens, and the trade-offs around injection location are discussed in the next subsection. The goal is to let the Agent discover which specialized capabilities are available without loading all Skill content into context.\n\nRouting depends heavily on the metadata's `description` field. It should be concise enough to keep the always-loaded token count low, but written as a routing rule rather than a feature summary. The clearest pattern is \"Use when / Do not use when,\" supported by **negative examples** that identify situations in which the Skill should not be triggered. Negative examples are not optional; they are essential to accurate Skill routing. Broad descriptions such as \"help with backend\" activate on unrelated tasks, while explicit exclusions make routing substantially more precise. For routing purposes, \"when to use me\" matters far more than \"what I can do.\"\n\n**Layer 2 (Core Workflow)**: When the Agent determines that a specific Skill is needed for a task, it loads the complete `SKILL.md` via a dedicated Skill tool, and the content appears in the conversation history as a tool result. Using the PPTX Skill[^ch2-4] as an example, it contains the core workflow for handling PowerPoint files: how to extract text via markitdown (Microsoft's open-source document-to-Markdown tool), how to unzip the PPTX file to access the raw XML structure, and the path conventions for key files.\n\n[^ch2-4]: Anthropic, \"PPTX Skill\", 2025. https://github.com/anthropics/skills/\n\n**Layer 3 (Details)**: File references allow deeper navigation into more detailed sub-documents. The main file references `html2pptx.md` (detailed workflow for creating PowerPoint from HTML templates), `reference.md` (format technical details), and others. The Agent selectively reads relevant sub-documents based on specific needs.\n\nSkills not only contain instructional documentation but can also bundle executable code tools and template files—turning them from pure knowledge transfer into operational capabilities.\n\nThe value of Skills lies not only in context management but also in providing a sustainable path for accumulating domain knowledge. Each Skill is a self-contained knowledge module that can be independently developed, tested, version-controlled, and shared. This modularity transforms Agent capability expansion from centralized system prompt editing into a distributed Skill ecosystem, similar in spirit to package managers such as Python's pip or Node.js's npm. Each Skill encapsulates best practices for a specific domain. Anthropic's official Skills repository already covers document processing (PPTX, PDF, DOCX), data analysis, code generation, and other domains, allowing developers to use, customize, or create entirely new Skills.\n\nThis reveals an important principle for Agent developers: **when choosing an Agent interaction mode, align with the interaction patterns the model and API are designed to support**. When building Agents with Claude, fully leverage Skills and structured system prompts; when using other models, follow the conventions optimized by that model vendor. The Agent usage patterns promoted by foundation model companies often reflect the modes those models are trained and evaluated to support.\n\n### Skills Implementation Methods and Trade-offs\n\nAfter defining Skills, the next question is a concrete engineering problem: where in the context should Skill content be placed? This design decision directly affects KV Cache efficiency and the model's ability to follow the Skill's instructions. In principle, there are two straightforward approaches, but both have significant costs. Production systems such as Claude Code use a third approach that avoids the main drawbacks of both.\n\n**Approach One: Inject into System Prompt (system message)**. Append Skill content directly to the system prompt. The model's instruction-following ability is strongest for content in the system position (because training heavily uses instructions in this position), so Skill execution is most effective. The problem: each time a new Skill is loaded, the system message content changes, invalidating the KV Cache prefix. If the Agent frequently switches Skills (e.g., a task requires first using a search Skill, then a document Skill), the cache is repeatedly invalidated, significantly increasing latency and cost.\n\n**Approach Two: Read as a regular file, with content appearing in the middle of the context**. The Agent reads the Skill file via a generic file-reading tool, and the file content appears as a tool result in the conversation history—i.e., the middle of the context. This approach does not affect the KV Cache at all (the system prompt remains unchanged), but it places higher demands on the model's **instruction-following** ability: the model needs to accurately identify and follow the instructions within the Skill in the middle of a long context, rather than treating it as ordinary tool output to reference. In practice, different models vary significantly in their support for this mode—Claude performs most reliably because its training heavily uses instruction-following data in the middle position; other models often degrade when following instructions injected in the middle of the context.\n\n**Approach Three (Production Implementation): Metadata provided as dynamic context, full content loaded on demand via a dedicated tool**. Claude Code's core approach is to separate Skill \"routing\" from \"execution\": the model first receives the metadata for the available Skills and uses it to determine whether the current task requires a particular Skill; only after a Skill is selected does it load the complete `SKILL.md`. This design balances context overhead, Prompt Cache reuse, and instruction-following ability.\n\n- **Metadata list**—the `name` + `description` of all installed Skills (usually only a few hundred tokens)—is made available to the model in advance, allowing it to determine which Skills are relevant to the current task. Importantly, **the message role used to inject this metadata into the context is an implementation detail of the Claude Code Agent Harness, not a fixed requirement of the Agent Skills mechanism itself**. In some historical versions of Claude Code, this type of dynamic context appeared as user-role content wrapped in `<system-reminder>`; newer implementation paths that support mid-conversation system messages can instead use an appended system-role context block. Regardless of the representation, the common goal is to make the model aware of the currently available Skills without repeatedly rewriting the stable context prefix.\n\n- **Full content**—once the model determines from the metadata that a Skill is suitable for the current task, it reads the corresponding `SKILL.md` on demand through the Skill tool, and the content then enters the current execution context. This avoids loading the complete instructions for every Skill at the beginning of the session, reducing the amount of irrelevant context.\n\nIt is therefore important to distinguish two levels: **\"Skill metadata must be visible to the model in advance\" is a relatively stable mechanism, while \"user role, system role, or a wrapper such as `<system-reminder>`\" is a version-specific implementation choice.** `<system-reminder>` is not a protocol format exclusive to Agent Skills; it is one representation used by the Claude Code Agent Harness to inject dynamic system context.\n\nNote that **dynamically adding system context during a conversation is not unique to Skills**. In addition to metadata about available Skills, an Agent may need to keep the model informed of the current task state, runtime environment, or other dynamic information. The next section on the **Agent Status Bar** will explore this mechanism further, and the Skill metadata list can be viewed as one concrete example.\n\nThe following two figures show the effect of this design from two perspectives: the position of Skills in the trajectory and the evolution of the KV Cache.\n\n![Figure 2-12: Complete Structure of the Agent Trajectory After Enabling Skills](images/fig2-12.svg){height=55%}\n\n![Figure 2-13: Evolution of KV Cache as the Agent Trajectory Grows](images/fig2-13.svg)\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第10/17部分]\n- **浪费的词元**:大多数内容与当前任务无关。\n- **分散的注意力**:上下文中过多不相关的信息会分散模型对关键内容的注意力(本章后面的上下文压缩部分将在“上下文老化”概念下详细讨论这一点)。\n\n这是从静态提示工程到动态提示的自然演进:**不是一次性将所有知识加载到智能体中,而是允许按需加载知识**。智能体技能系统是这一理念的工程实现。\n\n### 技能:领域能力的可组合单元\n智能体技能的核心思想是将智能体的能力模块化,成为独立的、可加载的知识包[^ch2-3]。每个技能本质上是一组提示词和包含专门领域指导的文件,就像特定任务的操作手册。与将所有指令放在单个系统提示中的传统方法不同,技能使用逐步披露:首先向智能体展示目录摘要,然后仅在需要时加载完整内容。框架提供一个目录,而不是一次性将所有领域手册加载到上下文中,让智能体根据需要检索相关手册。\n\n[^ch2-3]Anthropic,“用智能体技能为现实世界装备智能体”,2025年。\n\n**第1层(元数据)**:每个技能必须包含一个`SKILL.md`文件,以YAML前置元数据(文件顶部由`---`界定的元数据块,类似于书籍的版权页)开头,包含`name`和`description`字段。智能体框架在启动时扫描所有已安装的技能,并将它们的`name`和`description`注入对话上下文。这通常只花费几百个词元,关于注入位置的权衡将在下一个小节讨论。目标是让智能体在不将所有技能内容加载到上下文中的情况下发现可用的专门能力。\n\n路由在很大程度上依赖于元数据的`description`字段。它应该足够简洁,以保持始终加载的词元数低,但应写成路由规则而不是功能摘要。最清晰的模式是“何时使用/何时不使用”,由**负面示例**支持,负面示例识别不应触发技能的情况。负面示例不是可选的;它们对于准确的技能路由至关重要。像“帮助后端”这样宽泛的描述会在不相关的任务上激活,而明确的排除能使路由更精确。出于路由目的,“何时使用我”比“我能做什么”重要得多。\n\n**第2层(核心工作流)**:当智能体确定任务需要特定技能时,它通过专用的技能工具加载完整的`SKILL.md`,内容作为工具结果出现在对话历史中。以PPTX技能[^ch2-4]为例,它包含处理PowerPoint文件的核心工作流:如何通过markitdown(微软的开源文档转Markdown工具)提取文本,如何解压缩PPTX文件以访问原始XML结构,以及关键文件的路径约定。\n\n[^ch2-4]Anthropic,“PPTX技能”,2025年。https://github.com/anthropics/skills/\n\n**第3层(细节)**:文件引用允许深入导航到更详细的子文档。主文件引用`html2pptx.md`(从HTML模板创建PowerPoint的详细工作流)、`reference.md`(格式技术细节)等。智能体根据特定需求有选择地读取相关子文档。\n\n技能不仅包含指导文档,还可以捆绑可执行代码工具和模板文件——将它们从纯知识传递转变为操作能力。\n\n技能的价值不仅在于上下文管理,还在于提供积累领域知识的可持续路径。每个技能是一个独立的知识模块,可以独立开发、测试、版本控制和共享。这种模块化将智能体能力扩展从集中式系统提示编辑转变为分布式技能生态系统,与Python的pip或Node.js的npm等包管理器在精神上相似。每个技能封装了特定领域的最佳实践。Anthropic的官方技能仓库已经涵盖文档处理(PPTX、PDF、DOCX)、数据分析、代码生成等领域,允许开发者使用、定制或创建全新的技能。\n\n这为智能体开发者揭示了一个重要原则:**在选择智能体交互模式时,要与模型和API设计支持的交互模式保持一致**。使用Claude构建智能体时,充分利用技能和结构化系统提示;使用其他模型时,遵循该模型供应商优化的约定。基础模型公司推广的智能体使用模式通常反映了这些模型训练和评估所支持的模式。\n\n### 技能实现方法及权衡\n定义技能后,下一个问题是具体的工程问题:技能内容应放置在上下文中的哪个位置?这个设计决策直接影响键值缓存(KV Cache)的效率和模型遵循技能指令的能力。原则上有两种直接方法,但都有显著成本。Claude Code等生产系统使用第三种方法,避免了两种方法的主要缺点。\n\n**方法一:注入系统提示(系统消息)**。将技能内容直接附加到系统提示中。模型在系统位置的内容的指令遵循能力最强(因为训练大量使用该位置的指令),所以技能执行最有效。问题在于:每次加载新技能时,系统消息内容改变,使KV Cache前缀失效。如果智能体频繁切换技能(例如,任务需要先使用搜索技能,然后使用文档技能),缓存会反复失效,显著增加时延和成本。\n\n**方法二:作为普通文件读取,内容出现在上下文中间**。智能体通过通用文件读取工具读取技能文件,文件内容作为工具结果出现在对话历史中——即上下文中间。这种方法完全不影响KV Cache(系统提示保持不变),但对模型的**指令遵循**能力提出了更高要求:模型需要准确识别并遵循上下文中间技能中的指令,而不是将其视为普通工具输出来引用。实际上,不同模型对这种模式的支持差异很大——Claude表现最可靠,因为其训练大量使用中间位置的指令遵循数据;其他模型在遵循上下文中间注入的指令时往往表现下降。\n\n**方法三(生产实现):元数据作为动态上下文,通过专用工具按需加载完整内容**。Claude Code的核心方法是将技能“路由”与“执行”分离:模型首先接收可用技能的元数据,并利用它确定当前任务是否需要特定技能;仅在选择技能后才加载完整的`SKILL.md`。这种设计平衡了上下文开销、提示缓存重用和指令遵循能力。\n\n- **元数据列表**——所有已安装技能的`name` + `description`(通常只有几百个词元)预先提供给模型,使其能够确定哪些技能与当前任务相关。重要的是,**将此元数据注入上下文的消息角色是Claude Code智能体框架的实现细节,而不是智能体技能机制本身的固定要求**。在Claude Code的一些历史版本中,这种动态上下文以包裹在`<system-reminder>`中的用户角色内容形式出现;支持会话中系统消息的较新实现路径可以改为使用附加的系统角色上下文块。无论表示形式如何,共同目标是让模型了解当前可用的技能,而无需反复重写稳定的上下文前缀。\n\n- **完整内容**——一旦模型从元数据中确定某个技能适合当前任务,它就通过技能工具按需读取相应的`SKILL.md`,内容随后进入当前执行上下文。这避免了在会话开始时加载所有技能的完整指令,减少了不相关上下文的数量。\n\n因此,区分两个层次很重要:**“技能元数据必须预先对模型可见”是相对稳定的机制,而“用户角色、系统角色或`<system-reminder>`等包装”是特定版本的实现选择**。`<system-reminder>`不是智能体技能专属的协议格式;它是Claude Code智能体框架注入动态系统上下文的一种表示形式。\n\n请注意,**在会话中动态添加系统上下文并非技能独有**。除了可用技能的元数据外,智能体可能需要让模型了解当前任务状态、运行时环境或其他动态信息。下一节关于**智能体状态栏**将进一步探讨这种机制,技能元数据列表可视为一个具体示例。\n\n以下两个图从两个角度展示了这种设计的效果:技能在轨迹中的位置和KV Cache的演进。\n\n![图2-12:启用技能后智能体轨迹的完整结构](images/fig2-12.svg){height=55%}\n\n![图2-13:智能体轨迹增长时KV Cache的演进](images/fig2-13.svg)\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第10/17部分]\n- **标记浪费**:大多数内容与当前任务无关。\n- **注意力稀释**:上下文中过多的无关信息稀释了模型对关键内容的注意力(本章后面的上下文压缩部分将在“上下文腐烂”概念下详细讨论这一点)。\n\n这是从静态提示工程到动态提示的自然演进:**不是一次性将所有知识加载到代理中,而是允许它按需加载知识**。代理技能系统是这一想法的工程实现。\n\n### 技能:领域能力的可组合单元\n代理技能的核心思想是将代理的能力模块化,形成独立的、可加载的知识包[^ch2-3]。每个技能本质上是一组包含专业领域指导的提示和文件,就像特定任务的操作手册。与将所有指令放入单个系统提示的传统方法不同,技能使用渐进披露:首先向代理展示目录摘要,然后仅在需要时加载完整内容。框架提供一个目录,让代理根据需要检索相关手册,而不是一次性将所有领域手册加载到上下文中。\n\n[^ch2-3]: Anthropic,“用代理技能为现实世界装备代理”,2025年。\n\n**第1层(元数据)**:每个技能必须包含一个`SKILL.md`文件,以YAML前matter开头(文件顶部由`---`分隔的元数据块,类似于书籍的版权页),包含`name`和`description`字段。代理框架在启动时扫描所有已安装的技能,并将它们的`name`和`description`注入对话上下文。这通常只花费几百个标记,下一小节将讨论注入位置的权衡。目标是让代理在不将所有技能内容加载到上下文中的情况下,发现可用的专业能力。\n\n路由在很大程度上依赖于元数据的`description`字段。它应该足够简洁,以保持始终加载的标记数低,但应写成路由规则而不是功能摘要。最清晰的模式是“何时使用/何时不使用”,由**否定示例**支持,这些示例识别不应触发技能的情况。否定示例不是可选的;它们对于准确的技能路由至关重要。像“帮助处理后端”这样宽泛的描述会在不相关的任务上激活,而明确的排除使路由大大更精确。出于路由目的,“何时使用我”比“我能做什么”重要得多。\n\n**第2层(核心工作流)**:当代理确定任务需要特定技能时,它通过专用的技能工具加载完整的`SKILL.md`,内容作为工具结果出现在对话历史中。以PPTX技能[^ch2-4]为例,它包含处理PowerPoint文件的核心工作流:如何通过markitdown(微软的开源文档转Markdown工具)提取文本,如何解压缩PPTX文件以访问原始XML结构,以及关键文件的路径约定。\n\n[^ch2-4]: Anthropic,“PPTX技能”,2025年。https://github.com/anthropics/skills/\n\n**第3层(细节)**:文件引用允许更深入地导航到更详细的子文档。主文件引用`html2pptx.md`(从HTML模板创建PowerPoint的详细工作流)、`reference.md`(格式技术细节)等。代理根据特定需求有选择地读取相关子文档。\n\n技能不仅包含说明性文档,还可以捆绑可执行代码工具和模板文件——将它们从纯知识传递转变为操作能力。\n\n技能的价值不仅在于上下文管理,还在于提供积累领域知识的可持续路径。每个技能是一个自包含的知识模块,可以独立开发、测试、版本控制和共享。这种模块化将代理能力扩展从集中式系统提示编辑转变为分布式技能生态系统,与Python的pip或Node.js的npm等包管理器精神相似。每个技能封装了特定领域的最佳实践。Anthropic的官方技能库已经涵盖文档处理(PPTX、PDF、DOCX)、数据分析、代码生成等领域,允许开发者使用、定制或创建全新的技能。\n\n这揭示了代理开发者的一个重要原则:**在选择代理交互模式时,与模型和API设计支持的交互模式保持一致**。使用Claude构建代理时,充分利用技能和结构化系统提示;使用其他模型时,遵循该模型供应商优化的约定。基础模型公司推广的代理使用模式通常反映了那些模型经过训练和评估支持的模式。\n\n### 技能实现方法及权衡\n定义技能后,下一个问题是具体的工程问题:技能内容应放置在上下文中的哪个位置?这个设计决策直接影响KV缓存效率和模型遵循技能指令的能力。原则上有两种直接方法,但都有显著成本。Claude Code等生产系统使用第三种方法,避免了两种方法的主要缺点。\n\n**方法一:注入系统提示(系统消息)**。将技能内容直接附加到系统提示中。模型在系统位置的内容的指令遵循能力最强(因为训练大量使用该位置的指令),因此技能执行最有效。问题:每次加载新技能时,系统消息内容更改,使KV缓存前缀失效。如果代理频繁切换技能(例如,任务需要首先使用搜索技能,然后使用文档技能),缓存会反复失效,显著增加延迟和成本。\n\n**方法二:作为普通文件读取,内容出现在上下文中间**。代理通过通用文件读取工具读取技能文件,文件内容作为工具结果出现在对话历史中——即上下文中间。这种方法完全不影响KV缓存(系统提示保持不变),但对模型的**指令遵循**能力提出了更高要求:模型需要在长上下文中准确识别并遵循技能中的指令,而不是将其视为普通工具输出来参考。实际上,不同模型对此模式的支持差异很大——Claude执行最可靠,因为其训练大量使用中间位置的指令遵循数据;其他模型在遵循注入上下文中间的指令时往往退化。\n\n**方法三(生产实现):元数据作为动态上下文,通过专用工具按需加载完整内容**。Claude Code的核心方法是将技能“路由”与“执行”分离:模型首先接收可用技能的元数据,并使用它来确定当前任务是否需要特定技能;仅在选择技能后才加载完整的`SKILL.md`。这种设计平衡了上下文开销、提示缓存重用和指令遵循能力。\n\n- **元数据列表**——所有已安装技能的`name`+`description`(通常只有几百个标记)——预先提供给模型,使其能够确定当前任务相关的技能。重要的是,**用于将此元数据注入上下文的消息角色是Claude Code代理框架的实现细节,而不是代理技能机制本身的固定要求**。在Claude Code的某些历史版本中,这种类型的动态上下文以包裹在`<system-reminder>`中的用户角色内容形式出现;支持会话中间系统消息的较新实现路径可以改为使用附加的系统角色上下文块。无论表示如何,共同目标是让模型在不重复重写稳定上下文前缀的情况下,了解当前可用的技能。\n\n- **完整内容**——一旦模型从元数据中确定技能适合当前任务,它通过技能工具按需读取相应的`SKILL.md`,内容随后进入当前执行上下文。这避免了在会话开始时加载每个技能的完整指令,减少了不相关上下文的数量。\n\n因此,区分两个层次很重要:**“技能元数据必须提前对模型可见”是相对稳定的机制,而“用户角色、系统角色或`<system-reminder>`等包装器”是特定版本的实现选择**。`<system-reminder>`不是代理技能独有的协议格式;它是Claude Code代理框架注入动态系统上下文的一种表示。\n\n注意,**在会话期间动态添加系统上下文并非技能独有**。除了可用技能的元数据,代理可能需要让模型了解当前任务状态、运行时环境或其他动态信息。下一节关于**代理状态栏**将进一步探讨该机制,技能元数据列表可视为一个具体示例。\n\n以下两个图从两个角度展示了该设计的效果:技能在轨迹中的位置和KV缓存的演进。\n\n![图2-12:启用技能后代理轨迹的完整结构](images/fig2-12.svg){height=55%}\n\n![图2-13:代理轨迹增长时KV缓存的演进](images/fig2-13.svg)"
},
{
"role": "assistant",
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"准确翻译了所有关键概念,如“context rot”译为“上下文老化”(符合技术文档规范),“Progressive Disclosure”译为“逐步披露”,未出现遗漏或误译。例如,“Negative examples are not optional; they are essential to accurate Skill routing.”译为“负面示例不是可选的;它们对于准确的技能路由至关重要。”完全忠实原文。\"},\"fluency\":{\"score\":5,\"evidence\":\"语言流畅自然,符合中文技术文档表达习惯。例如,“the model's instruction-following ability is strongest for content in the system position”译为“模型在系统位置的内容的指令遵循能力最强”,语句通顺,无语法错误。”},\" ,\"terminology\":{\"score\":5,\"evidence\":\"术语一致且准确,如“Agent”统一译为“智能体”,“KV Cache”译为“键值缓存(KV Cache)”并保持一致,“token”译为“词元”,“instruction-following”译为“指令遵循”,符合技术翻译规范。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"所有Markdown元素保留完整,包括标题层级(###)、列表项(-)、引用块(**...**)、图片链接(![Figure...](images/...))、脚注([^ch2-3])及代码格式(`SKILL.md`),未出现格式丢失或错乱。\"},\"Y\":{\"accuracy\":{\"score\":4,\"evidence\":\"存在少量术语误译,如“context rot”译为“上下文腐烂”(非标准技术术语,应为“上下文老化”或“上下文衰退”),“token”译为“标记”(技术文档中更常用“词元”)。“Negative examples”译为“否定示例”(应为“负面示例”更准确)。\"},\"fluency\":{\"score\":4,\"evidence\":\"整体流畅,但部分表达略显生硬。例如,“the model needs to accurately identify and follow the instructions within the Skill in the middle of a long context”译为“模型需要在长上下文中准确识别并遵循技能中的指令”,“长上下文”表述稍显冗余,可优化为“上下文中间”。\"},\"terminology\":{\"score\":4,\"evidence\":\"部分术语不一致或不规范,如“Agent”译为“代理”(技术文档中“智能体”更通用),“KV Cache”译为“KV缓存”(缺少括号标注原文),“instruction-following”译为“指令遵循”(正确),但“token”译为“标记”(不标准)。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown格式保留完整,标题、列表、图片、脚注、代码格式等均未丢失,与原文结构一致。\"}},\"preferred\":\"X\",\"preference_evidence\":\"X在准确性(无关键术语误译)、术语规范性(“词元”“负面示例”“智能体”等标准译法)和流畅度(更自然的中文表达)上优于Y,且Markdown格式保真度相同。Y存在“上下文腐烂”“标记”等不规范术语,影响技术文档的专业性。\"} } } }"
},
{
"role": "user",
"content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys."
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853908779553ea4fa82759f0cae4f2ffa723d4f9dd63080d0",
"model": "doubao-seed-1-6-250615",
"created": 1785390890,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"准确翻译了所有关键概念,如“context rot”译为“上下文老化”(符合技术文档规范),“Progressive Disclosure”译为“逐步披露”,未出现遗漏或误译。例如,“Negative examples are not optional; they are essential to accurate Skill routing.”译为“负面示例不是可选的;它们对于准确的技能路由至关重要。”完全忠实原文。\"},\"fluency\":{\"score\":5,\"evidence\":\"语言流畅自然,符合中文技术文档表达习惯。例如,“the model's instruction-following ability is strongest for content in the system position”译为“模型在系统位置的内容的指令遵循能力最强”,语句通顺,无语法错误。\"},\"terminology\":{\"score\":5,\"evidence\":\"术语一致且准确,如“Agent”统一译为“智能体”,“KV Cache”译为“键值缓存(KV Cache)”并保持一致,“token”译为“词元”,“instruction-following”译为“指令遵循”,符合技术翻译规范。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"所有Markdown元素保留完整,包括标题层级(###)、列表项(-)、引用块(**...**)、图片链接(![Figure...](images/...))、脚注([^ch2-3])及代码格式(`SKILL.md`),未出现格式丢失或错乱。\"}},\"Y\":{\"accuracy\":{\"score\":4,\"evidence\":\"存在少量术语误译,如“context rot”译为“上下文腐烂”(非标准技术术语,应为“上下文老化”或“上下文衰退”),“token”译为“标记”(技术文档中更常用“词元”)。“Negative examples”译为“否定示例”(应为“负面示例”更准确)。\"},\"fluency\":{\"score\":4,\"evidence\":\"整体流畅,但部分表达略显生硬。例如,“the model needs to accurately identify and follow the instructions within the Skill in the middle of a long context”译为“模型需要在长上下文中准确识别并遵循技能中的指令”,“长上下文”表述稍显冗余,可优化为“上下文中间”。\"},\"terminology\":{\"score\":4,\"evidence\":\"部分术语不一致或不规范,如“Agent”译为“代理”(技术文档中“智能体”更通用),“KV Cache”译为“KV缓存”(缺少括号标注原文),“instruction-following”译为“指令遵循”(正确),但“token”译为“标记”(不标准)。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown格式保留完整,标题、列表、图片、脚注、代码格式等均未丢失,与原文结构一致。\"}}},\"preferred\":\"X\",\"preference_evidence\":\"X在准确性(无关键术语误译)、术语规范性(“词元”“负面示例”“智能体”等标准译法)和流畅度(更自然的中文表达)上优于Y,且Markdown格式保真度相同。Y存在“上下文腐烂”“标记”等不规范术语,影响技术文档的专业性。\" }",
"usage": {
"prompt_tokens": 6695,
"completion_tokens": 681,
"total_tokens": 7376
}
},
"latency_milliseconds": 12523,
"validation": {
"valid": true
}
}
]
}