Files
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

43 lines
31 KiB
JSON
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n> **Experiment 2-8 ★★: Several Useful Agent Status Bar Techniques**\n>\n> The `agent-status-bar` experimental framework implements five status bar techniques, each of which can be independently enabled or disabled:\n>\n> **Timestamp Tracking**: Adds a prefix in the format `[2025-09-14 10:30:45]` to user messages and tool responses (note: not placed in the system prompt, as that would break the KV Cache). This enables the Agent to understand temporal relationships and provides information for debugging and auditing. This technique also implements a time simulation feature, allowing the Agent to understand relationships like \"yesterday's files\" and \"today's modifications.\"\n>\n> **Tool Call Counter**: Maintains a global dictionary recording the number of times each tool has been called, annotating responses with \"Tool call #3 for 'read_file'.\" This explicit counting encourages the model to change strategy after repeated failures: after the first failure, check the path; after the second failure, list the directory; after the third, stop retrying and seek an alternative. Its deeper value lies in implicit cost awareness: the Agent can infer that it has already spent too many attempts on a particular operation.\n>\n> **TODO List Management**: Inspired by Manus's concept of \"manipulating attention through restatement,\" TODO List Management provides two dedicated tools: `rewrite_todo_list` and `update_todo_status`. Each TODO item includes a unique identifier, content, status (pending/in_progress/completed/cancelled), and a timestamp. From the perspective of cognitive load theory, the TODO list serves as external memory—just as humans write checklists when handling complex projects, the Agent also needs a place to record \"what has been done and what remains.\" Experimental data show that Agents with TODO support complete tasks in an average of 15 iterations, while those without it require 21 iterations and often miss subtasks.\n>\n> **Detailed Error Information**: Contains four layers—error type and description, full parameter JSON, call stack information, and targeted fix suggestions (e.g., when encountering a FileNotFoundError, suggest verifying the path, checking the working directory, and using absolute paths). When enabled, this information raises the Agent's error-recovery success rate from 60% to 95%. Instead of retrying blindly, the Agent can diagnose the failure and choose an alternative.\n>\n> **System State Awareness**: Injects information such as the current time, working directory, operating system type, shell environment, and Python version. Tracking the working directory is particularly critical—it is automatically updated after the Agent executes a `cd` command, ensuring subsequent operations are performed in the correct context. Operating system information enables the Agent to make platform-specific decisions (e.g., using `apt` on Linux, `brew` on macOS).\n>\n> These techniques produce an emergent effect when working together (i.e., limited effectiveness when used individually, but unexpectedly powerful results when combined). The combination of timestamps and tool counters allows the Agent to understand the frequency and temporal distribution of operations; the combination of TODO lists and system state enables the Agent to adjust task strategies based on the environment; and the combination of detailed error information and tool counters allows the Agent not only to change strategies after multiple failures but also to understand the reasons for failure.\n>\n> An Agent with all these techniques enabled is not merely a tool that executes instructions mechanically; it becomes a state-aware assistant. When a file is not found, it first checks the directory, then lists available files, and if still not found, marks the task as cancelled in the TODO and adds an alternative task. This adaptive behavior is something no single technique can achieve alone.\n>\n\n### From Readings to Strategy: The Agent's Perception of Physical Time\n\nAmong the five techniques in Experiment 2-8, timestamp tracking and the tool call counter may look like unrelated pieces of meta-information. Together, however, they point to a more fundamental capability: enabling the Agent to condition its behavior on physical time and adjust its pace accordingly. When a person is asked to \"write a paragraph in three minutes\" versus \"write a paragraph in thirty minutes,\" the output differs. For today's cutting-edge Agents, however, the output is often almost identical. The Agent struggles to determine whether a job is finished, whether an obstacle is permanent or temporary, or whether a tool call that has been running for three minutes is still making progress or has stalled. The author and collaborators refer to this missing capability as **time sense** and break it down into three measurable axes[^ch2-8]:\n\n- **Urgency**—The budget axis: Matching effort to the clock. When time is tight, deliver decisively under uncertainty; when time is ample, dig deeper, verify more, and polish further. It is bidirectional: low urgency does not mean \"do less,\" but rather \"do not stop yet; keep going.\"\n- **Persistence**—The endpoint axis: Distinguishing genuine blockers from transient ones and knowing whether a task is finished. Failure occurs at both extremes: repeatedly retrying an unrecoverable error (retrying a 410 Gone endpoint five times) or abandoning a recoverable failure too soon (asserting \"information not found\" after only two searches).\n- **Vigilance**—The monitoring axis: Treating unexpected timing in tool responses as evidence worth investigating. A call that should return in 500ms but takes 5 seconds, and a call that \"succeeds\" in 1ms but returns an empty body, are both signals—provided the Agent is monitoring these readings.\n\nThis three-axis framework maps directly onto the status bar: timestamps provide signals for urgency and vigilance, while the tool call counter provides signals for persistence. However, **simply showing the model these readings is not enough to change its behavior**. A benchmark compared four conditions: no timing information, raw timestamps only, timestamps plus instructions on how to interpret them, and an Agent-generated pace assessment. Raw timestamps performed almost the same as no timing information, differing by only two to three percentage points. What raised the pass rate from just over 10% to 4050%, an increase of 19 to 49 percentage points, was the operational guidance. In other words, the model can see `elapsed_ms=5000 expected_ms=500`, but it will not automatically adjust its pace. What it lacks is not the reading, but the **strategy for acting on that reading**.\n\nThis fills the gap left earlier in the section. The tool call counter can correct behavior with the single reading \"This is call #3 (3/3)\" because the decision rule is obvious: stop when the limit is reached. For pace judgments such as \"how much effort to spend\" or \"whether to route around this obstacle,\" the rules are less obvious, and the model cannot reliably infer the correct action from raw readings alone. An effective \"pace status bar\" therefore needs both the **reading** (how long the task has taken, whether this tool is slow, how many times this obstacle has been encountered) and a short **operational strategy** (deliver when time is tight, diagnose slow calls, route around hard blockers). Neither is sufficient alone. Explicit readings are raw material; the model also needs guidance that translates readings into actions.\n\nThis gap is not specific to any one model. Across six models from four vendor families—from Claude, Gemini, GPT to Qwen—without operational guidance, the pass rate stayed just above 10%. This suggests that current post-training often fails to teach time-sensitive control behavior, rather than showing a lack of intelligence in any particular model. The gap can be addressed at inference time with the \"status bar + operational guidance\" approach described above. If a smaller model needs this sense of pacing without relying on prompts, it can also be distilled into the weights. Chapter 7, on post-training, discusses this training path and an important contrast: sparse outcome rewards failed to induce the behavior, while dense token-level signals succeeded.\n\n[^ch2-8]: Li, Bojie and Noah Shi. *Agents That Sense Physical Time: Urgency, Persistence, and Vigilance as Missing Controls for LLM Agents.* 2026. https://01.me/research/physical-time-agent\n\n### Design Philosophy\n\nThis set of techniques has a practical advantage: all meta-information appears in the context in a human-readable form, allowing developers to inspect what information the Agent received and what decisions it made. More importantly, the approach requires no changes to the model. No fine-tuning is needed; the techniques work with any language model and can be tested individually or combined as needed.\n\n## Context Compression Strategies\n\nThe previous sections discussed what to include in context: prompt engineering determines what to write, Skills determine what to load on demand, and the Agent Status Bar determines what meta-information to inject. As multi-turn interactions deepen, however, the context keeps expanding. This section turns to the opposite problem: **how to reduce content in the context**—when to compress, how to compress, and why compression can be useful even before the context window is full.\n\n### Why Compression Is Needed: Not Just a Length Issue\n\nContext compression has two distinct motivations. Understanding both is crucial for designing an effective compression strategy.\n\n**First, addressing length and cost constraints.** This is the most intuitive reason: the context window is limited (e.g., 128K tokens), tool call results routinely run to tens of thousands of characters, and a few rounds of interaction can fill the window and cut the task short. More tokens also mean higher API costs and sharply higher inference latency.\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第14/17部分]\n\n#### 实验2-8★★:几种有用的Agent状态栏技术\n\n`agent-status-bar`实验框架实现了五种状态栏技术,每种技术都可以独立启用或禁用:\n\n**时间戳跟踪**:在用户消息和工具响应前添加格式为`[2025-09-14 10:30:45]`的前缀(注意:不放在系统提示中,否则会破坏键值缓存KV Cache)。这使Agent能够理解时间关系,并为调试和审计提供信息。该技术还实现了时间模拟功能,让Agent能够理解“昨天的文件”“今天的修改”等关系。\n\n**工具调用计数器**:维护一个全局字典记录每个工具被调用的次数,并用“对'read_file'的第3次工具调用”标注响应。这种明确的计数鼓励模型在多次失败后改变策略:第一次失败后检查路径;第二次失败后列出目录;第三次后停止重试并寻求替代方案。其更深层价值在于隐含的成本意识:Agent可以推断在特定操作上已经花费了太多尝试。\n\n**待办事项列表管理**:受马努斯“通过重述操纵注意力”的概念启发,待办事项列表管理提供了两个专用工具:`rewrite_todo_list`和`update_todo_status`。每个待办事项包括唯一标识符、内容、状态(待处理/进行中/已完成/已取消)和时间戳。从认知负荷理论角度看,待办事项列表充当外部记忆——就像人类处理复杂项目时写清单一样,Agent也需要记录“已做之事和待做之事”的地方。实验数据显示,支持待办事项的Agent平均15次迭代完成任务,而不支持的需要21次迭代且常遗漏子任务。\n\n**详细错误信息**:包含四层内容——错误类型和描述、完整参数JSON、调用栈信息和针对性修复建议(例如遇到FileNotFoundError时,建议验证路径、检查工作目录、使用绝对路径)。启用时,该信息将Agent的错误恢复成功率从60%提高到95%。Agent不再盲目重试,而是能够诊断失败并选择替代方案。\n\n**系统状态感知**:注入当前时间、工作目录、操作系统类型、shell环境和Python版本等信息。跟踪工作目录尤为关键——Agent执行`cd`命令后会自动更新,确保后续操作在正确上下文中进行。操作系统信息使Agent能够做出特定平台的决策(例如在Linux上使用`apt`,在macOS上使用`brew`)。\n\n这些技术共同作用时会产生涌现效应(即单独使用时效果有限,但组合使用时却有意想不到的强大结果)。时间戳和工具计数器的组合使Agent能够理解操作的频率和时间分布;待办事项列表和系统状态的组合使Agent能够根据环境调整任务策略;详细错误信息和工具计数器的组合使Agent不仅能在多次失败后改变策略,还能理解失败原因。\n\n启用所有这些技术的Agent不仅仅是机械执行指令的工具,它变成了一个状态感知型助手。当文件未找到时,它首先检查目录,然后列出可用文件,如果仍未找到,就在待办事项中标记任务为已取消并添加替代任务。这种自适应行为是任何单一技术都无法单独实现的。\n\n#### 从阅读到策略:Agent对物理时间的感知\n\n在实验2-8的五种技术中,时间戳跟踪和工具调用计数器看似是不相关的元信息,但它们共同指向一个更根本的能力:使Agent能够根据物理时间调整行为并相应调整节奏。当要求一个人“在三分钟内写一段文字”与“在三十分钟内写一段文字”时,输出会不同。然而,对于当今最先进的Agent来说,输出往往几乎相同。Agent难以确定工作是否完成、障碍是永久还是暂时、运行了三分钟的工具调用是仍在进展还是已停滞。作者及其合作者将这种缺失的能力称为**时间感知**,并将其分解为三个可衡量的维度[^ch2-8]:\n\n- **紧急性**——预算维度:根据时钟调整努力程度。时间紧迫时,在不确定情况下果断交付;时间充裕时,深入挖掘、更多验证、进一步完善。它是双向的:低紧急性不意味着“少做”,而是“还没完成;继续进行”。\n- **持久性**——终点维度:区分真正的障碍和短暂的障碍,知道任务是否完成。两种极端都会导致失败:反复重试不可恢复的错误(对410 Gone端点重试五次)或过早放弃可恢复的失败(仅搜索两次就断言“未找到信息”)。\n- **警觉性**——监控维度:将工具响应中的意外时间视为值得调查的证据。应该在500毫秒内返回但耗时5秒的调用,以及“成功”在1毫秒内返回但返回空体的调用,都是信号——前提是Agent在监控这些读数。\n\n这个三维框架直接映射到状态栏:时间戳提供紧急性和警觉性的信号,而工具调用计数器提供持久性的信号。然而,**仅仅向模型展示这些读数不足以改变其行为**。一项基准测试比较了四种情况:没有时间信息、只有原始时间戳、时间戳加上如何解释它们的说明,以及Agent生成的节奏评估。原始时间戳的表现几乎与没有时间信息相同,仅相差两到三个百分点。将通过率从刚超过10%提高到40-50%(提高了19到49个百分点)的是操作指导。换句话说,模型可以看到`elapsed_ms=5000 expected_ms=500`,但它不会自动调整节奏。它缺乏的不是读数,而是**对该读数采取行动的策略**。\n\n这填补了本节前面留下的空白。工具调用计数器可以用“这是第3次调用(3/3)”这一单一读数纠正行为,因为决策规则很明确:达到限制时停止。对于“花费多少精力”或“是否绕过这个障碍”等节奏判断,规则不那么明确,模型仅从原始读数无法可靠推断出正确行动。因此,有效的“节奏状态栏”需要既有**读数**(任务已花费多长时间、这个工具是否缓慢、遇到这个障碍多少次),又有简短的**操作策略**(时间紧迫时交付、诊断缓慢的调用、绕过顽固障碍)。两者单独都不充分。明确的读数是原材料;模型还需要将读数转化为行动的指导。\n\n这个空白并非特定于任何一个模型。在来自四个厂商家族的六个模型中——从Claude、Gemini、GPT到Qwen——没有操作指导时,通过率仅略高于10%。这表明当前的后训练往往未能教授时间敏感的控制行为,而不是任何特定模型缺乏智能。可以在推理时通过上述“状态栏+操作指导”的方法解决这个空白。如果较小的模型需要这种节奏感知而不依赖提示,也可以将其提炼到权重中。第7章关于后训练的内容将讨论这条训练路径以及一个重要对比:稀疏结果奖励未能诱导出这种行为,而密集词元级信号成功了。\n\n[^ch2-8]: 李博杰和诺亚·石。《感知物理时间的Agent:紧急性、持久性和警觉性是大语言模型Agent缺失的控制》。2026年。https://01.me/research/physical-time-agent\n\n#### 设计理念\n\n这套技术有一个实际优势:所有元信息都以人类可读的形式出现在上下文中,允许开发者检查Agent接收到的信息和做出的决策。更重要的是,该方法不需要对模型进行修改。不需要微调;这些技术适用于任何语言模型,可以根据需要单独测试或组合使用。\n\n### 上下文压缩策略\n\n前面的章节讨论了上下文中应包含什么:提示工程决定写什么,技能决定按需加载什么,Agent状态栏决定注入什么元信息。然而,随着多轮交互的深入,上下文不断扩展。本节转向相反的问题:**如何减少上下文中的内容**——何时压缩、如何压缩,以及为什么即使在上下文窗口未满时压缩也可能有用。\n\n#### 为什么需要压缩:不仅仅是长度问题\n\n上下文压缩有两个不同的动机。理解这两者对于设计有效的压缩策略至关重要。\n\n**首先,应对长度和成本限制**。这是最直观的原因:上下文窗口有限(例如128K词元),工具调用结果通常长达数万字符,几轮交互就能填满窗口并中断任务。更多词元也意味着更高的API成本和急剧增加的推理时延。\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第14/17部分]\n> **实验2-8 ★★:几种有用的代理状态栏技术**\n>\n> `agent-status-bar`实验框架实现了五种状态栏技术,每种技术都可以独立启用或禁用:\n>\n> **时间戳跟踪**:在用户消息和工具响应中添加格式为`[2025-09-14 10:30:45]`的前缀(注意:不放在系统提示中,因为那样会破坏KV缓存)。这使代理能够理解时间关系,并为调试和审计提供信息。该技术还实现了时间模拟功能,允许代理理解“昨天的文件”和“今天的修改”等关系。\n>\n> **工具调用计数器**:维护一个全局字典记录每个工具被调用的次数,用“对'read_file'的第3次工具调用”标注响应。这种显式计数鼓励模型在多次失败后改变策略:第一次失败后,检查路径;第二次失败后,列出目录;第三次后,停止重试并寻求替代方案。其更深层的价值在于隐含的成本意识:代理可以推断出它在特定操作上已经花费了太多尝试。\n>\n> **待办事项列表管理**:受Manus的“通过重述操纵注意力”概念启发,待办事项列表管理提供两个专用工具:`rewrite_todo_list`和`update_todo_status`。每个待办事项包括唯一标识符、内容、状态(待办/进行中/已完成/已取消)和时间戳。从认知负荷理论的角度看,待办事项列表充当外部记忆——就像人类处理复杂项目时编写清单一样,代理也需要一个记录“已完成和剩余事项”的地方。实验数据显示,支持待办事项的代理平均在15次迭代中完成任务,而没有该功能的代理需要21次迭代且经常遗漏子任务。\n>\n> **详细错误信息**:包含四层内容——错误类型和描述、完整参数JSON、调用栈信息和针对性修复建议(例如,遇到FileNotFoundError时,建议验证路径、检查工作目录并使用绝对路径)。启用时,该信息将代理的错误恢复成功率从60%提高到95%。代理不再盲目重试,而是可以诊断失败并选择替代方案。\n>\n> **系统状态感知**:注入当前时间、工作目录、操作系统类型、shell环境和Python版本等信息。跟踪工作目录尤其关键——代理执行`cd`命令后会自动更新,确保后续操作在正确上下文中进行。操作系统信息使代理能够做出特定平台的决策(例如,在Linux上使用`apt`,在macOS上使用`brew`)。\n>\n> 这些技术一起使用时会产生涌现效应(即单独使用时效果有限,但组合使用时意外强大)。时间戳和工具计数器的组合使代理能够理解操作的频率和时间分布;待办事项列表和系统状态的组合使代理能够根据环境调整任务策略;详细错误信息和工具计数器的组合使代理不仅能在多次失败后改变策略,还能理解失败的原因。\n>\n> 启用所有这些技术的代理不仅仅是机械执行指令的工具;它成为一个状态感知助手。当文件未找到时,它首先检查目录,然后列出可用文件,如果仍未找到,将任务标记为已取消并添加替代任务。这种自适应行为是任何单一技术都无法单独实现的。\n>\n\n### 从读数到策略:代理对物理时间的感知\n在实验2-8的五种技术中,时间戳跟踪和工具调用计数器看似是不相关的元信息。然而,它们共同指向一个更基本的能力:使代理能够根据物理时间调整行为并相应调整节奏。当一个人被要求“在三分钟内写一段”与“在三十分钟内写一段”时,输出不同。然而,对于当今最先进的代理来说,输出往往几乎相同。代理难以确定工作是否完成、障碍是永久还是暂时、运行了三分钟的工具调用是仍在进展还是已停滞。作者及其合作者将这种缺失的能力称为**时间感知**,并将其分解为三个可衡量的轴[^ch2-8]:\n\n- **紧急程度**——预算轴:根据时钟匹配努力。时间紧迫时,在不确定情况下果断交付;时间充裕时,深入挖掘、更多验证、进一步完善。它是双向的:低紧急程度不意味着“少做”,而是“还没停止;继续进行”。\n- **持久性**——终点轴:区分真正的障碍和短暂的障碍,并知道任务是否完成。两种极端都会导致失败:反复重试不可恢复的错误(五次重试410 Gone端点)或过早放弃可恢复的失败(仅两次搜索后断言“未找到信息”)。\n- **警觉性**——监控轴:将工具响应中的意外时间视为值得调查的证据。应该在500ms内返回但花费5秒的调用,以及“成功”返回但主体为空的调用,都是信号——前提是代理在监控这些读数。\n\n这个三轴框架直接映射到状态栏:时间戳提供紧急程度和警觉性的信号,而工具调用计数器提供持久性的信号。然而,**仅向模型展示这些读数不足以改变其行为**。一个基准比较了四种条件:没有时间信息、仅原始时间戳、时间戳加上如何解释它们的指令、代理生成的节奏评估。原始时间戳的表现几乎与没有时间信息相同,仅相差两到三个百分点。将通过率从略高于10%提高到40-50%(提高了19到49个百分点)的是操作指南。换句话说,模型可以看到`elapsed_ms=5000 expected_ms=500`,但它不会自动调整节奏。它缺少的不是读数,而是**对该读数采取行动的策略**。\n\n这填补了本节前面留下的空白。工具调用计数器可以用“这是第3次调用(3/3)”的单一读数纠正行为,因为决策规则很明显:达到限制时停止。对于“花费多少努力”或“是否绕过此障碍”等节奏判断,规则不太明显,模型仅从原始读数无法可靠推断出正确行动。因此,有效的“节奏状态栏”需要既有**读数**(任务已花费多长时间、此工具是否缓慢、遇到此障碍多少次),又有简短的**操作策略**(时间紧迫时交付、诊断缓慢调用、绕过硬障碍)。两者单独都不充分。显式读数是原材料;模型还需要将读数转化为行动的指导。\n\n这个空白不是任何一个模型特有的。在来自四个供应商家族的六个模型中——从Claude、Gemini、GPT到Qwen——没有操作指南时,通过率仅略高于10%。这表明当前的后训练通常未能教授时间敏感的控制行为,而不是任何特定模型缺乏智能。可以通过上述“状态栏+操作指南”的方法在推理时解决这个空白。如果较小的模型需要这种节奏感知而不依赖提示,也可以将其提炼到权重中。第7章关于后训练的内容讨论了这种训练路径和一个重要对比:稀疏结果奖励未能诱导出该行为,而密集标记级信号成功了。\n\n[^ch2-8]: Li, Bojie and Noah Shi. *Agents That Sense Physical Time: Urgency, Persistence, and Vigilance as Missing Controls for LLM Agents.* 2026. https://01.me/research/physical-time-agent\n\n### 设计理念\n这套技术有一个实际优势:所有元信息都以人类可读的形式出现在上下文中,允许开发者检查代理收到了什么信息以及做出了什么决策。更重要的是,该方法不需要对模型进行更改。不需要微调;这些技术适用于任何语言模型,可以根据需要单独测试或组合使用。\n\n## 上下文压缩策略\n前面的章节讨论了上下文中应包含什么:提示工程决定写什么,技能决定按需加载什么,代理状态栏决定注入什么元信息。然而,随着多轮交互的深入,上下文不断扩展。本节转向相反的问题:**如何减少上下文中的内容**——何时压缩、如何压缩,以及为什么即使在上下文窗口填满之前压缩也可能有用。\n\n### 为什么需要压缩:不仅仅是长度问题\n上下文压缩有两个不同的动机。理解这两者对于设计有效的压缩策略至关重要。\n\n**首先,解决长度和成本约束**。这是最直观的原因:上下文窗口有限(例如128K标记),工具调用结果通常长达数万字符,几轮交互就可能填满窗口并缩短任务。更多标记也意味着更高的API成本和急剧增加的推理延迟。"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853909371043ea4fa82759f0cae4f2ffa723d4f9dd6a7436a",
"model": "doubao-seed-1-6-250615",
"created": 1785390952,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims found. For example, 'timestamp tracking' is accurately translated as '时间戳跟踪', and technical details like 'FileNotFoundError' are preserved. The experimental data '15 iterations vs 21 iterations' is correctly retained.\"},\"fluency\":{\"score\":4,\"evidence\":\"Most sentences are fluent, e.g., '从认知负荷理论角度看,待办事项列表充当外部记忆——就像人类处理复杂项目时写清单一样,Agent也需要记录“已做之事和待做之事”的地方。' However, '这使Agent能够理解时间关系' could be more natural as '这让智能体能够理解时间关系' (using '智能体' instead of 'Agent' for consistency).\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms: 'KV Cache' (键值缓存), 'call stack' (调用栈), 'cognitive load theory' (认知负荷理论), 'emergent effect' (涌现效应). 'TODO List Management' is consistently translated as '待办事项列表管理'.\"},\"markdown_code_fidelity\":{\"score\":3,\"evidence\":\"Headings are demoted (e.g., original '## Context Compression Strategies' becomes '### 上下文压缩策略' in X). Fenced code blocks and equations are preserved, but heading hierarchy is altered, reducing structural fidelity.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All content is accurately translated without omissions or inventions. For instance, 'error-recovery success rate from 60% to 95%' is correctly translated, and 'Manus' is retained as 'Manus'. The citation '[^ch2-8]' is properly preserved.\"},\"fluency\":{\"score\":5,\"evidence\":\"Highly fluent and natural, e.g., '当文件未找到时,它首先检查目录,然后列出可用文件,如果仍未找到,将任务标记为已取消并添加替代任务。' '代理' (instead of 'Agent') is consistently used, enhancing readability.\"},\"terminology\":{\"score\":4,\"evidence\":\"Most terms are correct, but 'persistence' is translated as '持久性' (accurate), while 'vigilance' is '警觉性' (correct). However, 'TODO item status 'pending'' is translated as '待办' instead of '待处理' (inconsistent with X's '待处理', though both are acceptable).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown structure is fully preserved: headings (e.g., '## 上下文压缩策略'), lists, code blocks (e.g., `rewrite_todo_list`), and citations ([^ch2-8]) maintain original hierarchy and formatting.\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher fluency (consistently using '代理' instead of mixing 'Agent' and '智能体') and perfect markdown_code_fidelity (preserving heading hierarchy), while matching X in accuracy and terminology. X's heading demotion reduces structural clarity compared to Y's faithful markdown preservation.\"}",
"usage": {
"prompt_tokens": 6080,
"completion_tokens": 668,
"total_tokens": 6748
}
},
"latency_milliseconds": 14972,
"validation": {
"valid": true
}
}
]
}