Files
ai-agent-book/chapter10/book-translation/validation/real_20260730T061500Z_v4/judge_receipts/unit-08.json
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

43 lines
31 KiB
JSON

{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nAutonomous Agents are well suited to open-ended problems—those where it is difficult or impossible to predict the number of steps required. Typical use cases include: Coding Agents solving SWE-bench (Software Engineering Benchmark, a benchmark for evaluating an Agent's ability to automatically fix real GitHub issues) tasks, \"Computer Use\" Agents operating computer interfaces like a human, and research tasks requiring iterative search and analysis.\n\nAutonomy also costs more and lets errors compound. Deploying an autonomous Agent therefore demands thorough testing in a sandbox, appropriate guardrails and monitoring, and human-in-the-loop checkpoints at critical decision points.\n\n#### Choosing and Mixing the Two Patterns\n\nIn practice, workflows and autonomous Agents are not mutually exclusive—many systems mix the two: critical processes with strict compliance requirements run as workflows for reliability, while the parts that need flexible decisions switch to autonomous mode. n8n, for example, is a mature open-source workflow automation framework in which developers build Agents by arranging functional components on a visual canvas—and workflow nodes and autonomous Agent nodes can coexist in the same system.\n\n![Figure 1-6: n8n workflow editor interface](images/n8n-workflow.png)\n\n#### Brief Comparison of Mainstream Agent Frameworks\n\nThe following table summarizes widely used Agent frameworks and platforms to help readers identify the right one for their scenario:\n\n| Harness Focus | Corresponding Chapter | Core Content | Security Concerns |\n|---------------|-----------------|------------------------------------|---------------------------|\n| Context design | Chapter 2 (Context Engineering) | Prompt Engineering, Agent Status Bar, Context Compression, Agent Skills | Prompt Injection and information leakage |\n| Context extension (knowledge persistence) | Chapter 3 (Knowledge Bases) | User memory, RAG, structured indexes, Agentic RAG | Exposure of sensitive information, privacy protection |\n| Tool design and security constraints | Chapter 4 (Tool Design) | Tool classification, permission control, MCP standards, asynchronous architecture | Misoperation, unauthorized access, irreversible operations |\n| Tool validation and correction | Chapter 5 (Code Generation) | Coding Agent Harnesses, test-driven development, rules encoded as code | Identity impersonation, attribution of responsibility |\n| System-level validation | Chapter 6 (Evaluation) | Evaluation environments, datasets, automated evaluation, observability | — |\n| Model-level correction | Chapter 7 (Post-training) | SFT (Supervised Fine-Tuning), reinforcement learning—writing feedback signals accumulated in the Harness into model parameters, which can be viewed as an extension of Harness engineering | Goal deviation, alignment, and robustness |\n| Experience-driven continuous correction | Chapter 8 (Continuous Evolution) | Trajectory learning signals; knowledge/instruction/program/parameter updates; self-modification; validation and rollback | Memory poisoning, unsafe self-modification, capability drift |\n| Multimodal context and tools | Chapter 9 (Multimodality and Real-Time Interaction) | Voice Agents, Computer Use, robotic manipulation | Security filtering of multimodal input, permission control in real-time interaction |\n| Constraints and correction among multiple Agents | Chapter 10 (Multi-Agent Collaboration) | Collaboration architectures, failure modes, Agent societies | Trust-boundary violations among Agents, conflicts over shared resources |\n\nAs the \"Model as Agent\" trend deepens, a framework's core value no longer lies in \"orchestrating LLM calls\"—models increasingly decide for themselves. What has grown more important is the Harness engineering around the model: context management, the tool ecosystem, security constraints, error recovery. When choosing a framework, the question is not how sophisticated the framework is, but whether it lets you focus on business logic through the thinnest possible layer of abstraction.\n\nOrchestration patterns solve the organization of context and tools within the Harness—how LLM calls, tools, and data flows connect. But task completion is not enough; tasks must also be completed correctly and safely. We therefore turn to the main way constrain, verify, and correct are implemented in practice: guardrails.\n\n### Guardrails and Safety\n\nThis section gives a high-level overview of guardrails to establish the big picture. Implementation details and practice follow in Chapter 2 (prompt injection protection), Chapter 4 (tool permission control), and Chapter 5 (code execution security); first-time readers do not need to follow every detail.\n\nGuardrails are how the \"constrain, verify, and correct\" layer of the Harness is primarily implemented—a layered defense that keeps Agent behavior safe and controllable. Well-designed **guardrails** help manage data privacy risks (for example, preventing system prompt leakage) and reputational risks (for example, keeping model behavior consistent with the brand). Start with guardrails for the risks you have already identified, then add new ones as new vulnerabilities surface.\n\nThink of guardrails as defense in depth. No single guardrail is likely to be sufficient on its own, but several specialized ones combined make a far more resilient Agent system.\n\n#### Types of Guardrails\n\nBased on where they sit in the execution flow, guardrails fall into three types: input-side, execution-side, and output-side.\n\n**Input-side** guardrails intercept requests before they reach the Agent, typically through four mechanisms. **Relevance classifiers** flag off-topic queries—for example, a coding assistant being asked, \"How tall is the Empire State Building?\" **Safety classifiers** detect jailbreaks (inducing the model to bypass its safety restrictions) and prompt injections (embedding malicious instructions in input). The key difference: in a jailbreak, the user tries to bypass the model's restrictions directly; in prompt injection, an attacker manipulates model behavior indirectly through external data (web content, documents). **Content moderation** flags harmful or inappropriate input, such as violent or discriminatory content. **Rule-based protections** apply deterministic measures—blacklists, input length limits, regular-expression filters—against known threats like SQL injection.\n\n**Execution-side** guardrails validate tool calls. The core is **tool risk rating**: based on whether an operation is reversible, its permission level, and financial impact, each tool is assigned a risk level (low/medium/high). High-risk operations require additional review or human confirmation.\n\n**Output-side** guardrails check the response before it is returned to the user. **PII filters** review the output for personally identifiable information (e.g., ID numbers, phone numbers) to prevent unnecessary exposure; **output validation** ensures the reply aligns with brand values through content checks.\n\nNote that some mechanisms (e.g., rule-based regex filtering) can be used on both the input and output sides; the above categorization follows the most common deployment locations.\n\nA representative industry practice of classifier-based guardrails is Anthropic's Constitutional Classifiers[^ch1-3]. Its design has three key elements. First, **rule-driven training**: a \"constitution\" written in natural language—which explicitly specifies what is allowed and what is not—is used to generate synthetic training data for the input and output classifiers. Second, **joint contextual judgment**: the new generation checks the user's question and the model's answer together, because some answers look perfectly fine on their own (e.g., \"how to use food flavorings\"), and only against the question does it become clear that \"food flavorings\" is code for chemical reagents. Third, **two-stage screening**: an extremely lightweight probe—which reads the model's internal activations at almost zero cost—checks every conversation first, and anything suspicious is escalated to a more powerful classifier for review rather than being refused outright. This way the first stage can tolerate more false positives without hurting the user experience, and the overall cost is greatly reduced.\n\n[^ch1-3]: Anthropic. \"Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks\", 2026. https://www.anthropic.com/research/next-generation-constitutional-classifiers; paper: Cunningham et al., \"Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks\", arXiv:2601.04603\n\n#### Human Intervention\n\n**Human-in-the-loop** intervention is a key protective measure: it lets an Agent improve real-world performance without degrading the user experience. It matters most in early deployment, when it helps identify failure modes, surface edge cases, and establish a robust evaluation cycle.\n\nWith a human-in-the-loop mechanism, an Agent that cannot complete a task can hand over control gracefully. In customer service, this means escalating to a human representative; for a Coding Agent, it means handing control back to the developer.\n\nThere are typically two main situations that trigger human intervention:\n\n**Exceeding Failure Thresholds**\nSet caps on the Agent's retries and operations. If the Agent exceeds those caps (for example, it still cannot infer the customer's intent after several attempts), escalate to a human.\n\n**High-Risk Operations**\nSensitive, irreversible, or high-risk operations should trigger human oversight—at least until the team has built enough confidence in the Agent's reliability. Typical examples: canceling a user's order, authorizing a large refund, processing a payment.\n\nWith the five Harness elements in mind, the rest of the book follows this structure.\n\n### This Book as a Practical Guide to Harness Engineering\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第8/9部分]\n自主代理非常适合开放式问题——那些难以或不可能预测所需步骤数量的问题。典型用例包括:解决SWE-bench(软件工程基准,评估代理自动修复真实GitHub问题能力的基准)任务的编码代理、像人类一样操作计算机界面的“计算机使用”代理,以及需要迭代搜索和分析的研究任务。\n\n自主性也成本更高,且会让错误累积。因此,部署自主代理需要在沙盒中进行彻底测试、设置适当的护栏和监控,并在关键决策点设置人工参与的检查点。\n\n#### 选择和混合两种模式\n在实践中,工作流和自主代理并非相互排斥——许多系统混合使用两者:具有严格合规要求的关键流程以工作流形式运行以确保可靠性,而需要灵活决策的部分切换到自主模式。例如,n8n是一个成熟的开源工作流自动化框架,开发者通过在可视化画布上排列功能组件来构建代理——工作流节点和自主代理节点可以在同一系统中共存。\n\n![图1-6:n8n工作流编辑器界面](images/n8n-workflow.png)\n\n#### 主流代理框架简要比较\n下表总结了广泛使用的代理框架和平台,帮助读者为自己的场景找到合适的框架:\n\n| 框架关注点 | 对应章节 | 核心内容 | 安全关注点 |\n|------------------|------------------------|------------------------------------------|----------------------|\n| 上下文设计 | 第2章(上下文工程) | 提示工程、代理状态栏、上下文压缩、代理技能 | 提示注入和信息泄露 |\n| 上下文扩展(知识持久化) | 第3章(知识库) | 用户记忆、RAG、结构化索引、代理式RAG | 敏感信息暴露、隐私保护 |\n| 工具设计和安全约束 | 第4章(工具设计) | 工具分类、权限控制、MCP标准、异步架构 | 误操作、未授权访问、不可逆转操作 |\n| 工具验证和纠正 | 第5章(代码生成) | 编码代理框架、测试驱动开发、编码规则 | 身份冒充、责任归属 |\n| 系统级验证 | 第6章(评估) | 评估环境、数据集、自动化评估、可观测性 | — |\n| 模型级纠正 | 第7章(后训练) | SFT(监督微调)、强化学习——将框架中积累的反馈信号写入模型参数,可视为框架工程的扩展 | 目标偏差、对齐和鲁棒性 |\n| 经验驱动的持续纠正 | 第8章(持续演进) | 轨迹学习信号;知识/指令/程序/参数更新;自我修改;验证和回滚 | 内存中毒、不安全自我修改、能力漂移 |\n| 多模态上下文和工具 | 第9章(多模态和实时交互) | 语音代理、计算机使用、机器人操作 | 多模态输入的安全过滤、实时交互中的权限控制 |\n| 多代理之间的约束和纠正 | 第10章(多代理协作) | 协作架构、失败模式、代理社会 | 代理之间的信任边界违反、共享资源冲突 |\n\n随着“模型即代理”趋势的深化,框架的核心价值不再在于“编排LLM调用”——模型越来越自行决策。变得更重要的是围绕模型的框架工程:上下文管理、工具生态系统、安全约束、错误恢复。选择框架时,问题不是框架有多复杂,而是它是否让你通过尽可能薄的抽象层专注于业务逻辑。\n\n编排模式解决框架内上下文和工具的组织方式——LLM调用、工具和数据流如何连接。但仅完成任务是不够的;任务还必须正确且安全地完成。因此,我们转向实践中实现约束、验证和纠正的主要方式:护栏。\n\n### 护栏和安全\n本节对护栏进行高层次概述,建立整体图景。实现细节和实践在第2章(提示注入保护)、第4章(工具权限控制)和第5章(代码执行安全)中后续介绍;首次阅读的读者无需关注每个细节。\n\n护栏是框架“约束、验证和纠正”层的主要实现方式——一种分层防御,保持代理行为安全可控。设计良好的**护栏**有助于管理数据隐私风险(例如,防止系统提示泄露)和声誉风险(例如,保持模型行为与品牌一致)。从已识别的风险开始设置护栏,然后随着新漏洞出现添加新的护栏。\n\n将护栏视为深度防御。单个护栏本身不太可能足够,但几个专门的护栏组合起来会形成更具弹性的代理系统。\n\n#### 护栏的类型\n根据它们在执行流程中的位置,护栏分为三种类型:输入侧、执行侧和输出侧。\n\n**输入侧**护栏在请求到达代理之前拦截它们,通常通过四种机制。**相关性分类器**标记离题查询——例如,编码助手被问到“帝国大厦有多高?”**安全分类器**检测越狱(诱导模型绕过其安全限制)和提示注入(在输入中嵌入恶意指令)。关键区别在于:在越狱中,用户直接尝试绕过模型的限制;在提示注入中,攻击者通过外部数据(网络内容、文档)间接操纵模型行为。**内容审核**标记有害或不适当的输入,例如暴力或歧视性内容。**基于规则的保护**应用确定性措施——黑名单、输入长度限制、正则表达式过滤——对抗已知威胁如SQL注入。\n\n**执行侧**护栏验证工具调用。核心是**工具风险评级**:根据操作是否可逆、权限级别和财务影响,每个工具被分配风险级别(低/中/高)。高风险操作需要额外审查或人类确认。\n\n**输出侧**护栏在响应返回给用户之前检查它。**PII过滤器**审查输出中的个人身份信息(例如,身份证号码、电话号码)以防止不必要的暴露;**输出验证**通过内容检查确保回复符合品牌价值。\n\n注意,一些机制(例如,基于规则的正则表达式过滤)可以在输入侧和输出侧使用;上述分类遵循最常见的部署位置。\n\n基于分类器的护栏的一个代表性行业实践是Anthropic的宪法分类器[^ch1-3]。其设计有三个关键要素。首先,**规则驱动训练**:用自然语言编写的“宪法”——明确指定允许和不允许的内容——用于为输入和输出分类器生成合成训练数据。其次,**联合上下文判断**:新一代检查用户的问题和模型的答案一起,因为有些答案单独看完全没问题(例如,“如何使用食品调味料”),只有结合问题才会清楚“食品调味料”是化学试剂的暗语。第三,**两阶段筛选**:一个极其轻量的探测器——几乎不费成本读取模型的内部激活——首先检查每个对话,任何可疑的都升级到更强大的分类器审查,而不是直接拒绝。这样第一阶段可以容忍更多假阳性而不影响用户体验,整体成本大大降低。\n\n[^ch1-3]: Anthropic. \"Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks\", 2026. https://www.anthropic.com/research/next-generation-constitutional-classifiers; paper: Cunningham et al., \"Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks\", arXiv:2601.04603\n\n#### 人工干预\n**人工参与**干预是关键的保护措施:它让代理在不降低用户体验的情况下提高真实世界性能。在早期部署中最重要,此时它有助于识别失败模式、暴露边缘情况并建立稳健的评估循环。\n\n通过人工参与机制,无法完成任务的代理可以优雅地移交控制权。在客户服务中,这意味着升级到人类代表;对于编码代理,这意味着将控制权交还给开发者。\n\n通常有两种主要情况触发人工干预:\n\n**超过失败阈值**\n设置代理重试和操作的上限。如果代理超过这些上限(例如,几次尝试后仍无法推断客户意图),升级到人类。\n\n**高风险操作**\n敏感、不可逆转或高风险的操作应触发人工监督——至少在团队对代理的可靠性建立足够信心之前。典型示例:取消用户订单、授权大额退款、处理支付。\n\n牢记五个框架要素,本书其余部分遵循此结构。\n\n### 本书作为框架工程的实用指南\n\nANONYMOUS CHINESE Y:\n### 与人工智能代理入门 [第8/9部分]\n\n自主代理非常适合解决开放式问题——那些难以或不可能预测所需步骤数量的问题。典型用例包括:解决SWE-bench(软件工程基准,用于评估代理自动修复真实GitHub问题能力的基准)任务的编码代理、像人类一样操作计算机界面的“计算机使用”代理,以及需要迭代搜索和分析的研究任务。\n\n自主性成本也更高,且会让错误累积。因此,部署自主代理需要在沙盒中进行彻底测试、设置适当的防护措施和监控,并在关键决策点设置人工参与的检查点。\n\n#### 选择和混合两种模式\n\n实际上,工作流和自主代理并非相互排斥——许多系统混合了两者:具有严格合规要求的关键流程作为工作流运行以确保可靠性,而需要灵活决策的部分则切换到自主模式。例如,n8n是一个成熟的开源工作流自动化框架,开发人员通过在可视化画布上排列功能组件来构建代理——工作流节点和自主代理节点可以在同一系统中共存。\n\n![图1-6:n8n工作流编辑器界面](images/n8n-workflow.png)\n\n#### 主流代理框架简要比较\n\n下表总结了广泛使用的代理框架和平台,以帮助读者为自己的场景找到合适的框架:\n\n| 框架聚焦点 | 对应章节 | 核心内容 | 安全关注点 |\n|------------------|------------------------|------------------------------------------|--------------------------|\n| 上下文设计 | 第2章(上下文工程) | 提示工程、代理状态栏、上下文压缩、代理技能 | 提示注入和信息泄露 |\n| 上下文扩展(知识持久化) | 第3章(知识库) | 用户内存、RAG、结构化索引、代理式RAG | 敏感信息暴露、隐私保护 |\n| 工具设计和安全约束 | 第4章(工具设计) | 工具分类、权限控制、MCP标准、异步架构 | 误操作、未经授权访问、不可逆转操作 |\n| 工具验证和纠正 | 第5章(代码生成) | 编码代理框架、测试驱动开发、编码编码的规则 | 身份冒充、责任归属 |\n| 系统级验证 | 第6章(评估) | 评估环境、数据集、自动化评估、可观测性 | — |\n| 模型级纠正 | 第7章(训练后) | SFT(监督微调)、强化学习——将框架中积累的反馈信号写入模型参数,可视为框架工程的扩展 | 目标偏离、对齐和鲁棒性 |\n| 经验驱动的持续纠正 | 第8章(持续演进) | 轨迹学习信号;知识/指令/程序/参数更新;自我修改;验证和回滚 | 内存中毒、不安全的自我修改、能力漂移 |\n| 多模态上下文和工具 | 第9章(多模态和实时交互) | 语音代理、计算机使用、机器人操作 | 多模态输入的安全过滤、实时交互中的权限控制 |\n| 多个代理之间的约束和纠正 | 第10章(多代理协作) | 协作架构、失败模式、代理社会 | 代理之间的信任边界违规、共享资源冲突 |\n\n随着“模型即代理”趋势的深化,框架的核心价值不再在于“编排大语言模型调用”——模型越来越多地自行决策。变得更重要的是围绕模型的框架工程:上下文管理、工具生态系统、安全约束、错误恢复。选择框架时,问题不在于框架有多复杂,而在于它是否能让你通过尽可能薄的抽象层专注于业务逻辑。\n\n编排模式解决框架内上下文和工具的组织问题——大语言模型调用、工具和数据流如何连接。但仅完成任务是不够的;任务还必须正确且安全地完成。因此,我们转向实践中实施约束、验证和纠正的主要方式:防护措施。\n\n### 防护措施和安全性\n\n本节从高层级概述防护措施以建立整体图景。实现细节和实践将在第2章(提示注入防护)、第4章(工具权限控制)和第5章(代码执行安全)中介绍;首次阅读的读者无需关注每个细节。\n\n防护措施是框架中“约束、验证和纠正”层的主要实现方式——一种分层防御,使代理行为安全可控。设计良好的**防护措施**有助于管理数据隐私风险(例如,防止系统提示泄露)和声誉风险(例如,保持模型行为与品牌一致)。从已识别的风险的防护措施开始,然后在新漏洞出现时添加新的防护措施。\n\n将防护措施视为深度防御。单个防护措施本身不太可能足够,但几个专门的防护措施结合起来会形成更具弹性的代理系统。\n\n#### 防护措施的类型\n\n根据防护措施在执行流程中的位置,可分为输入侧、执行侧和输出侧三种类型。\n\n**输入侧**防护措施在请求到达代理之前拦截请求,通常通过四种机制。**相关性分类器**标记离题查询——例如,编码助手被问到“帝国大厦有多高?”。**安全分类器**检测越狱(诱导模型绕过其安全限制)和提示注入(在输入中嵌入恶意指令)。关键区别在于:在越狱中,用户试图直接绕过模型的限制;在提示注入中,攻击者通过外部数据(网络内容、文档)间接操纵模型行为。**内容审核**标记有害或不适当的输入,例如暴力或歧视性内容。**基于规则的防护**针对已知威胁(如SQL注入)应用确定性措施——黑名单、输入长度限制、正则表达式过滤。\n\n**执行侧**防护措施验证工具调用。核心是**工具风险评级**:根据操作是否可逆、其权限级别和财务影响,为每个工具分配风险级别(低/中/高)。高风险操作需要额外审查或人工确认。\n\n**输出侧**防护措施在响应返回给用户之前检查响应。**PII过滤器**检查输出中的个人身份信息(例如,身份证号码、电话号码)以防止不必要的暴露;**输出验证**通过内容检查确保回复符合品牌价值。\n\n请注意,一些机制(例如,基于规则的正则表达式过滤)可以在输入侧和输出侧使用;上述分类遵循最常见的部署位置。\n\n基于分类器的防护措施的一个代表性行业实践是Anthropic的宪法分类器[^ch1-3]。其设计有三个关键要素。首先,**规则驱动训练**:用自然语言编写的“宪法”——明确规定允许和不允许的内容——用于为输入和输出分类器生成合成训练数据。其次,**联合上下文判断**:新一代同时检查用户的问题和模型的答案,因为有些答案本身看起来完全没问题(例如,“如何使用食品香料”),只有结合问题才能清楚“食品香料”是化学试剂的暗语。第三,**两阶段筛选**:一个极其轻量级的探测器——几乎不花费成本读取模型的内部激活——首先检查每个对话,任何可疑的内容都会升级到更强大的分类器进行审查,而不是直接拒绝。这样第一阶段可以容忍更多的假阳性而不影响用户体验,整体成本大大降低。\n\n[^ch1-3]: Anthropic. “下一代宪法分类器:更高效地抵御通用越狱”, 2026. https://www.anthropic.com/research/next-generation-constitutional-classifiers; 论文:Cunningham等人,“宪法分类器++:高效的生产级抵御通用越狱的防御”, arXiv:2601.04603\n\n#### 人工干预\n\n**人工参与**干预是关键的保护措施:它让代理在不降低用户体验的情况下提高实际性能。在早期部署中最为重要,此时它有助于识别失败模式、暴露边缘情况并建立稳健的评估周期。\n\n通过人工参与机制,无法完成任务的代理可以优雅地移交控制权。在客户服务中,这意味着升级到人类代表;对于编码代理,这意味着将控制权交还给开发人员。\n\n通常有两种主要情况会触发人工干预:\n\n**超过失败阈值**\n设置代理重试和操作的上限。如果代理超过这些上限(例如,经过几次尝试仍无法推断客户意图),则升级到人类。\n\n**高风险操作**\n敏感、不可逆转或高风险的操作应触发人工监督——至少在团队对代理的可靠性建立足够信心之前。典型示例:取消用户订单、批准大额退款、处理付款。\n\n牢记框架的五个要素,本书其余部分遵循此结构。\n\n### 本书作为框架工程的实用指南"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853906385123ea4fa82759f0cae4f2ffa723d4f9dd614d6cb",
"model": "doubao-seed-1-6-250615",
"created": 1785390654,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, the definition of 'SWE-bench' is accurately translated as '软件工程基准,评估代理自动修复真实GitHub问题能力的基准' (matches the original). The table rows and technical details like 'MCP standards' are preserved.\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent, e.g., '自主性也成本更高,且会让错误累积' flows naturally. Minor awkwardness: '代理状态栏' (Agent Status Bar) is less intuitive but technically correct.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms: 'guardrails' consistently translated as '护栏', 'Harness' as '框架', 'RAG' retained, 'prompt injection' as '提示注入' throughout.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"All Markdown elements preserved: headings (####), image syntax, table structure, list formatting, and citation [^ch1-3] with correct URL and paper title in English.\"}},\"Y\":{\"accuracy\":{\"score\":4,\"evidence\":\"Incorrect translation of the citation: the original English title 'Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks' is mistranslated to Chinese as '下一代宪法分类器:更高效地抵御通用越狱' instead of retaining the original English title. This changes the reference's presentation.\"},\"fluency\":{\"score\":5,\"evidence\":\"Smoother phrasing in some sections, e.g., '解决开放式问题' (solving open-ended problems) and '大语言模型' (LLM) for clarity. Natural flow in '防护措施是框架中“约束、验证和纠正”层的主要实现方式'.\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent 'guardrails' translation: mostly '防护措施' but '护栏' in the table (e.g., '安全关注点' column). 'LLM' translated as '大语言模型' which is acceptable but less technical than X's retention of 'LLM'.\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Image syntax and table structure preserved, but the citation [^ch1-3] incorrectly translates the paper title into Chinese instead of keeping the original English, violating markdown code fidelity for references.\"}}},\"preferred\":\"X\",\"preference_evidence\":\"X has higher accuracy (no mistranslated citation), consistent terminology ('guardrails' as '护栏'), and full markdown_code_fidelity (original English citation preserved). Y's fluency is slightly better but is outweighed by X's superior accuracy, terminology consistency, and markdown fidelity.\"}",
"usage": {
"prompt_tokens": 5921,
"completion_tokens": 565,
"total_tokens": 6486
}
},
"latency_milliseconds": 15807,
"validation": {
"valid": true
}
}
]
}