{ "schema_version": 1, "credential_free": true, "attempts": [ { "attempt": 1, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nA concrete example shows the value of the Harness. Suppose you ask an Agent to refund a user's order placed 3 days ago. **Without a Harness**: the model does not receive the refund policy (no context), does not know which API to call (no tools), fabricates a refund result for the user (no verification), and the user discovers the refund never happened (no correction). **With a Harness**: the system prompt specifies the 7-day refund policy (context), the Agent calls the `query_order` and `process_refund` tools to perform the operation (tools), the framework checks that the refund does not exceed the order total (constrain), confirms against the database that the refund went through (verify), and automatically retries if the API call times out (correct). Same model, substantially different results.\n\nIn short, a model without a Harness may be highly capable, but it lacks the surrounding controls needed for reliable task completion.\n\nMore precisely, all infrastructure outside the model belongs to the Harness. The core of the Harness is Context and Tools, around which three types of engineering safeguards are built:\n\n| Function | One-Sentence Responsibility | Relationship with Context/Tools |\n|----------|-------------------------------------------|------------------------------------------|\n| **Context** | Provides the model with relevant information | Core capability |\n| **Tools** | Provides the model with action interfaces | Core capability |\n| **Constrain** | Sets behavioral boundaries—what can and cannot be done | Safety boundary built around context and tools |\n| **Verify** | Automatically judges the correctness of tool execution results | Checking mechanism built around tool execution results |\n| **Correct** | Automatically recovers or rolls back when problems are found | Recovery mechanism built around tool call failures |\n\nContext and Tools let the Agent complete tasks—understand the task and act on it. Constrain, Verify, and Correct make sure it does so reliably and safely—not as something apart from Context and Tools, but as the engineering that keeps them working reliably in production. Along the maturity curve of Agent products, the emphasis between these two groups shifts.\n\nEarly Agent frameworks focused on Context and Tools: give the model tools, give it context, and let it complete tasks. Production-grade systems have shifted their center of gravity to Constrain, Verify, and Correct: making sure tool calls are safe, context is managed, and errors are recoverable.\n\nTake Claude Code. The vast majority of its Harness code does Constrain, Verify, and Correct, not Context and Tools—the tools themselves (file read/write, command execution, search) are only a small part; the safeguards built around them are the true core. These mechanisms include:\n\n- **Process State Management**: Tracks which step the Agent is currently executing\n- **Multi-Layer Context Compression**: Automatically prunes information when there is too much\n- **Permission Classification**: Controls which operations require user confirmation\n- **Circuit Breaker**: Automatically stops retrying after repeated errors so one failing operation does not cascade through the whole system\n- **Error Recovery Mechanisms**: Catches exceptions, rolls back to the last stable state, retries, or hands off to a human\n\n**The industry is shifting from task completion to reliable task completion, making Harness Engineering the core competitive advantage of Agent systems.**\n\n### From Prompt Engineering to Loop Engineering: The Evolution of Engineering Paradigms\n\nLooking back at the development of AI application engineering, a clear evolutionary arc emerges:\n\n**Software Engineering** is the foundation—traditional system design, architecture, testing, and deployment. **Prompt Engineering** was the first wave of innovation—improving output quality by refining the natural-language instructions fed to the model. **Context Engineering** was the second wave—the realization that optimizing the prompt alone is not enough: the model's working context (system instructions, tool definitions, conversation history, external knowledge) has to be managed systematically. **Harness Engineering** was the third wave—it widens the view from \"what information the model receives\" to \"what kind of system the model runs in,\" taking in all infrastructure outside the model: constraint mechanisms, verification methods, feedback loops, error recovery. **Loop Engineering** came next, widening the view from a single run to sustained autonomous operation across runs: who discovers the next piece of work, when to verify, and when the task counts as truly done (Chapter 10 develops this alongside multi-agent collaboration systems).\n\nIn July 2026, the industry began using **Graph Engineering** for a higher-level orchestration perspective: organizing Agent loops, deterministic programs, and human approvals into an explicit execution graph, where nodes provide capabilities, edges define routing and dependencies, and structured state travels along those edges and is persisted at key boundaries.[^ch1-graph-engineering] Graph Engineering is not a replacement for Loop Engineering, nor should it simply be treated as a \"sixth layer\" in the progression above. A loop is itself a graph with a back edge, and a node in the graph can still run ReAct or another Agent loop internally. The name has not yet stabilized, so this book treats it as emerging terminology for existing orchestration and Harness practices; Chapter 10 develops the multi-agent part. Here, \"graph\" means a control-flow or execution graph, not the knowledge graph used by GraphRAG.\n\n[^ch1-graph-engineering]: Josh C. Simmons used the name explicitly in his July 4, 2026 article *We Are Entering the Graph Engineering Phase*, summarizing it in terms of nodes, typed edges, and checkpointed state. On July 18, Peter Steinberger's question about whether the discussion had shifted from loops to graphs helped the name spread further. The practices predate the label: the official documentation for LangGraph, Microsoft Agent Framework, and Google ADK describes them as graph orchestration or graph-based workflows. See https://www.drjoshcsimmons.com/writing/we-are-entering-the-graph-engineering-phase, https://x.com/steipete/status/2078277297791189132, https://docs.langchain.com/oss/python/langgraph/overview, https://learn.microsoft.com/en-us/agent-framework/workflows/, and https://adk.dev/workflows/.\n\nThese five stages are not replacements but nested layers: Prompt Engineering is a subset of Context Engineering, which is a subset of Harness Engineering, which is a subset of Loop Engineering. Each layer widens the engineer's scope of concern and influence beyond the last. **As models converge in capability and stop being the decisive differentiator, competitive advantage shifts to the engineering outside the model.** Recent engineering practice supports this view. LangChain's work on Terminal Bench 2.0 (a benchmark evaluating an Agent's ability to complete complex tasks in a terminal environment) is a striking example: their Coding Agent improved from 52.8% to 66.5% (jumping from outside the top 30 to the top 5 on the leaderboard). What changed was not the model but the Harness—having the Agent check its own execution results, detect when it was stuck in a repetitive loop, and refine its reasoning strategy. OpenAI's engineering team has shared a similar experience: 3 engineers completed approximately one million lines of code and nearly 1500 PRs in 5 months, about 10 times traditional development speed. The main driver was not a stronger model; it was getting the Harness right.\n\n### Core Principles of the Five Harness Functions\n\nThe earlier table listed the Harness's five functions. The table below adds each function's core design principle and where this book treats it, mapping concept to practice:\n\n| Function | Core Principle | Practical Example | See Chapter |\n|----------|------------------------------------------|----------------------------------|---------|\n| **Context** | Information Sufficiency: Ensure the Agent makes decisions based on sufficient information at every decision point | System prompts, knowledge bases, Agent status bars, Sidecar bypass queries | Chapters 2 & 3 |\n| **Tools** | Clear Interface: Tool names are intuitive, parameters have examples, boundaries are explained | MCP tools, code interpreter, search tools | Chapter 4 |\n| **Constrain** | Fail-Safe Defaults: All capabilities are off by default and must be explicitly enabled (similar to mobile app permission management) | In Claude Code, every tool requires user authorization by default before execution | Chapter 4 |\n| **Verify** | Input Isolation: Security checks only look at structured data (e.g., JSON fields returned by tools), not free-form text generated by the model (because attackers might manipulate model output through prompt injection) | Linter checks, type systems, tool call result validation | Chapters 5 & 6 |\n| **Correct** | Do not expose intermediate states until a failure is confirmed unrecoverable (e.g., silently retry a failed tool call instead of showing the user a half-finished result) | Silent retries, continuation generation, fallback to human judgment upon consecutive failures (circuit breaker mechanism) | Chapters 2 & 5 |\n\nThe five functions form a closed loop: Context and Tools support decision-making, Constrain prevents errors, Verify detects deviations, and Correct closes the cycle. If any link is missing, the system develops a reliability gap. Before examining specific orchestration patterns and guardrail designs, we first lay out the core principles for building effective Agents and for choosing a model—the foundation for every design decision that follows.\n\n### Core Principles for Building Effective Agents\n\nBased on Anthropic's experience, successful Agent systems follow three core principles.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第6/9部分]\n一个具体的例子展示了框架的价值。假设你让一个代理退还用户3天前下的订单。**没有框架**:模型没有收到退款政策(没有上下文),不知道调用哪个API(没有工具),为用户编造退款结果(没有验证),用户发现退款从未发生(没有纠正)。**有框架**:系统提示指定了7天退款政策(上下文),代理调用`query_order`和`process_refund`工具执行操作(工具),框架检查退款金额不超过订单总额(约束),与数据库确认退款已完成(验证),如果API调用超时自动重试(纠正)。同样的模型,结果大不相同。\n\n简而言之,没有框架的模型可能能力很强,但缺乏可靠完成任务所需的周围控制。\n\n更准确地说,模型之外的所有基础设施都属于框架。框架的核心是上下文和工具,围绕它们构建了三类工程保障:\n\n| 功能 | 一句话职责 | 与上下文/工具的关系 |\n|----------|----------------------------------|---------------------------|\n| **上下文** | 为模型提供相关信息 | 核心能力 |\n| **工具** | 为模型提供行动接口 | 核心能力 |\n| **约束** | 设定行为边界——能做和不能做的事 | 围绕上下文和工具的安全边界 |\n| **验证** | 自动判断工具执行结果的正确性 | 围绕工具执行结果的检查机制 |\n| **纠正** | 发现问题时自动恢复或回滚 | 围绕工具调用失败的恢复机制 |\n\n上下文和工具让代理完成任务——理解任务并采取行动。约束、验证和纠正确保它可靠且安全地完成任务——不是与上下文和工具分离的东西,而是让它们在生产中可靠运行的工程。随着代理产品的成熟度曲线,这两组之间的重点发生转移。\n\n早期的代理框架专注于上下文和工具:给模型工具,给它上下文,让它完成任务。生产级系统已将重心转移到约束、验证和纠正:确保工具调用安全、上下文得到管理、错误可恢复。\n\n以Claude Code为例。它的框架代码绝大多数用于约束、验证和纠正,而不是上下文和工具——工具本身(文件读写、命令执行、搜索)只是很小的一部分;围绕它们构建的保障措施才是真正的核心。这些机制包括:\n\n- **进程状态管理**:跟踪代理当前执行的步骤\n- **多层上下文压缩**:信息过多时自动修剪\n- **权限分类**:控制哪些操作需要用户确认\n- **断路器**:重复错误后自动停止重试,防止一个失败操作级联影响整个系统\n- **错误恢复机制**:捕获异常、回滚到最后稳定状态、重试或移交人类处理\n\n**行业正在从完成任务转向可靠完成任务,使框架工程成为代理系统的核心竞争力。**\n\n### 从提示工程到循环工程:工程范式的演进\n回顾人工智能应用工程的发展,出现了一条清晰的演进弧线:\n\n**软件工程**是基础——传统的系统设计、架构、测试和部署。**提示工程**是第一波创新——通过完善喂给模型的自然语言指令来提高输出质量。**上下文工程**是第二波——意识到仅优化提示是不够的:模型的工作上下文(系统指令、工具定义、对话历史、外部知识)必须系统地管理。**框架工程**是第三波——它将视角从“模型接收什么信息”扩展到“模型运行在什么样的系统中”,纳入模型之外的所有基础设施:约束机制、验证方法、反馈循环、错误恢复。**循环工程**紧随其后,将视角从单次运行扩展到跨运行的持续自主操作:谁发现下一个工作、何时验证、何时任务才算真正完成(第10章与多代理协作系统一起发展这部分)。\n\n2026年7月,行业开始使用**图工程**从更高层次的编排视角:将代理循环、确定性程序和人类审批组织成显式的执行图,其中节点提供能力,边定义路由和依赖关系,结构化状态沿这些边传递并在关键边界持久化[^ch1-graph-engineering]。图工程不是循环工程的替代品,也不应简单地视为上述演进中的“第六层”。循环本身就是带有回边的图,图中的节点仍然可以内部运行ReAct或其他代理循环。名称尚未稳定,所以本书将其视为现有编排和框架实践的新兴术语;第10章发展多代理部分。这里的“图”指控制流或执行图,而不是GraphRAG使用的知识图。\n\n[^ch1-graph-engineering]: Josh C. Simmons在2026年7月4日的文章《我们正在进入图工程阶段》中明确使用了这个名称,用节点、类型化边和检查点状态来总结。7月18日,Peter Steinberger关于讨论是否从循环转向图的问题进一步推动了该名称的传播。这些实践早于标签出现:LangGraph、微软代理框架和谷歌ADK的官方文档将它们描述为图编排或基于图的工作流。见https://www.drjoshcsimmons.com/writing/we-are-entering-the-graph-engineering-phase,https://x.com/steipete/status/2078277297791189132,https://docs.langchain.com/oss/python/langgraph/overview,https://learn.microsoft.com/en-us/agent-framework/workflows/,以及https://adk.dev/workflows/。\n\n这五个阶段不是替代品,而是嵌套的层:提示工程是上下文工程的子集,上下文工程是框架工程的子集,框架工程是循环工程的子集。每一层都扩大了工程师的关注范围和影响力。**随着模型在能力上趋于一致,不再是决定性的差异化因素,竞争优势转移到模型之外的工程上。** 最近的工程实践支持这一观点。LangChain在Terminal Bench 2.0(评估代理在终端环境中完成复杂任务能力的基准)上的工作是一个显著的例子:他们的编码代理从52.8%提高到66.5%(从排行榜前30名之外跃升至前5名)。改变的不是模型,而是框架——让代理检查自己的执行结果,检测何时陷入重复循环,并完善其推理策略。OpenAI的工程团队分享了类似的经验:3名工程师在5个月内完成了约100万行代码和约1500个PR,约是传统开发速度的10倍。主要驱动力不是更强的模型;而是把框架搞对了。\n\n### 框架五个功能的核心原则\n前面的表格列出了框架的五个功能。下表添加了每个功能的核心设计原则以及本书对其的处理位置,将概念映射到实践:\n\n| 功能 | 核心原则 | 实践示例 | 见第章 |\n|----------|----------------------------------|------------------------------|-------|\n| **上下文** | 信息充足:确保代理在每个决策点基于充足信息做决策 | 系统提示、知识库、代理状态栏、Sidecar旁路查询 | 第2章和第3章 |\n| **工具** | 接口清晰:工具名称直观、参数有示例、边界有说明 | MCP工具、代码解释器、搜索工具 | 第4章 |\n| **约束** | 故障安全默认:所有能力默认关闭,必须显式启用(类似移动应用权限管理) | 在Claude Code中,每个工具默认执行前需要用户授权 | 第4章 |\n| **验证** | 输入隔离:安全检查仅查看结构化数据(例如工具返回的JSON字段),不查看模型生成的自由形式文本(因为攻击者可能通过提示注入操纵模型输出) | 林特检查、类型系统、工具调用结果验证 | 第5章和第6章 |\n| **纠正** | 直到确认故障不可恢复才暴露中间状态(例如静默重试失败的工具调用,而不是向用户显示未完成的结果) | 静默重试、继续生成、连续失败时移交人类判断(断路器机制) | 第2章和第5章 |\n\n五个功能形成一个闭环:上下文和工具支持决策,约束防止错误,验证检测偏差,纠正闭合循环。如果任何环节缺失,系统就会出现可靠性缺口。在审视具体的编排模式和护栏设计之前,我们首先列出构建有效代理和选择模型的核心原则——这是后续每个设计决策的基础。\n\n### 构建有效代理的核心原则\n基于Anthropic的经验,成功的代理系统遵循三个核心原则。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第6/9部分]\n\n一个具体的例子展示了框架的价值。假设你让一个代理退还用户3天前下的订单。**没有框架**:模型没有收到退款政策(没有上下文),不知道调用哪个API(没有工具),为用户编造退款结果(没有验证),用户发现退款从未发生(没有纠正)。**有框架**:系统提示指定了7天退款政策(上下文),代理调用`query_order`和`process_refund`工具执行操作(工具),框架检查退款不超过订单总额(约束),根据数据库确认退款已完成(验证),如果API调用超时自动重试(纠正)。同样的模型,结果大不相同。\n\n简而言之,没有框架的模型可能功能强大,但缺乏可靠完成任务所需的周围控制。\n\n更准确地说,模型之外的所有基础设施都属于框架。框架的核心是上下文和工具,围绕它们构建了三种工程保障:\n\n| 功能 | 一句话职责 | 与上下文/工具的关系 |\n|------------|--------------------------------|------------------------------|\n| **上下文** | 为模型提供相关信息 | 核心能力 |\n| **工具** | 为模型提供行动接口 | 核心能力 |\n| **约束** | 设置行为边界——能做和不能做的事 | 围绕上下文和工具构建的安全边界 |\n| **验证** | 自动判断工具执行结果的正确性 | 围绕工具执行结果构建的检查机制 |\n| **纠正** | 发现问题时自动恢复或回滚 | 围绕工具调用失败构建的恢复机制 |\n\n上下文和工具让代理完成任务——理解任务并采取行动。约束、验证和纠正确保其可靠且安全地完成任务——不是与上下文和工具分离的东西,而是确保它们在生产中可靠运行的工程。随着代理产品的成熟曲线,这两组之间的重点发生转移。\n\n早期的代理框架侧重于上下文和工具:给模型工具,给它上下文,让它完成任务。生产级系统已将重心转移到约束、验证和纠正上:确保工具调用安全、上下文得到管理、错误可恢复。\n\n以Claude Code为例。它的框架代码绝大多数都在进行约束、验证和纠正,而不是上下文和工具——工具本身(文件读写、命令执行、搜索)只是很小的一部分;围绕它们构建的保障措施才是真正的核心。这些机制包括:\n\n- **进程状态管理**:跟踪代理当前执行的步骤\n- **多层上下文压缩**:信息过多时自动修剪\n- **权限分类**:控制哪些操作需要用户确认\n- **断路器**:重复错误后自动停止重试,防止一个失败操作波及整个系统\n- **错误恢复机制**:捕获异常,回滚到最后稳定状态,重试或移交人工处理\n\n**行业正在从完成任务转向可靠完成任务,这使得框架工程成为代理系统的核心竞争优势。**\n\n### 从提示工程到循环工程:工程范式的演变\n\n回顾人工智能应用工程的发展,出现了一条清晰的演进弧线:\n\n**软件工程**是基础——传统的系统设计、架构、测试和部署。**提示工程**是第一波创新——通过完善喂给模型的自然语言指令来提高输出质量。**上下文工程**是第二波——意识到仅优化提示词是不够的:必须系统地管理模型的工作上下文(系统指令、工具定义、对话历史、外部知识)。**框架工程**是第三波——将视角从“模型接收什么信息”拓宽到“模型运行在什么样的系统中”,纳入模型之外的所有基础设施:约束机制、验证方法、反馈循环、错误恢复。接下来是**循环工程**,将视角从单次运行拓宽到跨运行的持续自主操作:谁发现下一个工作,何时验证,何时任务才算真正完成(第10章与多代理协作系统一起展开)。\n\n2026年7月,行业开始使用**图工程**从更高层次的编排视角:将代理循环、确定性程序和人工审批组织成显式的执行图,其中节点提供能力,边定义路由和依赖关系,结构化状态沿边传递并在关键边界持久化。[^ch1-graph-engineering] 图工程不是循环工程的替代品,也不应简单地视为上述演进中的“第六层”。循环本身就是带有回边的图,图中的节点仍可以内部运行ReAct或其他代理循环。名称尚未稳定,所以本书将其视为现有编排和框架实践的新兴术语;第10章展开多代理部分。这里的“图”指控制流或执行图,不是GraphRAG使用的知识图。\n\n[^ch1-graph-engineering]: Josh C. Simmons在2026年7月4日的文章《我们正在进入图工程阶段》中明确使用了这个名称,从节点、类型化边和检查点状态的角度进行了总结。7月18日,Peter Steinberger关于讨论是否从循环转向图的问题进一步推动了该名称的传播。这些实践早于该标签:LangGraph、微软代理框架和谷歌ADK的官方文档将它们描述为图编排或基于图的工作流。参见https://www.drjoshcsimmons.com/writing/we-are-entering-the-graph-engineering-phase,https://x.com/steipete/status/2078277297791189132,https://docs.langchain.com/oss/python/langgraph/overview,https://learn.microsoft.com/en-us/agent-framework/workflows/,以及https://adk.dev/workflows/。\n\n这五个阶段不是替代品而是嵌套的层次:提示工程是上下文工程的子集,上下文工程是框架工程的子集,框架工程是循环工程的子集。每一层都扩大了工程师的关注范围和影响力。**随着模型在能力上趋于一致,不再是决定性的差异化因素,竞争优势转移到模型之外的工程上。** 最近的工程实践支持这一观点。LangChain在Terminal Bench 2.0(评估代理在终端环境中完成复杂任务能力的基准)上的工作就是一个显著的例子:他们的编码代理从52.8%提高到66.5%(从排行榜前30名之外跃升至前5名)。改变的不是模型,而是框架——让代理检查自己的执行结果,检测何时陷入重复循环,并完善其推理策略。OpenAI的工程团队也分享了类似的经验:3名工程师在5个月内完成了大约100万行代码和近1500个拉取请求,约为传统开发速度的10倍。主要驱动力不是更强的模型;而是正确构建了框架。\n\n### 框架五个功能的核心原则\n\n前面的表格列出了框架的五个功能。下面的表格添加了每个功能的核心设计原则以及本书对其的处理方式,将概念映射到实践:\n\n| 功能 | 核心原则 | 实践示例 | 参见章节 |\n|------------|------------------------------------------|--------------------------------------|----------|\n| **上下文** | 信息充分性:确保代理在每个决策点都基于充分的信息做出决策 | 系统提示、知识库、代理状态栏、辅助程序绕过查询 | 第2章和第3章 |\n| **工具** | 清晰接口:工具名称直观,参数有示例,边界有说明 | MCP工具、代码解释器、搜索工具 | 第4章 | \n| **约束** | 故障安全默认值:所有能力默认关闭,必须明确启用(类似于移动应用权限管理) | 在Claude Code中,每个工具默认在执行前需要用户授权 | 第4章 |\n| **验证** | 输入隔离:安全检查只看结构化数据(例如工具返回的JSON字段),不看模型生成的自由文本(因为攻击者可能通过提示注入操纵模型输出) | 代码检查器、类型系统、工具调用结果验证 | 第5章和第6章 |\n| **纠正** | 在确认故障不可恢复之前不暴露中间状态(例如静默重试失败的工具调用,而不是向用户显示半完成的结果) | 静默重试、续生生成、连续失败时移交人工判断(断路器机制) | 第2章和第5章 |\n\n五个功能形成一个闭环:上下文和工具支持决策,约束防止错误,验证检测偏差,纠正闭合循环。如果任何环节缺失,系统就会出现可靠性缺口。在研究具体的编排模式和护栏设计之前,我们首先列出构建有效代理和选择模型的核心原则——这是后续所有设计决策的基础。\n\n### 构建有效代理的核心原则\n\n基于Anthropic的经验,成功的代理系统遵循三个核心原则。" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853905933903ea4fa82759f0cae4f2ffa723d4f9dd6361a2d", "model": "doubao-seed-1-6-250615", "created": 1785390623, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Minor omission: In the table under 'Core Principles of the Five Harness Functions', X translates 'Sidecar bypass queries' as 'Sidecar旁路查询' (retaining 'Sidecar'), while Y translates it as '辅助程序绕过查询' (explaining 'Sidecar' as '辅助程序'). The original uses 'Sidecar' as a technical term, so X is accurate here, but Y's explanation is not an error. However, X has no critical omissions/inventions, hence 4.\"},\"fluency\":{\"score\":4,\"evidence\":\"Most sentences flow naturally, e.g., '上下文和工具让代理完成任务——理解任务并采取行动' (Context and Tools let the Agent complete tasks—understand the task and act on it). Occasional awkward phrasing: '不是作为与上下文和工具分离的东西' (not as something apart from Context and Tools) is slightly rigid compared to Y's '不是与上下文和工具分离的东西' (same meaning but smoother in Y).\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of key terms: 'Harness'译为'框架' (Framework) consistently; 'Constrain'译为'约束', 'Verify'译为'验证', 'Correct'译为'纠正' throughout. Technical terms like '断路器' (Circuit Breaker), 'MCP tools'保留原样 are accurate and consistent.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Tables, headings (e.g., '### 从提示工程到循环工程:工程范式的演进'), code blocks (e.g., `query_order`), footnotes ([^ch1-graph-engineering]), and links are preserved exactly as in the source. No formatting errors in tables or markdown structure.\"}},\"Y\":{\"accuracy\":{\"score\":4,\"evidence\":\"Minor addition: In 'OpenAI's engineering team has shared a similar experience', Y adds '也' (also) in 'OpenAI的工程团队也分享了类似的经验', which is not in the original but does not change meaning. No critical inaccuracies, hence 4.\"},\"fluency\":{\"score\":5,\"evidence\":\"Superior flow in complex sentences: '生产级系统已将重心转移到约束、验证和纠正上:确保工具调用安全、上下文得到管理、错误可恢复' (Production-grade systems have shifted their center of gravity to Constrain, Verify, and Correct: making sure tool calls are safe, context is managed, and errors are recoverable) is more idiomatic than X's '生产级系统已将重心转移到约束、验证和纠正:确保工具调用安全、上下文得到管理、错误可恢复' (missing '上' for natural flow).\"},\"terminology\":{\"score\":4,\"evidence\":\"Mostly consistent, but 'Sidecar bypass queries'译为'辅助程序绕过查询' (explaining 'Sidecar' as '辅助程序') deviates from the original's technical term 'Sidecar'. 'PRs'译为'拉取请求' (Pull Requests) is accurate, but less consistent than X's retention of 'PR' in some contexts (though X also uses 'PRs' as 'PR').\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown elements fully preserved: tables with correct alignment, headings, code fences (e.g., `process_refund`), footnotes, and links. The table under 'Core Principles' has identical structure to the source, with no missing rows/columns.\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y demonstrates superior fluency with more natural phrasing (e.g., adding '上' in '重心转移到...上' for idiomatic flow) and accurate handling of technical terms with appropriate explanations (e.g., 'PRs'译为'拉取请求' for clarity). While X is strong in terminology consistency, Y's fluency edge makes it more readable without sacrificing accuracy.\"}", "usage": { "prompt_tokens": 6096, "completion_tokens": 830, "total_tokens": 6926 } }, "latency_milliseconds": 30423, "validation": { "valid": true } } ] }