Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
43 lines
30 KiB
JSON
43 lines
30 KiB
JSON
{
|
||
"schema_version": 1,
|
||
"credential_free": true,
|
||
"attempts": [
|
||
{
|
||
"attempt": 1,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n**Keep it simple.** Start with the simplest solution and add complexity only when truly necessary. Direct API calls are preferable to complex frameworks; clear code is preferable to clever abstraction—every extra layer of abstraction is a new blind spot during debugging.\n\n**Keep it transparent.** Show the Agent's planning steps, execution logs, and decision trajectory clearly. This is not just a debugging convenience; it is a precondition for user trust—an error inside a black box is hard to locate or fix from outside.\n\n**Design a well-structured tool interface (ACI, Agent-Computer Interface).** ACI means designing the interface from the Agent's perspective—easy for the Agent to understand and use—rather than from the programmer's perspective, as in traditional APIs. Tool names and parameters should be intuitive, and wherever misuse is likely the design should make the mistake impossible from the start: a SIM card's notched corner lets it slide into the tray in only one orientation, and a microwave refuses to heat while its door is open. Manufacturing calls this \"designing errors out\" philosophy **Poka-yoke**, a term from the Toyota Production System. A poorly designed tool can cause even the strongest model to fail repeatedly: the interface is the only channel between model and tool, and a vague interface gets amplified into systemic error.\n\nThe next three sections address three freestanding but important topics in Harness engineering: model selection, orchestration patterns, and guardrails and safety. None belongs to the five Harness elements proper, but all are unavoidable in engineering practice.\n\n### How to Choose a Model\n\nBefore discussing orchestration patterns, we first need to answer a practical question: what kind of model should drive your Agent?\n\nThe model is the foundation of the Agent's intelligence, and choosing the right one often matters more than any amount of prompt tuning. Model releases move too quickly for specific version recommendations to stay useful, so this section offers directions instead.\n\n**Know the \"Big Three.\"** The three most commonly used closed-source model providers in current Agent development are OpenAI (GPT/o series), Anthropic (Claude series), and Google (Gemini series). Each has its strengths: Claude excels in complex reasoning, coding, and tool calling, making it a popular choice for Agent development; Gemini offers an ultra-long context window and powerful multimodal capabilities, making it suitable for long texts and multimedia scenarios such as images and videos; the GPT/o series offers broadly balanced capabilities and has the largest user base. When selecting a model, do not rely only on leaderboards; **evaluate it on your own tasks** (see Chapter 6).\n\n**Chinese Models.** If your application is deployed in China or you are on a tight budget, models from Chinese vendors are a pragmatic choice. ByteDance's Doubao series offers extremely low latency within China, suitable for real-time interaction; Moonshot AI's Kimi is among the stronger Chinese models for Agent capabilities; open-source models like Qwen and DeepSeek have advantages in cost and customizability. Note that models differ widely in tool-calling ability, so be sure to test in your specific scenario before committing. Chinese models are typically accessed via APIs from platforms like Volcano Engine (Doubao) and SiliconFlow (open-source models), while non-Chinese models can be accessed through aggregator services such as OpenRouter.\n\n**Open Source vs. Closed Source.** Closed-source models generally lead in capability but are more expensive and constrained by the vendor's API policies. Open-source models are low-cost, support private deployment, and allow fine-tuning customization, making them suitable for cost-sensitive scenarios or those with data compliance requirements.\n\n**Most Agents Need a Model that Supports Reasoning.** Agents make complex decisions—multi-step reasoning, tool selection—and models without reasoning tend to perform poorly on them. The exceptions are few: a single simple step, or Computer Use GUI operations that amount to clicking a fixed position, where a non-reasoning model may suffice. The moment multi-step reasoning or dynamic decision-making enters, a reasoning model is essential.\n\n**Consider Output Speed and Multimodal Capabilities.** Beyond cost, two dimensions are easy to overlook. One is **output token speed**: Agents typically run many rounds of inference, and each round must finish before the next can start, so output speed directly determines end-to-end latency—a 20-round Agent task that runs 2 seconds slower per round means an extra 40 seconds of waiting. The other is **multimodal support**: if your Agent needs to understand images, audio, or video, multimodal capability is a hard requirement, and models differ widely here.\n\n### Orchestration Patterns: Workflow vs. Autonomous\n\nOrchestration patterns are how the Harness organizes its \"context and tools\" layer—they determine how context flows between LLM calls, how tools are scheduled, and whether the Agent's execution path is fixed in advance or generated dynamically. Agent orchestration has evolved from simple to complex, and each pattern has suitable use cases and trade-offs. In Anthropic's experience working with dozens of teams building LLM Agents, the most successful implementations rarely use complex frameworks; they use simple, composable patterns.\n\nWhen building an LLM application, progress from simple to complex. Start with a single LLM call—if better prompts and in-context examples solve the problem, do not build an Agent system. When multiple steps are needed and the task decomposes cleanly into fixed sub-tasks, use a workflow. Use an autonomous Agent only when you need dynamic decisions and a flexible execution path. And remember: Agent systems typically trade latency and cost for better task performance—evaluate carefully whether that trade is worth it.\n\n#### Workflow Pattern: Deterministic Orchestration\n\nA **workflow** is a system that orchestrates LLMs and tools through predefined code paths. Its execution path is deterministic and designed in advance by the developer—the behavior of each step and transition is defined in code; the LLM handles only the understanding and generation inside each node.\n\nFor example, a flight-booking Agent can use a workflow with four fixed nodes:\n\n1. **Verify User Identity**—Call the identity verification API to confirm who the user is.\n2. **Search for Available Flights**—Query the flight database based on user requirements.\n3. **Complete Payment**—Call the payment interface to deduct the amount.\n4. **Confirm Booking**—Call the booking API to lock the seat and send a confirmation to the user.\n\nAn LLM can be used within each node (e.g., using natural language to understand the user's travel needs), but the flow sequence between nodes is fixed by code—the system will not book a seat before payment is completed, nor will it start searching for flights before identity verification.\n\nThe workflow pattern has two core advantages. First, **strict process control**: the developer can guarantee that critical steps are never skipped or run out of order—business rules like \"no booking before payment\" are enforced by code, not left to the LLM's judgment. Second, **security**: because the execution path is deterministic, prompt injection or a model error can at most affect the processing inside the current node; it cannot make the Agent jump to a branch it should not reach. The attack surface is confined to a single node.\n\nThe main limitation of a workflow is its **lack of flexibility**. When an unanticipated event occurs—for example, the user changes the booking during payment, or a flight is canceled and the system needs to recommend an alternative—the fixed path cannot adapt on its own; it can only follow a preset exception branch or hand control back to a human.\n\n#### Autonomous Agent: Runtime Decision-Making\n\nWhen the fixed path of a workflow is insufficient, we need an **autonomous Agent**. The core difference between an autonomous Agent and a workflow is that the execution path is not predefined but is determined at runtime by the Agent based on **environmental feedback**.\n\nReturning to the flight example, an autonomous Agent needs no four predefined nodes. The user says, \"Book me a flight to Shanghai next Wednesday,\" and the Agent determines the sequence dynamically: it searches for flights, discovers that login is required, verifies identity, and resumes the search. If the cheapest flight has a layover, it can ask whether that is acceptable; if the user says no, it adjusts the search criteria.\n\nAn autonomous Agent therefore has to plan for itself—choose its own execution steps—and recognize failure and change strategy rather than simply halting on error. But autonomy is not unbounded: explicit **stopping conditions** must be designed in (task complete, maximum iterations reached, unrecoverable error hit), or the Agent can enter infinite loops or continue executing after the task is already done.\n\nFrom an implementation perspective, an autonomous Agent is essentially an LLM using tools in a loop, continuously obtaining environmental feedback to make progress on the task—this is the ReAct loop introduced earlier. Common exit conditions include: calling a final output tool, the model returning a response without any tool calls, or encountering an error or reaching the maximum number of rounds.\n\n\n\n\n\nANONYMOUS CHINESE X:\n### 入门指南:AI 代理(第 7/9 部分)\n\n**保持简单**。从最简单的解决方案开始,仅在真正必要时增加复杂性。直接的 API 调用比复杂的框架更可取;清晰的代码比巧妙的抽象更可取——每一层额外的抽象在调试时都是新的盲点。\n\n**保持透明**。清晰展示代理的规划步骤、执行日志和决策轨迹。这不仅是调试的便利;也是用户信任的前提——黑盒内的错误很难从外部定位或修复。\n\n**设计结构良好的工具接口(ACI,代理-计算机接口)**。ACI 是从代理的角度设计接口——让代理易于理解和使用——而不是像传统 API 那样从程序员的角度设计。工具名称和参数应直观,并且在可能出现误用的地方,设计应从一开始就避免错误:SIM 卡的缺口角使其只能以一种方向滑入托盘中,微波炉门打开时无法加热。制造业将此称为“消除错误”理念 **防错法(Poka-yoke)**,这是丰田生产系统中的一个术语。设计不佳的工具甚至会导致最强的模型反复失败:接口是模型和工具之间的唯一通道,模糊的接口会被放大为系统性错误。\n\n接下来的三个部分讨论了框架工程中三个独立但重要的主题:模型选择、编排模式以及防护措施和安全性。这些都不属于框架的五个适当元素,但在工程实践中都是不可避免的。\n\n### 如何选择模型\n\n在讨论编排模式之前,我们首先需要回答一个实际问题:什么样的模型应该驱动你的代理?\n\n模型是代理智能的基础,选择合适的模型往往比任何数量的提示调整都重要。模型发布更新太快,特定版本的推荐很难保持有用,所以本节提供方向而非具体推荐。\n\n**了解“三大巨头”**。当前代理开发中最常用的三个闭源模型提供商是 OpenAI(GPT/o 系列)、Anthropic(Claude 系列)和 Google(Gemini 系列)。每个都有其优势:Claude 在复杂推理、编码和工具调用方面表现出色,是代理开发的热门选择;Gemini 提供超长上下文窗口和强大的多模态能力,适合长文本和图像、视频等多媒体场景;GPT/o 系列能力均衡且用户基数最大。选择模型时,不要仅依赖排行榜;**在自己的任务上进行评估**(见第 6 章)。\n\n**中文模型**。如果你的应用部署在中国或预算有限,中国供应商的模型是务实之选。字节跳动的豆包系列在中国内延迟极低,适合实时交互;摩斯智算的 Kimi 在代理能力方面是较强的中文模型之一;通义千问、深度求索等开源模型在成本和可定制性方面有优势。请注意,模型在工具调用能力上差异很大,所以在投入使用前一定要在具体场景中测试。中文模型通常通过火山引擎(豆包)、硅基智能(开源模型)等平台的 API 访问,而非中文模型可以通过 OpenRouter 等聚合服务访问。\n\n**开源与闭源**。闭源模型通常能力领先,但成本更高且受供应商 API 政策限制。开源模型成本低,支持私有部署,允许微调定制,适合成本敏感场景或有数据合规要求的场景。\n\n**大多数代理需要支持推理的模型**。代理要做出复杂决策——多步推理、工具选择等,没有推理能力的模型在这些方面往往表现不佳。例外情况很少:单一简单步骤,或相当于点击固定位置的计算机使用 GUI 操作,此时非推理模型可能够用。一旦涉及多步推理或动态决策,推理模型就至关重要。\n\n**考虑输出速度和多模态能力**。除了成本,还有两个容易忽视的维度。一是**输出词元速度**:代理通常要运行多轮推理,每一轮必须在前一轮完成后才能开始,所以输出速度直接决定端到端时延——一个 20 轮的代理任务,每轮慢 2 秒,就会多等 40 秒。二是**多模态支持**:如果你的代理需要理解图像、音频或视频,多模态能力是硬性要求,而模型在这方面差异很大。\n\n### 编排模式:工作流与自主式\n\n编排模式是框架组织其“上下文和工具”层的方式——它们决定上下文在大语言模型调用之间如何流动,工具如何调度,以及代理的执行路径是预先固定还是动态生成。代理编排从简单到复杂不断演变,每种模式都有合适的用例和权衡。根据 Anthropic 与数十个构建大语言模型代理的团队合作经验,最成功的实现很少使用复杂框架;它们使用简单、可组合的模式。\n\n构建大语言模型应用时,要从简单到复杂推进。从单个大语言模型调用开始——如果更好的提示和上下文示例能解决问题,就不要构建代理系统。当需要多个步骤且任务能清晰分解为固定子任务时,使用工作流。只有当需要动态决策和灵活执行路径时,才使用自主式代理。并且记住:代理系统通常以时延和成本换取更好的任务性能——要仔细评估这种交换是否值得。\n\n#### 工作流模式:确定性编排\n\n**工作流**是通过预定义代码路径编排大语言模型和工具的系统。其执行路径是确定性的,由开发者预先设计——每个步骤和转换的行为都在代码中定义;大语言模型仅处理每个节点内的理解和生成。\n\n例如,一个航班预订代理可以使用包含四个固定节点的工作流:\n\n1. **验证用户身份**——调用身份验证 API 确认用户身份。\n2. **搜索可用航班**——根据用户需求查询航班数据库。\n3. **完成支付**——调用支付接口扣款。\n4. **确认预订**——调用预订 API 锁定座位并向用户发送确认。\n\n每个节点内都可以使用大语言模型(例如用自然语言理解用户的出行需求),但节点之间的流程顺序由代码固定——系统不会在支付完成前预订座位,也不会在身份验证前开始搜索航班。\n\n工作流模式有两个核心优势。首先,**严格的流程控制**:开发者可以保证关键步骤绝不会被跳过或顺序错误——“支付前不预订”等业务规则由代码强制执行,而不是交由大语言模型判断。其次,**安全性**:因为执行路径是确定性的,提示注入或模型错误最多影响当前节点内的处理;不会让代理跳转到不应到达的分支。攻击面局限在单个节点内。\n\n工作流的主要局限是**缺乏灵活性**。当出现意外事件时——例如用户在支付时更改预订,或航班取消系统需要推荐替代方案——固定路径无法自行适应;只能遵循预设的异常分支或将控制权交回给人类。\n\n#### 自主式代理:运行时决策\n\n当工作流的固定路径不足时,我们需要**自主式代理**。自主式代理与工作流的核心区别在于,执行路径不是预先定义的,而是由代理在运行时根据**环境反馈**确定的。\n\n回到航班示例,自主式代理不需要四个预定义节点。用户说“给我订下周三去上海的航班”,代理动态确定顺序:搜索航班,发现需要登录,验证身份,然后继续搜索。如果最便宜的航班有经停,它可以询问是否可以接受;如果用户说不行,它就调整搜索标准。\n\n因此,自主式代理必须自己规划——选择自己的执行步骤——并识别失败并改变策略,而不是简单地在错误时停止。但自主性不是无边界的:必须设计明确的**停止条件**(任务完成、达到最大迭代次数、遇到不可恢复错误),否则代理可能进入无限循环或在任务已完成后继续执行。\n\n从实现角度看,自主式代理本质上是在循环中使用工具的大语言模型,不断获取环境反馈以推进任务——这就是前面介绍的 ReAct 循环。常见的退出条件包括:调用最终输出工具、模型返回没有任何工具调用的响应,或遇到错误或达到最大轮次。\n\n\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门[第7/9部分]\n**保持简单**。从最简单的解决方案开始,只有在真正必要时才增加复杂性。直接的API调用优于复杂的框架;清晰的代码优于巧妙的抽象——每一层额外的抽象都是调试时的新盲点。\n\n**保持透明**。清晰展示代理的规划步骤、执行日志和决策轨迹。这不仅是调试的便利,更是用户信任的前提——黑盒内的错误很难从外部定位或修复。\n\n**设计良好结构的工具接口(ACI,代理-计算机接口)**。ACI意味着从代理的角度设计接口——让代理易于理解和使用——而不是像传统API那样从程序员的角度设计。工具名称和参数应直观,在可能误用的地方,设计应从一开始就杜绝错误:SIM卡的缺口角使其只能以一种方向滑入托槽,微波炉门打开时拒绝加热。制造业将这种“消除错误”的设计理念称为**防错法(Poka-yoke)**,这是丰田生产系统中的一个术语。设计不佳的工具甚至会导致最强的模型反复失败:接口是模型和工具之间的唯一通道,模糊的接口会被放大为系统错误。\n\n接下来的三个部分讨论框架工程中三个独立但重要的主题:模型选择、编排模式以及护栏和安全。它们不属于五个框架要素本身,但在工程实践中都不可避免。\n\n### 如何选择模型\n在讨论编排模式之前,我们首先需要回答一个实际问题:你的代理应该由什么样的模型驱动?\n\n模型是代理智能的基础,选择合适的模型往往比任何数量的提示调整都重要。模型发布速度太快,特定版本的推荐难以保持有用,所以本节提供方向。\n\n**了解“三大”**。当前代理开发中最常用的三个闭源模型提供商是OpenAI(GPT/o系列)、Anthropic(Claude系列)和谷歌(Gemini系列)。每个都有其优势:Claude在复杂推理、编码和工具调用方面表现出色,是代理开发的热门选择;Gemini提供超长上下文窗口和强大的多模态能力,适合长文本和图像、视频等多媒体场景;GPT/o系列能力广泛平衡且用户基数最大。选择模型时,不要仅依赖排行榜;**在自己的任务上评估它**(见第6章)。\n\n**中文模型**。如果你的应用部署在中国或预算紧张,中国供应商的模型是务实的选择。字节跳动的豆包系列在中国内具有极低延迟,适合实时交互;月之暗面科技的Kimi在代理能力方面是较强的中国模型之一;Qwen和DeepSeek等开源模型在成本和可定制性方面有优势。注意模型在工具调用能力上差异很大,所以在承诺使用前一定要在具体场景中测试。中文模型通常通过火山引擎(豆包)和硅基流动(开源模型)等平台的API访问,而非中文模型可以通过OpenRouter等聚合服务访问。\n\n**开源与闭源**。闭源模型通常在能力上领先,但成本更高且受供应商API政策限制。开源模型成本低,支持私有部署,允许微调定制,适合成本敏感场景或有数据合规要求的场景。\n\n**大多数代理需要支持推理的模型**。代理进行复杂决策——多步骤推理、工具选择——没有推理能力的模型在这些方面往往表现不佳。例外很少:单个简单步骤,或相当于点击固定位置的计算机使用GUI操作,此时非推理模型可能足够。一旦涉及多步骤推理或动态决策,推理模型就至关重要。\n\n**考虑输出速度和多模态能力**。除了成本,有两个维度容易被忽视。一个是**输出标记速度**:代理通常运行多轮推理,每轮必须在前一轮完成后才能开始,所以输出速度直接决定端到端延迟——一个20轮的代理任务每轮慢2秒意味着额外等待40秒。另一个是**多模态支持**:如果你的代理需要理解图像、音频或视频,多模态能力是硬性要求,而模型在这方面差异很大。\n\n### 编排模式:工作流与自主\n编排模式是框架组织其“上下文和工具”层的方式——它们决定LLM调用之间上下文如何流动、工具如何调度,以及代理的执行路径是预先固定还是动态生成。代理编排从简单到复杂演进,每种模式都有合适的用例和权衡。根据Anthropic与数十个构建大语言模型代理的团队合作经验,最成功的实现很少使用复杂框架;它们使用简单、可组合的模式。\n\n构建大语言模型应用时,从简单到复杂推进。从单个LLM调用开始——如果更好的提示和上下文内示例解决了问题,就不要构建代理系统。当需要多个步骤且任务清晰分解为固定子任务时,使用工作流。仅当需要动态决策和灵活执行路径时,使用自主代理。并且记住:代理系统通常以延迟和成本换取更好的任务性能——仔细评估这种交换是否值得。\n\n#### 工作流模式:确定性编排\n**工作流**是通过预定义代码路径编排LLM和工具的系统。其执行路径是确定性的,由开发者预先设计——每个步骤和过渡的行为在代码中定义;LLM仅处理每个节点内的理解和生成。\n\n例如,一个航班预订代理可以使用具有四个固定节点的工作流:\n\n1. **验证用户身份**——调用身份验证API确认用户身份。\n2. **搜索可用航班**——根据用户需求查询航班数据库。\n3. **完成支付**——调用支付接口扣款。\n4. **确认预订**——调用预订API锁定座位并向用户发送确认。\n\n每个节点内可以使用LLM(例如用自然语言理解用户的旅行需求),但节点之间的流程顺序由代码固定——系统不会在支付完成前预订座位,也不会在身份验证前开始搜索航班。\n\n工作流模式有两个核心优势。首先,**严格流程控制**:开发者可以保证关键步骤永远不会被跳过或顺序错误——“支付前不预订”等业务规则由代码强制执行,而不是留给LLM判断。其次,**安全性**:因为执行路径是确定性的,提示注入或模型错误最多影响当前节点内的处理;它不会让代理跳转到不应到达的分支。攻击面局限于单个节点。\n\n工作流的主要限制是**缺乏灵活性**。当出现意外事件时——例如用户在支付期间更改预订,或航班取消系统需要推荐替代方案——固定路径无法自行适应;它只能遵循预设的异常分支或将控制权交还给人类。\n\n#### 自主代理:运行时决策\n当工作流的固定路径不足时,我们需要**自主代理**。自主代理与工作流的核心区别在于执行路径不是预先定义的,而是由代理根据**环境反馈**在运行时确定的。\n\n回到航班示例,自主代理不需要四个预定义节点。用户说“给我预订下周三去上海的航班”,代理动态确定顺序:它搜索航班,发现需要登录,验证身份,然后继续搜索。如果最便宜的航班有经停,它可以询问是否可接受;如果用户说不可接受,它调整搜索标准。\n\n因此,自主代理必须自行规划——选择自己的执行步骤——并识别失败和改变策略,而不是简单地在错误时停止。但自主性不是无界的:必须设计明确的**停止条件**(任务完成、达到最大迭代次数、遇到不可恢复错误),否则代理可能进入无限循环或在任务已完成后继续执行。\n\n从实现角度看,自主代理本质上是在循环中使用工具的LLM,不断获取环境反馈以推进任务——这就是前面介绍的ReAct循环。常见的退出条件包括:调用最终输出工具、模型返回没有任何工具调用的响应,或遇到错误或达到最大轮次。\n\n"
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "0217853906238163ea4fa82759f0cae4f2ffa723d4f9dd6c4d4c1",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390638,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions, inventions, or changed claims. For example, 'Poka-yoke' is correctly translated as '防错法(Poka-yoke)' with the original explanation preserved; 'output token speed' is accurately rendered as '输出词元速度' without altering meaning.\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural and idiomatic Chinese. Examples: '保持简单' (Keep it simple), '代理智能的基础' (the foundation of the Agent's intelligence) flow smoothly; complex sentences like '每一层额外的抽象在调试时都是新的盲点' (every extra layer of abstraction is a new blind spot during debugging) are well-structured.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical terms. 'Agent' is uniformly '代理', 'workflow' is '工作流', 'orchestration patterns' is '编排模式'; 'Poka-yoke' is consistently '防错法(Poka-yoke)'; 'ReAct loop' is 'ReAct 循环'.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"All Markdown elements preserved: headings (### 如何选择模型), bold text (**保持简单**), list items (1. **验证用户身份**), and image link  are correctly retained.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"No omissions or inventions. Key claims like 'evaluate it on your own tasks' are translated as '在自己的任务上评估它'; 'Chinese Models' section details are fully preserved.\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly less natural in some phrases. For example, '人工智能代理入门[第7/9部分]' uses brackets instead of the more standard '(第7/9部分)'; '护栏和安全' (guardrails and safety) is less idiomatic than X's '防护措施和安全性'.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent terminology: 'Agent' as '代理', 'closed-source' as '闭源', 'multimodal capabilities' as '多模态能力'; 'LLM' is retained as 'LLM' consistently.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown structure intact: headings, bold, lists, and image link  are correctly preserved.\"}}},\"preferred\":\"X\",\"preference_evidence\":\"X has slightly better fluency with more idiomatic expressions (e.g., '防护措施和安全性' vs. '护栏和安全') and consistent use of standard punctuation (parentheses for section numbers: '(第7/9部分)' vs. '[第7/9部分]'). Both are accurate, but X's fluency edge makes it preferable.\"}",
|
||
"usage": {
|
||
"prompt_tokens": 5683,
|
||
"completion_tokens": 630,
|
||
"total_tokens": 6313
|
||
}
|
||
},
|
||
"latency_milliseconds": 14693,
|
||
"validation": {
|
||
"valid": true
|
||
}
|
||
}
|
||
]
|
||
}
|