Files
ai-agent-book/chapter10/book-translation/validation/real_20260730T061500Z_v4/judge_receipts/unit-02.json
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

43 lines
31 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n| Agent Product | Working Context | Action Interfaces | Strategy |\n|-----------------|------------------------|--------------------------|-----------------------------|\n| **Coding Agents (e.g., Cursor)** | Requirements documents, codebase, terminal environment | Open-ended (internal reasoning, code search, file read/write, command execution, etc.) | Incremental development: understand requirements → search relevant code → edit code → test and verify → debug and fix |\n| **Search Agents (e.g., Deep Research)** | Web resources, academic databases, local files | Open-ended (internal reasoning, search queries, web reading, summary generation) | Iterative deepening: adjust search direction based on existing information, gradually synthesize a complete report |\n| **Computer Control Agents (e.g., Browser Use)** | Computer screen, browser pages, file system | Open-ended (internal reasoning, clicking, typing, scrolling, screenshots, code execution, etc.) | Visual perception + operation: observe screen → identify target elements → perform actions → verify results |\n| **Phone Assistant Agents (e.g., Doubao)** | Phone screen, installed apps | Open-ended (internal reasoning, clicking, swiping, typing, opening apps, etc.) | Intent understanding + App control: understand user needs → locate target app → perform actions → confirm completion |\n| **Personal Task Agents (e.g., Pine AI)** | User account information, historical bills, service provider knowledge base | Open-ended (internal reasoning, making calls, sending emails, filling forms, confirming with user) | Multi-step task execution: gather information → formulate negotiation strategy → contact service provider → negotiate → report results |\n\nThese systems share three features: an **open-ended action space**—not picking from a fixed set of buttons but generating arbitrary natural language and code; **internal reasoning**—planning before acting; and **continuous interaction**—adjusting strategy based on environmental feedback. These capabilities come precisely from the interplay of the reasoning engine, working context, and action interfaces—that is, LLM, context, and tools.\n\n### Tools: The Agent's Action Interfaces\n\nTools are the Agent's bridge to the outside world. They turn the Agent from a passive observer into an active system that can search, write files, run code, call APIs, send messages, or operate interfaces. Without tools, an Agent is limited to text generation; with them, it can act on external systems.\n\nTo discuss tools systematically, we can sort them into five types by the direction of the Agent's interaction with the world. At this stage, a brief overview of each type's representative scenarios is enough to establish the overall picture; later chapters treat each in depth.\n\n**Perception Tools** allow the Agent to access information: search engines provide real-time web data, file systems read local documents, and APIs and databases connect to external services and enterprise core data.\n\n**Execution Tools** allow the Agent to act on external systems: code execution, file operations, system commands, and external API calls turn decisions into concrete actions.\n\n**Collaboration Tools** allow the Agent to divide work with other Agents: delegating specialized tasks to sub-agents, requesting human confirmation at key decision points, or coordinating actions in multi-agent systems.\n\n**Event Trigger Tools** are invoked in a fundamentally different way from the first three categories: the Agent does not call them; they arrive as external inputs that trigger the Agent to begin work. A new email comes in, a scheduled time arrives, or another system fires a Webhook callback; the event activates the Agent and initiates reasoning and action. The Agent never calls these itself, yet they are still a channel through which it interacts with the outside world, so we count them in the broad tool system.\n\n**User Communication Tools** are the channels through which the Agent communicates with the user. Where execution tools change the external world, communication tools carry information—delivering the Agent's progress, or a proactive check-in, by text message, voice call, email, and so on.\n\nChapter 4 covers the full taxonomy and design principles for these five types. The quality of tool design directly determines what an Agent can reliably accomplish: define interfaces vaguely and the model will misuse them; handle errors poorly and a single failed tool can leave the Agent stuck; scope permissions too broadly and one Agent error can become irreversible. As the MCP (Model Context Protocol) standard spreads, integrating a tool is becoming as easy as installing a plugin—the ecosystem is expanding rapidly, but the design principles will not go out of date.\n\n**Tool Calling** (also known as Function Calling) is a core capability of modern LLM Agents: it lets the model invoke external tools in a structured way, transforming the LLM from a pure text generator into an intelligent system that can act through external interfaces. This book uses the term \"tool calling\" throughout.\n\nTool calling proceeds in four steps: first, the context tells the model which tools are available (names, purposes, parameters); then the model decides on its own whether to call a tool, which tool to call, and with what arguments; next, once the tool has run, its result is appended to the context; finally, the model decides its next move based on that result. This loop is the foundation of ReAct, introduced later in the chapter.\n\nFor a weather query, the simplified representation of the four-step process at the API level is as follows:\n\n```\nStep 1: Declare tools Step 2: Model decides to call\ntools: [{ assistant: {\n name: \"get_weather\", tool_calls: [{\n parameters: { function: \"get_weather\",\n city: \"string\" arguments: {city: \"Beijing\"}\n } }]\n}] }\n\nStep 3: Result appended to context Step 4: Model responds based on result\ntool: { assistant: {\n tool_call_id: \"call_1\", content: \"Today in Beijing: 28°C, sunny.\"\n content: '{\"temp\":28,\"sky\":\"clear\"}' }\n} }\n```\n\nThe developer only defines the tools and executes the calls; the model itself decides whether to call, which tool to call, and what arguments to pass. Chapter 2 examines this API structure in detail.\n\nWhen designing tools for an Agent, start with the narrowest capability the task needs, then expand gradually as the task grows more complex. If the task only requires basic arithmetic, a calculator with clearly defined parameters is enough; when it grows to reading spreadsheets, cleaning missing values, computing statistics, and plotting charts, a constrained Python code interpreter is easier to combine and explore with than an ever-growing collection of specialized tools. But generality also increases the risk of errors and expands the attack surface: code must run in an isolated sandbox, with network access disabled by default, no access to files outside the authorized working directory, and limits on execution time, CPU, memory, and output size.\n\nLikewise, a single logging tool is suitable for recording one execution; for long-running tasks that take hours or even days, a controlled virtual working directory can preserve plans, intermediate results, execution logs, and final artifacts so the Agent can resume across multiple runs. This directory should also restrict readable and writable paths, storage capacity, and file types, and prevent path traversal instead of exposing the entire host file system to the Agent.\n\nGeneral-purpose tools are not always better than specialized ones. High-risk operations or those governed by strict business constraints—such as payments, data deletion, sending email, and production deployment—should still be exposed as dedicated tools with explicit parameters, restricted permissions, and end-to-end auditability, with previews and human confirmation added when necessary. The core principle of tool design is therefore: **use general-purpose foundational capabilities for composition and exploration; use specialized tools to constrain high-risk operations and enforce strict business rules**.\n\n### LLM: The Agent's Reasoning Engine\n\nThe Large Language Model (LLM) is the Agent's decision-making core. Given a user request, it first has to infer the real intent (what users say is often not what they actually want), then break a vague or complex task into executable steps. Throughout execution it keeps making decisions: what to do next, whether to call a tool, which one, and with what arguments. This understandplanexecute capability comes from knowledge accumulated during pre-training, and it is the foundation that workflows and autonomous Agents alike depend on.\n\nA distinctive capability of LLM Agents is **internal reasoning**—before acting, the Agent can plan and reason through the task. This does not change the external environment, yet it markedly improves the actions that follow. This ability comes from pre-training (the initial training on massive amounts of internet text, through which the model learns language patterns and world knowledge): the model draws on reasoning patterns encoded in human knowledge, including mathematical laws, causal relationships, and strategies for decomposing problems. An Agent's reasoning is therefore not blind trial and error; it builds on a structured body of knowledge.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第2/9部分]\n| 代理产品 | 工作上下文 | 行动接口 | 策略 |\n|------------------|--------------------------|-----------------------------------|--------------------------------------|\n| **编码代理(例如Cursor)** | 需求文档、代码库、终端环境 | 开放式(内部推理、代码搜索、文件读写、命令执行等) | 增量式开发:理解需求→搜索相关代码→编辑代码→测试验证→调试修复 |\n| **搜索代理(例如Deep Research)** | 网络资源、学术数据库、本地文件 | 开放式(内部推理、搜索查询、网页阅读、摘要生成) | 迭代深化:根据现有信息调整搜索方向,逐步合成完整报告 |\n| **计算机控制代理(例如浏览器使用)** | 计算机屏幕、浏览器页面、文件系统 | 开放式(内部推理、点击、输入、滚动、截图、代码执行等) | 视觉感知+操作:观察屏幕→识别目标元素→执行操作→验证结果 |\n| **手机助手代理(例如豆包)** | 手机屏幕、已安装应用 | 开放式(内部推理、点击、滑动、输入、打开应用等) | 意图理解+应用控制:理解用户需求→定位目标应用→执行操作→确认完成 |\n| **个人任务代理(例如Pine AI)** | 用户账户信息、历史账单、服务提供商知识库 | 开放式(内部推理、打电话、发邮件、填表、与用户确认等) | 多步骤任务执行:收集信息→制定协商策略→联系服务提供商→协商→报告结果 |\n\n这些系统有三个共同特征:**开放式行动空间**——不是从固定的按钮集合中选择,而是生成任意的自然语言和代码;**内部推理**——在行动前进行规划;以及**连续交互**——根据环境反馈调整策略。这些能力恰好来自推理引擎、工作上下文和行动接口的相互作用——也就是大语言模型、上下文和工具的相互作用。\n\n### 工具:代理的行动接口\n\n工具是代理与外部世界的桥梁。它们将代理从被动的观察者转变为能够搜索、写入文件、运行代码、调用API、发送消息或操作接口的主动系统。没有工具,代理仅限于文本生成;有了工具,它可以对外部系统采取行动。\n\n为了系统地讨论工具,我们可以根据代理与世界交互的方向将其分为五类。在这个阶段,简要概述每种类型的代表性场景就足以建立整体图景;后面的章节将深入探讨每种类型。\n\n**感知工具**允许代理访问信息:搜索引擎提供实时网络数据,文件系统读取本地文档,API和数据库连接外部服务和企业核心数据。\n\n**执行工具**允许代理对外部系统采取行动:代码执行、文件操作、系统命令和外部API调用将决策转化为具体行动。\n\n**协作工具**允许代理与其他代理分工:将专门任务委托给子代理,在关键决策点请求人类确认,或在多代理系统中协调行动。\n\n**事件触发工具**以与前三类根本不同的方式被调用:代理不调用它们;它们作为外部输入到达,触发代理开始工作。新邮件到来、预定时间到达或另一个系统触发Webhook回调;事件激活代理并启动推理和行动。代理本身从不调用这些工具,但它们仍然是代理与外部世界交互的通道,所以我们将它们计入广义的工具系统中。\n\n**用户通信工具**是代理与用户通信的通道。执行工具改变外部世界,而通信工具传递信息——通过短信、语音通话、电子邮件等传递代理的进度或主动询问。\n\n第4章将涵盖这五类的完整分类法和设计原则。工具设计的质量直接决定了代理能够可靠完成的任务:接口定义模糊,模型会误用它们;错误处理不佳,单个失败的工具可能会让代理陷入困境;权限范围过广,一个代理错误可能会不可逆转。随着MCP(模型上下文协议)标准的传播,集成工具变得像安装插件一样简单——生态系统正在迅速扩展,但设计原则不会过时。\n\n**工具调用**(也称为函数调用)是现代大语言模型代理的核心能力:它让模型以结构化的方式调用外部工具,将大语言模型从纯文本生成器转变为能够通过外部接口行动的智能系统。本书通篇使用“工具调用”这个术语。\n\n工具调用分为四个步骤:首先,上下文告诉模型哪些工具可用(名称、用途、参数);然后模型自行决定是否调用工具、调用哪个工具以及使用什么参数;接下来,工具运行后,其结果附加到上下文中;最后,模型根据该结果决定下一步行动。这个循环是后面章节介绍的ReAct的基础。\n\n以天气查询为例,API层面四步过程的简化表示如下:\n\n```\n步骤1:声明工具 步骤2:模型决定调用\ntools: [{ assistant: {\n name: \"get_weather\", tool_calls: [{\n parameters: { function: \"get_weather\",\n city: \"string\" arguments: {city: \"Beijing\"}\n } }]\n}] }\n\n步骤3:结果附加到上下文 步骤4:模型根据结果响应\ntool: { assistant: {\n tool_call_id: \"call_1\", content: \"Today in Beijing: 28°C, sunny.\"\n content: '{\"temp\":28,\"sky\":\"clear\"}' }\n} }\n```\n\n开发者只需定义工具并执行调用;模型本身决定是否调用、调用哪个工具以及传递什么参数。第2章将详细检查这个API结构。\n\n为代理设计工具时,从任务所需的最窄能力开始,然后随着任务变得更复杂逐步扩展。如果任务只需要基本算术,具有明确定义参数的计算器就足够了;当任务发展到读取电子表格、清理缺失值、计算统计数据和绘制图表时,受约束的Python代码解释器比不断增长的专门工具集合更容易组合和探索。但通用性也会增加错误风险并扩大攻击面:代码必须在隔离的沙盒中运行,默认禁用网络访问,无法访问授权工作目录之外的文件,并且对执行时间、CPU、内存和输出大小有限制。\n\n同样,单个日志工具适合记录一次执行;对于耗时数小时甚至数天的长时间运行任务,受控的虚拟工作目录可以保存计划、中间结果、执行日志和最终工件,以便代理可以在多次运行中恢复。这个目录还应该限制可读和可写路径、存储容量和文件类型,并防止路径遍历,而不是将整个主机文件系统暴露给代理。\n\n通用工具并不总是比专门工具更好。高风险操作或受严格业务约束的操作——例如支付、数据删除、发送电子邮件和生产部署——仍然应该作为具有明确参数、受限权限和端到端可审计性的专用工具暴露,必要时添加预览和人类确认。因此,工具设计的核心原则是:**使用通用基础能力进行组合和探索;使用专门工具约束高风险操作并强制执行严格业务规则**。\n\n### 大语言模型:代理的推理引擎\n\n大语言模型(LLM)是代理的决策核心。给定用户请求,它首先必须推断真实意图(用户所说的往往不是他们真正想要的),然后将模糊或复杂的任务分解为可执行步骤。在整个执行过程中,它不断做出决策:下一步做什么、是否调用工具、调用哪个工具以及使用什么参数。这种理解-规划-执行能力来自预训练期间积累的知识,是工作流和自主代理都依赖的基础。\n\n大语言模型代理的一个独特能力是**内部推理**——在行动前,代理可以规划和推理任务。这不会改变外部环境,但会显著改善后续行动。这种能力来自预训练(在大量互联网文本上的初始训练,通过它模型学习语言模式和世界知识):模型利用编码在人类知识中的推理模式,包括数学定律、因果关系和分解问题的策略。因此,代理的推理不是盲目试错;它建立在结构化的知识体系之上。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第2/9部分]\n\n| 代理产品 | 工作上下文 | 动作接口 | 策略 |\n|------------------------|--------------------------------|------------------------------|--------------------------------------------------------------|\n| **编码代理(例如Cursor)** | 需求文档、代码库、终端环境 | 开放式(内部推理、代码搜索、文件读写、命令执行等) | 增量式开发:理解需求→搜索相关代码→编辑代码→测试验证→调试修复 |\n| **搜索代理(例如Deep Research)** | 网络资源、学术数据库、本地文件 | 开放式(内部推理、搜索查询、网页阅读、摘要生成) | 迭代深化:根据现有信息调整搜索方向,逐步合成完整报告 |\n| **计算机控制代理(例如浏览器使用)** | 计算机屏幕、浏览器页面、文件系统 | 开放式(内部推理、点击、打字、滚动、截图、代码执行等) | 视觉感知+操作:观察屏幕→识别目标元素→执行操作→验证结果 |\n| **手机助手代理(例如豆包)** | 手机屏幕、已安装应用 | 开放式(内部推理、点击、滑动、打字、打开应用等) | 意图理解+应用控制:理解用户需求→定位目标应用→执行操作→确认完成 |\n| **个人任务代理(例如Pine AI)** | 用户账户信息、历史账单、服务提供商知识库 | 开放式(内部推理、打电话、发邮件、填表、与用户确认) | 多步骤任务执行:收集信息→制定谈判策略→联系服务提供商→谈判→报告结果 |\n\n这些系统具有三个共同特征:**开放式动作空间**——不是从固定的按钮集合中选择,而是生成任意自然语言和代码;**内部推理**——在行动前进行规划;以及**连续交互**——根据环境反馈调整策略。这些能力正是源于推理引擎、工作上下文和动作接口的相互作用——也就是大语言模型(LLM)、上下文和工具的相互作用。\n\n### 工具:代理的动作接口\n\n工具是代理与外部世界的桥梁。它们将代理从被动观察者转变为能够搜索、写入文件、运行代码、调用API、发送消息或操作界面的主动系统。没有工具,代理仅限于文本生成;有了工具,它就能对外部系统采取行动。\n\n为了系统地讨论工具,我们可以根据代理与世界交互的方向将其分为五类。在这个阶段,简要概述每种类型的代表性场景就足以建立整体图景;后续章节将深入探讨每种类型。\n\n**感知工具**使代理能够获取信息:搜索引擎提供实时网络数据,文件系统读取本地文档,API和数据库连接外部服务和企业核心数据。\n\n**执行工具**使代理能够对外部系统采取行动:代码执行、文件操作、系统命令和外部API调用将决策转化为具体行动。\n\n**协作工具**使代理能够与其他代理分工合作:将专门任务委派给子代理,在关键决策点请求人类确认,或在多代理系统中协调行动。\n\n**事件触发工具**的调用方式与前三种类别根本不同:代理不会调用它们;它们作为外部输入到达,触发代理开始工作。新邮件到来、预定时间到达或另一个系统触发Webhook回调;事件激活代理并启动推理和行动。代理本身不会调用这些工具,但它们仍然是代理与外部世界交互的渠道,因此我们将其计入广义的工具系统中。\n\n**用户通信工具**是代理与用户通信的渠道。执行工具改变外部世界,而通信工具传递信息——通过短信、语音通话、电子邮件等传递代理的进展或主动签到。\n\n第4章将涵盖这五类工具的完整分类法和设计原则。工具设计的质量直接决定了代理能够可靠完成的任务:接口定义模糊,模型会误用它们;错误处理不佳,单个工具失败可能会让代理陷入困境;权限范围过广,一个代理错误可能会不可逆转。随着MCP(模型上下文协议)标准的推广,集成工具变得像安装插件一样简单——生态系统正在迅速扩展,但设计原则不会过时。\n\n**工具调用**(也称为函数调用)是现代大语言模型代理的核心能力:它让模型以结构化方式调用外部工具,将大语言模型从纯文本生成器转变为能够通过外部接口行动的智能系统。本书通篇使用“工具调用”这一术语。\n\n工具调用分为四个步骤:首先,上下文告知模型可用的工具(名称、用途、参数);然后,模型自行决定是否调用工具、调用哪个工具以及使用什么参数;接下来,工具运行后,其结果附加到上下文中;最后,模型根据该结果决定下一步行动。这个循环是后续章节介绍的ReAct循环的基础。\n\n以天气查询为例,API级别四步过程的简化表示如下:\n\n```\n步骤1:声明工具 步骤2:模型决定调用\ntools: [{ assistant: {\n name: \"get_weather\", tool_calls: [{\n parameters: { function: \"get_weather\",\n city: \"string\" arguments: {city: \"Beijing\"}\n } }]\n}] }\n\n步骤3:结果附加到上下文 步骤4:模型根据结果回应\ntool: { assistant: {\n tool_call_id: \"call_1\", content: \"Today in Beijing: 28°C, sunny.\"\n content: '{\"temp\":28,\"sky\":\"clear\"}' }\n} }\n```\n\n开发者只需定义工具并执行调用;模型本身决定是否调用、调用哪个工具以及传递什么参数。第2章将详细检查这种API结构。\n\n为代理设计工具时,从任务所需的最窄能力开始,然后随着任务变得更复杂逐步扩展。如果任务只需要基本算术,具有明确定义参数的计算器就足够了;当任务扩展到读取电子表格、清理缺失值、计算统计数据和绘制图表时,受限制的Python代码解释器比不断增长的专门工具集合更容易组合和探索。但通用性也会增加错误风险并扩大攻击面:代码必须在隔离沙箱中运行,默认禁用网络访问,无法访问授权工作目录外的文件,并且对执行时间、CPU、内存和输出大小有限制。\n\n同样,单个日志工具适合记录一次执行;对于耗时数小时甚至数天的长期任务,受控虚拟工作目录可以保存计划、中间结果、执行日志和最终工件,以便代理在多次运行中恢复。该目录还应限制可读和可写路径、存储容量和文件类型,并防止路径遍历,而不是将整个主机文件系统暴露给代理。\n\n通用工具并不总是比专门工具更好。高风险操作或受严格业务约束的操作——例如支付、数据删除、发送电子邮件和生产部署——仍应作为具有明确参数、受限权限和端到端可审计性的专用工具公开,并在必要时添加预览和人类确认。因此,工具设计的核心原则是:**使用通用基础能力进行组合和探索;使用专用工具约束高风险操作并强制执行严格业务规则**。\n\n### 大语言模型:代理的推理引擎\n\n大语言模型(LLM)是代理的决策核心。给定用户请求,它首先必须推断真实意图(用户所说的往往不是他们实际想要的),然后将模糊或复杂的任务分解为可执行步骤。在整个执行过程中,它不断做出决策:下一步做什么、是否调用工具、调用哪个工具以及使用什么参数。这种理解-规划-执行能力来自预训练期间积累的知识,是工作流和自主代理都依赖的基础。\n\n大语言模型代理的一个显著能力是**内部推理**——在行动前,代理可以规划和推理任务。这不会改变外部环境,但会显著改善后续行动。这种能力来自预训练(在大量互联网文本上的初始训练,通过该训练模型学习语言模式和世界知识):模型利用人类知识中编码的推理模式,包括数学定律、因果关系和分解问题的策略。因此,代理的推理不是盲目试错,而是基于结构化知识体系。"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "021785390223732e4feb8d226730241afb8b5c4c4c35b32d61f16",
"model": "doubao-seed-1-6-250615",
"created": 1785390240,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":5,\"evidence\":\"All table content, feature descriptions, tool classifications, and technical details are accurately translated without omissions or inventions. For example, the core principle of tool design is correctly rendered as '使用通用基础能力进行组合和探索;使用专门工具约束高风险操作并强制执行严格业务规则'.\"},\"fluency\":{\"score\":4,\"evidence\":\"Most sentences are fluent, but phrases like '主动询问' (X) for 'proactive check-in' are slightly less natural compared to Y's '主动签到'. The term '迭代深化' is consistent but slightly rigid.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of technical terms: '开放式行动空间' (open-ended action space), '内部推理' (internal reasoning), '工具调用' (tool calling), and 'MCP (模型上下文协议)' (Model Context Protocol).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Table structure, headings (###), code blocks (```), and formatting (**bold**) are fully preserved. The code example for weather query steps is identical to the source.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All critical information is accurately translated, including the three shared features of agents and tool design principles. The phrase '端到端可审计性' (end-to-end auditability) is precisely rendered.\"},\"fluency\":{\"score\":5,\"evidence\":\"More natural phrasing: '主动签到' (proactive check-in) instead of '主动询问', and '委派给子代理' (delegating to sub-agents) flows better than '委托给子代理'. Sentences like '意图理解+应用控制' are idiomatic.\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent terminology matching X, e.g., '开放式动作空间' (open-ended action space), '沙盒' (sandbox), and '路径遍历' (path traversal). 'MCP (模型上下文协议)' is correctly retained.\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Markdown elements are fully preserved, including table alignment, heading levels (###), code blocks, and bolded terms like '内部推理' (internal reasoning).\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y demonstrates slightly better fluency with more natural expressions such as '主动签到' (proactive check-in) and '委派给子代理' (delegating to sub-agents), while maintaining equal accuracy, terminology consistency, and markdown fidelity as X.\"}",
"usage": {
"prompt_tokens": 5650,
"completion_tokens": 533,
"total_tokens": 6183
}
},
"latency_milliseconds": 16790,
"validation": {
"valid": true
}
}
]
}