Files
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

43 lines
28 KiB
JSON

{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nSeen through the lens of Harness engineering, each chapter of this book systematically builds out one component of the Harness. Security, meanwhile, belongs to no single chapter; it is a cross-cutting concern of the whole book (a cross-cutting concern touches many parts of a system at once—the way logging, in software engineering, has to thread through every module). The table below presents the Harness functions, security aspects, and corresponding chapters in a single view:\n\n| Harness Focus | Corresponding Chapter | Core Content | Security Concerns |\n|--------------------|--------------------|-------------------------------|------------------------|\n| Context Design | Chapter 2 (Context Engineering) | Prompt engineering, Agent status bar, context compression, Agent Skills | Prompt injection and information leakage |\n| Context Expansion (Knowledge Persistence) | Chapter 3 (Knowledge Base) | User memory, RAG, structured indexing, agentic RAG | Sensitive information exposure, privacy protection |\n| Tool Design and Security Constraints | Chapter 4 (Tool Design) | Tool classification, permission control, MCP standard, asynchronous architecture | Misoperation, unauthorized access, irreversible operations |\n| Tool Verification and Correction | Chapter 5 (Code Generation) | Coding Agent's Harness, test-driven development, codified rules | Identity impersonation, responsibility attribution |\n| System-Level Verification | Chapter 6 (Evaluation) | Evaluation environment, datasets, automated evaluation, observability | — |\n| Model-Level Correction | Chapter 7 (Post-Training) | SFT (Supervised Fine-Tuning), Reinforcement Learning—encoding feedback signals accumulated by the Harness into model parameters, as an extension of Harness engineering | Goal misalignment, alignment and robustness |\n| System-Level Correction | Chapter 8 (Self-Evolution) | Externalized learning, tool creation, experience accumulation | — |\n| Multimodal Context and Tools | Chapter 9 (Multimodal and Real-Time Interaction) | Voice Agent, Computer Use, robotic operation | Security filtering of multimodal input, permission control in real-time interaction |\n| Constraints and Corrections Among Multiple Agents | Chapter 10 (Multi-Agent Collaboration) | Collaboration architecture, failure modes, Agent society | Trust boundary violations between Agents, shared resource conflicts |\n\nAnthropic's practice in building long-running Agents shows how Harness design can solve problems the model itself cannot. They split complex tasks between an \"Initialization Agent\" (setting up the environment, decomposing the task list) and an \"Execution Agent\" (making incremental progress each session and leaving clear handover artifacts), using a structured Harness to tackle the two failure modes of long tasks: running out of context and declaring the task done prematurely. The chapters ahead work through the Harness component by component—Chapter 2 begins with the most central one, context engineering, and Chapter 5 lays out the complete practice of Harness engineering in Coding Agents.\n\n## Chapter Summary\n\nThis chapter has built a practice-first framework for understanding and constructing AI Agents.\n\n**Agent = Reasoning Engine + Working Context + Action Interfaces**: The LLM provides reasoning and decision-making, context supplies the working set of information available at decision time, and tools provide the action interfaces. None of the three is dispensable.\n\n**Expanding Context and Tools Is the Primary Capability Lever**: Once the model is fixed, redefining or enlarging the observation and action spaces—that is, expanding context and tools—can often turn an unsolvable task into a solvable one directly. The evolution from Manus to OpenClaw shows that much of generality comes from broadening the interface boundary; that expansion must remain on-demand and be paired with permissions and verification.\n\n**Context Is the Decisive Factor**: Context consists of a static prefix (system prompt + tool definitions) and a dynamic trajectory (message history). Ablation shows that removing any component degrades the system markedly. The essence of the ReAct loop is appending to the trajectory, over and over, so the model keeps advancing the task.\n\n**Harness Is the Competitive Advantage**: Model capability is commoditizing; the real differentiator is the Harness—the constrain, verify, and correct mechanisms built around context and tools that enable reliable task completion. In production-grade Agent systems, the vast majority of Harness code goes into these safeguards, not the context and tools alone.\n\n**From Workflow to Autonomous Agent**: Prompts first, then workflows, autonomous Agents last—that ordering is the most practical way to reduce unexpected behavior. Every orchestration pattern has situations where it fits; no single pattern is best everywhere.\n\n**Security Is an Architectural Issue**: Guardrails, human-in-the-loop intervention, alignment (keeping the model's behavior consistent with human intent)—security has to be designed in from the first line of code, not patched on before launch. It spans five levels: model, context, tools, collaboration, and society.\n\nThe next chapter examines the Harness's most central component in depth: context engineering. Chapter 7 covers the Agent concept's academic roots in reinforcement learning and compares traditional RL with modern LLM Agents.\n\nThe thought questions below are designed to take the chapter's core concepts a level deeper.\n\n## Thought Questions\n\n1. ★★ If you could only add one capability to an Agent system—a stronger model, richer context, or more tools—which would you choose? Under what conditions would your choice change?\n2. ★★★ In the ReAct loop, each of the Agent's LLM calls receives the full history trajectory, so as the trajectory grows, the cost of this design grows quadratically. Can that quadratic growth be broken without losing critical information?\n3. ★★ The \"Model as Agent\" paradigm means models are becoming more autonomous in tool-calling decisions. However, this chapter argues that the importance of Harness engineering is actually increasing. How can these two trends coexist? Where does the future core value of Agent frameworks lie?\n4. ★★ In the ablation experiment, the absence of \"tool result feedback\" caused the Agent to fall into an infinite loop. In a production environment, besides missing tool results, what other situations could cause an Agent to loop? What detection and termination mechanisms would you design?\n5. ★ This chapter analyzed five Agent products along three dimensions: working context, action interfaces, and strategy. Pick an AI product you use daily, analyze it along the same three dimensions, and judge whether its architecture is appropriate. If you were designing it, what would you improve?\n6. ★★ If you were to design a customer service system specifically for booking flights, would you choose a workflow pattern or an autonomous Agent pattern? Is it possible to mix both patterns in the same system?\n7. ★★★ The guardrails section mentioned tool risk ratings. If a tool is generally low-risk but becomes high-risk with specific parameter combinations (e.g., `delete_file` deleting a normal file vs. deleting a system file), how would you design dynamic risk assessment?\n8. ★★ In the Agent product table in this chapter, all Agents have an \"open-ended\" action space. In what scenarios would a constrained action space (e.g., only being able to choose from predefined options) be superior to an open-ended one?\n9. ★★ The human-in-the-loop intervention mechanism requires the Agent to \"gracefully hand over control.\" However, in practice, the user might be offline, respond slowly, or give vague instructions. What should the Agent do in such cases?\n10. ★★★ The introduction states that \"good design principles should transcend model iteration cycles.\" Give an example of a current Agent design principle that you believe might become obsolete as models improve, and explain your reasoning.\n\n\nANONYMOUS CHINESE X:\n### 从Harness工程视角看[第9/9部分]AI代理入门\n\n从Harness工程的角度来看,本书的每一章都系统地构建了Harness的一个组成部分。与此同时,安全问题并不属于单一章节;它是贯穿整本书的跨领域关注点(跨领域关注点会同时涉及系统的多个部分,就像软件工程中的日志必须贯穿每个模块一样)。下表以单一视图呈现了Harness功能、安全方面及相应章节:\n\n| Harness关注点 | 对应章节 | 核心内容 | 安全关注点 |\n|--------------------|--------------------|-------------------------------|------------------------|\n| 上下文设计 | 第2章(上下文工程) | 提示工程、Agent状态栏、上下文压缩、Agent技能 | 提示注入和信息泄露 |\n| 上下文扩展(知识持久化) | 第3章(知识库) | 用户记忆、RAG、结构化索引、代理式RAG | 敏感信息暴露、隐私保护 |\n| 工具设计与安全约束 | 第4章(工具设计) | 工具分类、权限控制、MCP标准、异步架构 | 误操作、未授权访问、不可逆操作 |\n| 工具验证与纠正 | 第5章(代码生成) | 编码Agent的Harness、测试驱动开发、编码规则 | 身份冒充、责任归属 |\n| 系统级验证 | 第6章(评估) | 评估环境、数据集、自动化评估、可观测性 | — |\n| 模型级纠正 | 第7章(训练后) | SFT(监督微调)、强化学习——将Harness积累的反馈信号编码到模型参数中,作为Harness工程的扩展 | 目标不一致、对齐性和鲁棒性 |\n| 系统级纠正 | 第8章(自我进化) | 外部化学习、工具创建、经验积累 | — |\n| 多模态上下文与工具 | 第9章(多模态与实时交互) | 语音Agent、计算机使用、机器人操作 | 多模态输入的安全过滤、实时交互中的权限控制 |\n| 多Agent之间的约束与纠正 | 第10章(多Agent协作) | 协作架构、失败模式、Agent社会 | Agent之间的信任边界违规、共享资源冲突 |\n\nAnthropic在构建长期运行的Agent时的实践展示了Harness设计如何解决模型自身无法解决的问题。他们在“初始化Agent”(设置环境、分解任务列表)和“执行Agent”(每次会话逐步推进并留下清晰的交接工件)之间分配复杂任务,使用结构化的Harness来应对长任务的两种失败模式:上下文耗尽和过早宣告任务完成。接下来的章节将逐个介绍Harness组件——第2章从最核心的部分开始,即上下文工程,第5章阐述了编码Agent中Harness工程的完整实践。\n\n### 章节总结\n\n本章构建了一个以实践为导向的框架,用于理解和构建AI代理。\n\n**Agent = 推理引擎 + 工作上下文 + 行动接口**:大语言模型提供推理和决策,上下文提供决策时可用的工作信息集,工具提供行动接口。三者缺一不可。\n\n**扩展上下文和工具是主要的能力杠杆**:一旦模型固定,重新定义或扩大观察和行动空间——即扩展上下文和工具——通常可以直接将无法解决的任务转化为可解决的任务。从Manus到OpenClaw的演变表明,很多通用性来自于扩展接口边界;这种扩展必须按需进行,并与权限和验证相结合。\n\n**上下文是决定性因素**:上下文由静态前缀(系统提示 + 工具定义)和动态轨迹(消息历史)组成。消融实验表明,移除任何组件都会显著降低系统性能。ReAct循环的本质是不断向轨迹中追加内容,从而使模型不断推进任务。\n\n**Harness是竞争优势**:模型能力正在商品化;真正的差异化因素是Harness——围绕上下文和工具构建的约束、验证和纠正机制,能够实现可靠的任务完成。在生产级Agent系统中,绝大多数Harness代码都用于这些保障措施,而不仅仅是上下文和工具。\n\n**从工作流到自主Agent**:先有提示词,然后是工作流,最后是自主Agent——这种顺序是减少意外行为的最实用方式。每种编排模式都有其适用的情况;没有一种模式在所有地方都是最好的。\n\n**安全是架构问题**:护栏、人工参与、对齐(使模型行为与人类意图保持一致)——安全必须从代码的第一行开始设计,而不是在发布前修补。它涵盖五个层面:模型、上下文、工具、协作和社会。\n\n下一章将深入探讨Harness最核心的组件:上下文工程。第7章将介绍Agent概念在强化学习中的学术根源,并比较传统RL与现代LLM Agent。\n\n以下是针对本章核心概念进一步深入的思考问题。\n\n### 思考问题\n\n1. ★★ 如果只能给Agent系统添加一种能力——更强的模型、更丰富的上下文或更多工具,你会选择哪一种?在什么条件下你的选择会改变?\n2. ★★★ 在ReAct循环中,Agent的每次大语言模型调用都会接收完整的历史轨迹,因此随着轨迹增长,这种设计的成本呈二次方增长。能否在不丢失关键信息的情况下打破这种二次方增长?\n3. ★★ “模型即Agent”范式意味着模型在工具调用决策上变得更加自主。然而,本章认为Harness工程的重要性实际上在增加。这两种趋势如何共存?Agent框架的未来核心价值在哪里?\n4. ★★ 在消融实验中,缺少“工具结果反馈”导致Agent陷入无限循环。在生产环境中,除了缺少工具结果,还有哪些情况可能导致Agent循环?你会设计哪些检测和终止机制?\n5. ★ 本章从工作上下文、行动接口和策略三个维度分析了五种Agent产品。选择一个你日常使用的AI产品,从相同的三个维度进行分析,并判断其架构是否合适。如果由你设计,你会如何改进?\n6. ★★ 如果你要专门设计一个用于预订航班的客户服务系统,你会选择工作流模式还是自主Agent模式?是否可能在同一系统中混合使用这两种模式?\n7. ★★★ 护栏部分提到了工具风险评级。如果一个工具通常风险较低,但在特定参数组合下变得风险较高(例如`delete_file`删除普通文件与删除系统文件),你会如何设计动态风险评估?\n8. ★★ 在本章的Agent产品表中,所有Agent都有“开放式”行动空间。在什么场景下,受限行动空间(例如只能从预定义选项中选择)比开放式行动空间更优?\n9. ★★ 人工参与干预机制要求Agent“优雅地移交控制权”。然而,在实践中,用户可能离线、响应缓慢或给出模糊指示。这种情况下Agent应该怎么做?\n10. ★★★ 引言中提到“好的设计原则应该超越模型迭代周期”。给出一个你认为随着模型改进可能过时的当前Agent设计原则,并解释原因。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门[第9/9部分]\n从框架工程的角度看,本书的每一章系统地构建了框架的一个组件。与此同时,安全不属于任何单一章节;它是贯穿全书的横切关注点(横切关注点同时触及系统的多个部分——在软件工程中,日志记录必须贯穿每个模块的方式)。下表以单一视图呈现了框架功能、安全方面及对应的章节:\n\n| 框架关注点 | 对应章节 | 核心内容 | 安全关注点 |\n|------------------|------------------------|------------------------------------|----------------------|\n| 上下文设计 | 第2章(上下文工程) | 提示工程、代理状态栏、上下文压缩、代理技能 | 提示注入和信息泄露 |\n| 上下文扩展(知识持久化) | 第3章(知识库) | 用户记忆、RAG、结构化索引、代理式RAG | 敏感信息暴露、隐私保护 |\n| 工具设计和安全约束 | 第4章(工具设计) | 工具分类、权限控制、MCP标准、异步架构 | 误操作、未授权访问、不可逆转操作 |\n| 工具验证和纠正 | 第5章(代码生成) | 编码代理的框架、测试驱动开发、编码规则 | 身份冒充、责任归属 |\n| 系统级验证 | 第6章(评估) | 评估环境、数据集、自动化评估、可观测性 | — |\n| 模型级纠正 | 第7章(后训练) | SFT(监督微调)、强化学习——将框架积累的反馈信号编码到模型参数中,作为框架工程的扩展 | 目标不一致、对齐和鲁棒性 |\n| 系统级纠正 | 第8章(自我演进) | 外部化学习、工具创建、经验积累 | — |\n| 多模态上下文和工具 | 第9章(多模态和实时交互) | 语音代理、计算机使用、机器人操作 | 多模态输入的安全过滤、实时交互中的权限控制 |\n| 多代理之间的约束和纠正 | 第10章(多代理协作) | 协作架构、失败模式、代理社会 | 代理之间的信任边界违反、共享资源冲突 |\n\nAnthropic在构建长期运行代理的实践展示了框架设计如何解决模型自身无法解决的问题。他们将复杂任务在“初始化代理”(设置环境、分解任务列表)和“执行代理”(每次会话逐步推进并留下清晰的交接工件)之间拆分,使用结构化框架应对长任务的两种失败模式:上下文耗尽和过早宣告任务完成。接下来的章节将逐个组件讲解框架——第2章从最核心的上下文工程开始,第5章阐述编码代理中框架工程的完整实践。\n\n## 章节总结\n本章构建了一个以实践为导向的理解和构建人工智能代理的框架。\n\n**代理=推理引擎+工作上下文+行动接口**:大语言模型提供推理和决策,上下文提供决策时可用的工作信息集,工具提供行动接口。三者缺一不可。\n\n**扩展上下文和工具是主要的能力杠杆**:一旦模型固定,重新定义或扩大观测空间和行动空间——即扩展上下文和工具——通常可以直接将无法解决的任务变为可解决的任务。从Manus到OpenClaw的演进表明,通用性很大程度上来自接口边界的拓宽;这种扩展必须按需进行,并与权限和验证相结合。\n\n**上下文是决定性因素**:上下文由静态前缀(系统提示+工具定义)和动态轨迹(消息历史)组成。消融实验表明,移除任何组件都会显著降低系统性能。ReAct循环的本质是不断向轨迹追加内容,使模型持续推进任务。\n\n**框架是竞争优势**:模型能力趋于商品化;真正的差异化在于框架——围绕上下文和工具构建的约束、验证和纠正机制,使任务能够可靠完成。在生产级代理系统中,框架代码的绝大部分用于这些保障措施,而不仅仅是上下文和工具。\n\n**从工作流到自主代理**:先提示,然后工作流,最后自主代理——这个顺序是减少意外行为的最实用方式。每种编排模式都有适用场景;没有一种模式在所有地方都最佳。\n\n**安全是架构问题**:护栏、人工参与干预、对齐(使模型行为与人类意图一致)——安全必须从第一行代码开始设计,而不是在发布前修补。它涵盖五个层面:模型、上下文、工具、协作和社会。\n\n下一章将深入探讨框架中最核心的组件:上下文工程。第7章涵盖代理概念在强化学习中的学术根源,并比较传统强化学习与现代大语言模型代理。\n\n以下思考问题旨在将本章核心概念进一步深化。\n\n## 思考问题\n\n1. ★★ 如果你只能给代理系统添加一种能力——更强的模型、更丰富的上下文或更多工具,你会选择哪一种?在什么条件下你的选择会改变?\n2. ★★★ 在ReAct循环中,代理的每次大语言模型调用都接收完整的历史轨迹,因此随着轨迹增长,这种设计的成本呈二次方增长。能否在不丢失关键信息的情况下打破这种二次方增长?\n3. ★★ “模型即代理”范式意味着模型在工具调用决策上变得更加自主。然而,本章认为框架工程的重要性实际上在增加。这两种趋势如何共存?代理框架的未来核心价值在哪里?\n4. ★★ 在消融实验中,“工具结果反馈”的缺失导致代理陷入无限循环。在生产环境中,除了缺少工具结果,还有哪些情况可能导致代理循环?你会设计什么检测和终止机制?\n5. ★ 本章从工作上下文、行动接口和策略三个维度分析了五种代理产品。挑选一个你日常使用的人工智能产品,沿这三个维度进行分析,并判断其架构是否合适。如果由你设计,你会如何改进?\n6. ★★ 如果你要专门设计一个用于航班预订的客户服务系统,你会选择工作流模式还是自主代理模式?在同一系统中混合使用两种模式是否可能?\n7. ★★★ 护栏部分提到了工具风险评级。如果一个工具通常风险较低,但在特定参数组合下变得风险较高(例如`delete_file`删除普通文件与删除系统文件),你会如何设计动态风险评估?\n8. ★★ 本章的代理产品表中,所有代理都有“开放式”行动空间。在什么场景下,受限行动空间(例如只能从预定义选项中选择)比开放式行动空间更优?\n9. ★★ 人工参与干预机制要求代理“优雅地移交控制权”。然而,在实践中,用户可能离线、响应缓慢或给出模糊指令。在这种情况下,代理应该怎么做?\n10. ★★★ 引言中提到“良好的设计原则应超越模型迭代周期”。给出一个你认为随着模型改进可能过时的当前代理设计原则,并解释你的理由。"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853906543223ea4fa82759f0cae4f2ffa723d4f9dd6975ed3",
"model": "doubao-seed-1-6-250615",
"created": 1785390677,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Accurately translates technical content with minor omissions. For example, the English source's opening line 'Seen through the lens of Harness engineering' is correctly translated as '从Harness工程的角度来看' in X, preserving the technical perspective. However, X adds an unnecessary section title '[第9/9部分]AI代理入门' which is not present in the original Markdown.\"},\"fluency\":{\"score\":4,\"evidence\":\"Generally fluent with natural phrasing. For instance, 'cross-cutting concern' is translated as '跨领域关注点' with the explanatory clause '就像软件工程中的日志必须贯穿每个模块一样' smoothly integrated. Minor awkwardness in '代理式RAG' (vs. Y's '代理式RAG' which is identical here).\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent and correct technical term usage. 'Harness' is retained as 'Harness' throughout (e.g., 'Harness工程', 'Harness组件'), 'ReAct loop' as 'ReAct循环', 'ablation' as '消融实验'. Terms like '提示注入' (prompt injection) and '身份冒充' (identity impersonation) are accurately translated and consistent.\"},\"markdown_code_fidelity\":{\"score\":3,\"evidence\":\"Table structure and basic Markdown elements are preserved, but heading levels are incorrect. Original uses '## Chapter Summary' and '## Thought Questions', while X demotes these to '### 章节总结' and '### 思考问题'. Additionally, X adds an extraneous top-level heading '### 从Harness工程视角看[第9/9部分]AI代理入门' not in the source.\"}},\"Y\":{\"accuracy\":{\"score\":3,\"evidence\":\"Significant terminology inconsistency undermines accuracy. 'Harness' is incorrectly translated as '框架' (framework) throughout, altering the core concept (e.g., '框架工程' instead of 'Harness工程'). This changes the author's intended meaning, as 'Harness' specifically refers to constrain/verify/correct mechanisms, not a general 'framework'. Other content is mostly accurate but marred by this critical mistranslation.\"},\"fluency\":{\"score\":4,\"evidence\":\"Fluent and natural phrasing overall. 'cross-cutting concern' is translated as '横切关注点' with the explanatory example smoothly rendered. Sentences like '安全必须从第一行代码开始设计,而不是在发布前修补' flow well. Comparable to X in fluency, with minor differences in style (e.g., '可逆转操作' vs. X's '不可逆操作'—both correct).\"},\"terminology\":{\"score\":2,\"evidence\":\"Critical mistranslation of 'Harness' as '框架' (framework) throughout, which is technically incorrect and inconsistent with the source's definition of 'Harness' as constrain/verify/correct mechanisms. Other terms are mostly consistent but this single error invalidates key technical claims.\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Heading levels are mostly correct: '## 章节总结' and '## 思考问题' match the original's '## Chapter Summary' and '## Thought Questions'. Table structure, list items, and ★ symbols are preserved. No extraneous headings added, unlike X. Minor issue: loss of the italicized '(Supervised Fine-Tuning)' in 'SFT (Supervised Fine-Tuning)' (translated as 'SFT(监督微调)' in both X and Y, which is acceptable).\"}},\"preferred\":\"X\",\"preference_evidence\":\"X is preferred because it accurately preserves the critical term 'Harness' and maintains the author's intended meaning, whereas Y's mistranslation of 'Harness' as '框架' (framework) fundamentally alters the core technical concept. X has better terminology consistency and higher accuracy despite minor Markdown heading errors, while Y's critical terminology error undermines its technical validity.\"} ,\"preferred\":\"X\",\"preference_evidence\":\"X is preferred because it accurately preserves the critical term 'Harness' and maintains the author's intended meaning, whereas Y's mistranslation of 'Harness' as '框架' (framework) fundamentally alters the core technical concept. X has better terminology consistency and higher accuracy despite minor Markdown heading errors, while Y's critical terminology error undermines its technical validity.\"}",
"usage": {
"prompt_tokens": 4975,
"completion_tokens": 918,
"total_tokens": 5893
}
},
"latency_milliseconds": 23526,
"validation": {
"valid": true
}
}
]
}