Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
43 lines
31 KiB
JSON
43 lines
31 KiB
JSON
{
|
||
"schema_version": 1,
|
||
"credential_free": true,
|
||
"attempts": [
|
||
{
|
||
"attempt": 1,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n# Context Engineering\n\nChapter 1 defined context as the Agent's working set of information at the moment of decision. Designing and managing that context—what we call **Context Engineering**—is central to building effective Agents. In practice, context includes everything the model receives for a given interaction: the conversation history, system instructions, tool definitions, retrieved documents, runtime state, and other task-specific information. From the Harness perspective introduced in Chapter 1, context engineering implements much of the Harness's \"Context and Tools\" layer: it decides what information the Agent sees at each decision point and how that information is organized. Good context design gives the model the right background, constraints, and action interfaces so its general reasoning ability can be applied effectively to the task.\n\n\n\n## Context: The Ceiling of Agent Capability\n\nLarge language models achieve strong results on standardized benchmarks, but often underperform in real-world business settings. The reason is straightforward: model capabilities are general-purpose, while concrete tasks depend on local knowledge such as product architecture, business rules, operational constraints, and internal conventions. This information is usually absent from the model's parameters.\n\nConsider a highly capable engineer joining a new team. They may have deep theoretical knowledge and strong programming ability, but they do not yet understand the product architecture, business logic, technical debt, or team norms. If key architectural decisions are scattered across individual memories and the codebase is poorly documented, even an exceptional engineer will struggle to deliver value quickly. Today's AI Agents face the same problem.\n\nConsider a Coding Agent. Given the same instruction, \"Help me fix this bug,\" the quality of the context the Agent receives determines whether it can complete the task:\n\n- **Code context**: The codebase structure, module responsibilities, core data structures, and coding standards. Without this information, the Agent may produce code that is syntactically correct but inconsistent with the project's style or architecture.\n- **Process requirements**: Git branching strategy, commit conventions, review process, and CI/CD requirements. Without this information, the Agent may commit untested code directly to the main branch.\n- **Environment configuration**: Development setup, test database connection strings, staging deployment procedures, and API key management practices. Without this information, a fix that works locally may fail immediately in the test environment.\n\nThese three categories—code, process, and environment—form the minimum context an Agent needs to work effectively. The model's inherent capability is only the foundation; context sets the ceiling for Agent capability. A moderately capable model with well-organized context can often outperform a stronger model operating with insufficient context.\n\nContext engineering is therefore central to building effective Agents with today's models. It is not merely a matter of adding more text to a prompt. It requires systematically designing, organizing, and providing the background knowledge the model needs to complete a task.\nContext engineering is a technical problem, but more fundamentally it is an organizational problem. In many teams, critical knowledge remains tacit: architectural decisions live in the memories of senior engineers, business rules are transmitted informally, and important context is buried in private chat logs. If the team itself is a poor information environment, even a strong AI Agent will be limited.\n\nTeams that work effectively in remote settings often also provide effective environments for AI Agents. Open-source projects such as the Linux kernel are instructive examples: developers distributed across the world have maintained the project for more than thirty years. This works because the project has a transparent, documentation-driven communication culture. Discussions are public, decisions are recorded, and newcomers can understand the evolution of the code by reading the history. The same working style naturally creates an AI-friendly environment: information is public, retrievable, and structured.\n\nTreat an AI Agent as a new team member each time it starts a task. With sufficient background, it can produce high-quality work; without that background, much of its intelligence is wasted. Building an AI-native team is therefore primarily a documentation effort, not merely a matter of deploying new tools.\n\nOpenAI researcher Jiayi Weng expressed this point clearly: **\"For both humans and models, the most important thing is Context.\"** Reflecting on his own work, he noted: \"My work at OpenAI isn't that difficult. If someone else had all my context, they could do it too.\" The same principle applies to Agents: the ceiling of Agent capability is not determined only by model size, but by the completeness and precision of the context provided at each decision point. Weng also observed that the central problem in teamwork is inconsistency of context, and that one reason AI cannot replace humans in the short term is that AI and humans do not share the same environment. Context engineering addresses exactly this problem: how to systematically deliver the structured background information an Agent needs to the model.\n\nThe next question is how this contextual information is provided to the LLM at the technical level.\n\n## How Agents Call LLMs: The API-Level Context Structure\n\nThis section uses OpenAI's Chat Completions API as a concrete example. Anthropic, Google, and other providers differ in details, but their Agent-facing APIs follow a similar pattern: each model call is constructed from a structured conversation history plus a set of available tool definitions. Understanding this structure is the foundation for the context engineering techniques discussed later in this chapter.\n\n### The Four Message Roles\n\nIn Chat Completions-style APIs, the core input is a **message list**, usually named `messages`. Each message has a `role` field that tells the model how to interpret the message and where it came from:\n\n- **system**: Developer-written instructions that define the Agent's identity, behavior, constraints, and workflow. The model treats this as a high-priority instruction. In most conversations, the system message appears once at the beginning of the message list.\n- **user**: Input from the end user, representing the request the Agent needs to handle.\n- **assistant**: Previous model outputs, including natural-language replies and tool call requests. In multi-turn interactions, these messages are included in later requests so the next stateless model call has access to the prior trajectory.\n- **tool**: Results returned after the Agent framework executes a tool. Each tool result is linked to the corresponding tool call through `tool_call_id`, allowing the model to associate each result with the request that produced it.\n\nTool definitions are not messages. They are provided in a separate `tools` field, which declares the tools available to the model and specifies the parameters each tool accepts.\n\n### Single-Turn Request: The Simplest API Call\n\n\n\nStart with the simplest case: a single request without tool calls. The user asks, \"Hello, who are you?\" The example uses a locally deployed Qwen3-0.6B model, connecting it to the local LLM deployment experiment later in this section. The timestamps in the example are for demonstration only and are unrelated to the book's timeline.\n\n```javascript\n// ═══ Request constructed by the Agent framework ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← Written by developer\n \"content\": \"You are a helpful coding assistant. Follow user instructions.\"\n },\n {\n \"role\": \"user\", // ← User input\n \"content\": \"Hello, who are you?\"\n }\n ]\n}\n```\n\n```javascript\n// ═══ Response returned by the API ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← Generated by model\n \"content\": \"Hi! I'm a coding assistant. I can help you write code, debug issues, and explain technical concepts. How can I help?\"\n }\n }]\n}\n```\n\nThis request contains only two messages: one system message containing rules written by the developer and one user message containing the user's input. The model returns an assistant message as the reply. This is the most basic LLM API interaction pattern: **each call is stateless, so the request's message list must contain all the information the model needs**.\n\n### Multi-Turn Interaction with Tool Calls: The Core Loop of an Agent\n\nReal Agent workflows are usually more complex than a single-turn Q&A. When a user asks, \"What's the current time and weather in Vancouver?\", the model needs access to dynamic external information: the current time and the latest weather. The following example walks through each interaction between the Agent framework and the model.\n\n\n\n**First API call — Agent framework sends the initial request:**\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第1/17部分]\n### 上下文工程\n\n第1章将上下文定义为代理在决策时刻的工作信息集。设计和管理该上下文——我们称之为**上下文工程**——是构建有效代理的核心。在实践中,上下文包括模型在给定交互中接收的所有内容:对话历史、系统指令、工具定义、检索到的文档、运行时状态和其他特定任务的信息。从第1章介绍的框架角度看,上下文工程实现了框架的大部分“上下文和工具”层:它决定代理在每个决策点看到什么信息以及这些信息如何组织。良好的上下文设计为模型提供正确的背景、约束和行动接口,使其通用推理能力能够有效地应用于任务。\n\n\n\n## 上下文:代理能力的上限\n大语言模型在标准化基准测试中取得了优异成绩,但在真实业务场景中往往表现不佳。原因很简单:模型能力是通用的,而具体任务依赖于本地知识,如产品架构、业务规则、操作约束和内部约定。这些信息通常不存在于模型的参数中。\n\n设想一位能力很强的工程师加入一个新团队。他们可能有深厚的理论知识和强大的编程能力,但他们还不了解产品架构、业务逻辑、技术债务或团队规范。如果关键架构决策分散在个人记忆中,代码库文档匮乏,即使是杰出的工程师也难以迅速创造价值。如今的人工智能代理面临同样的问题。\n\n以编码代理为例。面对同样的指令“帮我修复这个漏洞”,代理接收的上下文质量决定了它能否完成任务:\n\n- **代码上下文**:代码库结构、模块职责、核心数据结构和编码标准。没有这些信息,代理可能生成语法正确但与项目风格或架构不一致的代码。\n- **流程要求**:Git分支策略、提交规范、审查流程和CI/CD要求。没有这些信息,代理可能直接将未经测试的代码提交到主分支。\n- **环境配置**:开发设置、测试数据库连接字符串、暂存部署程序和API密钥管理实践。没有这些信息,本地运行正常的修复可能在测试环境中立即失败。\n\n这三个类别——代码、流程和环境——构成了代理有效工作所需的最小上下文。模型固有的能力只是基础;上下文设定了代理能力的上限。具有良好组织上下文的中等能力模型往往能胜过在上下文不足情况下运行的更强模型。\n\n因此,上下文工程是用当今模型构建有效代理的核心。这不仅仅是向提示中添加更多文本的问题。它需要系统地设计、组织并提供模型完成任务所需的背景知识。上下文工程是一个技术问题,但从根本上说是一个组织问题。在许多团队中,关键知识仍然是隐性的:架构决策存在于高级工程师的记忆中,业务规则非正式传递,重要上下文埋藏在私人聊天记录中。如果团队本身是一个糟糕的信息环境,即使强大的人工智能代理也会受到限制。\n\n在远程环境中有效工作的团队通常也为人工智能代理提供了有效的环境。像Linux内核这样的开源项目就是有启发性的例子:分布在世界各地的开发者维护该项目已有三十多年。这之所以可行,是因为该项目具有透明的、以文档为驱动的沟通文化。讨论是公开的,决策被记录,新人可以通过阅读历史了解代码的演进。同样的工作风格自然创造了对人工智能友好的环境:信息是公开的、可检索的且结构化的。\n\n每次代理开始任务时,将其视为一个新的团队成员。有了足够的背景,它可以产出高质量的工作;没有背景,其大部分智能都会被浪费。因此,构建人工智能原生团队主要是一项文档工作,而不仅仅是部署新工具的问题。\n\nOpenAI研究员翁佳怡清晰地表达了这一点:**“对人类和模型来说,最重要的是上下文。”** 回顾自己的工作,他指出:“我在OpenAI的工作并不难。如果其他人拥有我所有的上下文,他们也能做到。” 同样的原则适用于代理:代理能力的上限不仅由模型大小决定,还由每个决策点提供的上下文的完整性和精确性决定。翁佳怡还观察到团队合作中的核心问题是上下文不一致,而人工智能短期内无法取代人类的一个原因是人工智能和人类没有共享相同的环境。上下文工程正是解决这个问题:如何系统地向模型提供代理所需的结构化背景信息。\n\n下一个问题是如何在技术层面将这些上下文信息提供给大语言模型。\n\n## 代理如何调用大语言模型:API级上下文结构\n本节以OpenAI的Chat Completions API为例进行具体说明。Anthropic、谷歌等提供商在细节上有所不同,但它们面向代理的API遵循类似模式:每次模型调用由结构化对话历史和一组可用工具定义构建而成。理解这种结构是本章后续讨论的上下文工程技术的基础。\n\n### 四种消息角色\n在Chat Completions风格的API中,核心输入是**消息列表**,通常命名为`messages`。每条消息有一个`role`字段,告诉模型如何解释消息以及它来自哪里:\n\n- **system**:开发者编写的指令,定义代理的身份、行为、约束和工作流。模型将其视为高优先级指令。在大多数对话中,系统消息在消息列表开头出现一次。\n- **user**:最终用户的输入,代表代理需要处理的请求。\n- **assistant**:之前的模型输出,包括自然语言回复和工具调用请求。在多轮交互中,这些消息包含在后续请求中,以便无状态的下一次模型调用能够访问之前的轨迹。\n- **tool**:代理框架执行工具后返回的结果。每个工具结果通过`tool_call_id`与相应的工具调用关联,使模型能够将每个结果与其产生的请求关联起来。\n\n工具定义不是消息。它们在单独的`tools`字段中提供,该字段声明模型可用的工具并指定每个工具接受的参数。\n\n### 单轮请求:最简单的API调用\n\n\n\n从最简单的情况开始:没有工具调用的单轮请求。用户问“你好,你是谁?”。示例使用本地部署的Qwen3-0.6B模型,与本节后面的本地大语言模型部署实验相关联。示例中的时间戳仅用于演示,与本书时间线无关。\n\n```javascript\n// ═══ 代理框架构造的请求 ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 开发者编写\n \"content\": \"You are a helpful coding assistant. Follow user instructions.\"\n },\n {\n \"role\": \"user\", // ← 用户输入\n \"content\": \"Hello, who are you?\"\n }\n ]\n}\n```\n\n```javascript\n// ═══ API返回的响应 ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 模型生成\n \"content\": \"Hi! I'm a coding assistant. I can help you write code, debug issues, and explain technical concepts. How can I help?\"\n }\n }]\n}\n```\n\n这个请求仅包含两条消息:一条系统消息包含开发者编写的规则,一条用户消息包含用户输入。模型返回一条助手消息作为回复。这是最基本的大语言模型API交互模式:**每次调用都是无状态的,所以请求的消息列表必须包含模型所需的所有信息**。\n\n### 带有工具调用的多轮交互:代理的核心循环\n真实的代理工作流通常比单轮问答复杂。当用户问“温哥华当前的时间和天气是什么?”时,模型需要访问动态外部信息:当前时间和最新天气。以下示例逐步展示代理框架与模型之间的每次交互。\n\n\n\n**第一次API调用——代理框架发送初始请求:**\n\nANONYMOUS CHINESE Y:\n### 上下文工程 [第1/17部分]\n# 上下文工程\n\n第1章将上下文定义为代理在决策时刻的工作信息集合。设计和管理该上下文——我们称之为**上下文工程**——是构建有效代理的核心。在实践中,上下文包括模型在给定交互中接收的所有内容:对话历史、系统指令、工具定义、检索到的文档、运行时状态和其他特定任务的信息。从第1章介绍的框架角度来看,上下文工程实现了框架的大部分“上下文和工具”层:它决定了代理在每个决策点看到的信息以及这些信息的组织方式。良好的上下文设计为模型提供正确的背景、约束和操作接口,使其通用推理能力能够有效地应用于任务。\n\n\n\n## 上下文:代理能力的上限\n大型语言模型在标准化基准测试中取得了优异成绩,但在现实世界的商业环境中往往表现不佳。原因很简单:模型能力是通用的,而具体任务依赖于本地知识,如产品架构、业务规则、操作约束和内部约定。这些信息通常不存在于模型的参数中。\n\n设想一位能力很强的工程师加入一个新团队。他们可能有深厚的理论知识和强大的编程能力,但还不了解产品架构、业务逻辑、技术债务或团队规范。如果关键架构决策分散在个人记忆中且代码库文档记录不佳,即使是优秀的工程师也难以快速创造价值。如今的AI代理也面临同样的问题。\n\n以编码代理为例。给定相同的指令“帮我修复这个错误”,代理接收的上下文质量决定了它能否完成任务:\n- **代码上下文**:代码库结构、模块职责、核心数据结构和编码标准。没有这些信息,代理可能生成语法正确但与项目风格或架构不一致的代码。\n- **流程要求**:Git分支策略、提交约定、审查流程和CI/CD要求。没有这些信息,代理可能直接将未经测试的代码提交到主分支。\n- **环境配置**:开发设置、测试数据库连接字符串、暂存部署程序和API密钥管理实践。没有这些信息,本地有效的修复可能在测试环境中立即失败。\n\n这三类——代码、流程和环境——构成了代理有效工作所需的最小上下文。模型的固有能力只是基础;上下文设定了代理能力的上限。具有良好组织上下文的中等能力模型往往能胜过在上下文不足情况下运行的更强模型。\n\n因此,上下文工程是用当今模型构建有效代理的核心。这不仅仅是向提示中添加更多文本的问题。它需要系统地设计、组织和提供模型完成任务所需的背景知识。上下文工程是一个技术问题,但从根本上说是一个组织问题。在许多团队中,关键知识仍然是隐性的:架构决策存在于高级工程师的记忆中,业务规则非正式传递,重要上下文埋藏在私人聊天记录中。如果团队本身是一个糟糕的信息环境,即使是强大的AI代理也会受到限制。\n\n在远程环境中有效工作的团队通常也为AI代理提供有效的环境。像Linux内核这样的开源项目就是有启发性的例子:分布在世界各地的开发者已经维护该项目三十多年。这之所以可行,是因为该项目具有透明的、以文档为驱动的沟通文化。讨论是公开的,决策被记录下来,新人可以通过阅读历史了解代码的演变。同样的工作方式自然创造了一个对AI友好的环境:信息是公开的、可检索的和结构化的。\n\n每次AI代理开始任务时,都将其视为新的团队成员。有了足够的背景,它可以产生高质量的工作;没有该背景,其大部分智能都会被浪费。因此,构建AI原生团队主要是一项文档工作,而不仅仅是部署新工具的问题。\n\nOpenAI研究员翁佳怡明确表达了这一点:**“对人类和模型来说,最重要的是上下文。”** 回顾自己的工作,他指出:“我在OpenAI的工作并不难。如果其他人拥有我所有的上下文,他们也能做到。” 同样的原则适用于代理:代理能力的上限不仅由模型大小决定,还由每个决策点提供的上下文的完整性和精确性决定。翁还观察到团队合作中的核心问题是上下文的不一致,而AI短期内无法取代人类的一个原因是AI和人类不共享相同的环境。上下文工程正是解决这个问题:如何系统地向模型提供代理所需的结构化背景信息。\n\n下一个问题是如何在技术层面将这些上下文信息提供给大语言模型(LLM)。\n\n## 代理如何调用LLM:API级上下文结构\n本节以OpenAI的聊天补全API为例。Anthropic、Google等提供商在细节上有所不同,但它们面向代理的API遵循类似模式:每个模型调用由结构化对话历史和一组可用工具定义构成。理解这种结构是本章后面讨论的上下文工程技术的基础。\n\n### 四种消息角色\n在聊天补全风格的API中,核心输入是一个**消息列表**,通常命名为`messages`。每个消息有一个`role`字段,告诉模型如何解释消息及其来源:\n- **system**:开发者编写的指令,定义代理的身份、行为、约束和工作流程。模型将其视为高优先级指令。在大多数对话中,系统消息在消息列表开头出现一次。\n- **user**:最终用户的输入,代表代理需要处理的请求。\n- **assistant**:之前的模型输出,包括自然语言回复和工具调用请求。在多轮交互中,这些消息包含在后续请求中,以便下一个无状态模型调用可以访问之前的轨迹。\n- **tool**:代理框架执行工具后返回的结果。每个工具结果通过`tool_call_id`与相应的工具调用链接,使模型能够将每个结果与其产生的请求关联起来。\n\n工具定义不是消息。它们在单独的`tools`字段中提供,该字段声明模型可用的工具并指定每个工具接受的参数。\n\n### 单轮请求:最简单的API调用\n\n\n\n从最简单的情况开始:没有工具调用的单轮请求。用户问“你好,你是谁?”示例使用本地部署的Qwen3-0.6B模型,与本节后面的本地LLM部署实验相连。示例中的时间戳仅用于演示,与本书时间线无关。\n\n```javascript\n// ═══ 代理框架构造的请求 ═══\n{\n \"model\": \"Qwen3-0.6B\",\n \"messages\": [\n {\n \"role\": \"system\", // ← 开发者编写\n \"content\": \"You are a helpful coding assistant. Follow user instructions.\"\n },\n {\n \"role\": \"user\", // ← 用户输入\n \"content\": \"Hello, who are you?\"\n }\n ]\n}\n```\n\n```javascript\n// ═══ API返回的响应 ═══\n{\n \"choices\": [{\n \"message\": {\n \"role\": \"assistant\", // ← 模型生成\n \"content\": \"Hi! I'm a coding assistant. I can help you write code, debug issues, and explain technical concepts. How can I help?\"\n }\n }]\n}\n```\n\n此请求仅包含两条消息:一条包含开发者编写规则的系统消息和一条包含用户输入的用户消息。模型返回一条助手消息作为回复。这是最基本的LLM API交互模式:**每次调用都是无状态的,因此请求的消息列表必须包含模型所需的所有信息**。\n\n### 带工具调用的多轮交互:代理的核心循环\n真实的代理工作流通常比单轮问答更复杂。当用户问“温哥华当前的时间和天气是多少?”时,模型需要访问动态外部信息:当前时间和最新天气。以下示例逐步介绍代理框架与模型之间的每次交互。\n\n\n\n**第一次API调用——代理框架发送初始请求:**"
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "0217853906778523ea4fa82759f0cae4f2ffa723d4f9dd6e0b6a7",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390699,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission: 'Context Engineering' heading in X is demoted to '###' instead of '#'. Inconsistency: 'Jiayi Weng' is translated as '翁佳怡' in X but '翁佳怡' in Y (same), but X retains 'LLM' untranslated in some places while Y translates it as '大语言模型(LLM)' consistently.\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '该上下文' (the context) is overused in X, e.g., '设计和管理该上下文' vs. Y's '设计和管理该上下文' (similar, but Y uses '该' more naturally). '产出高质量的工作' in X is less idiomatic than Y's '产生高质量的工作'.\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation: 'Agent' is sometimes '代理' (correct) but '人工智能代理' (redundant) in X; Y consistently uses '代理'. 'Context window' in Figure 2-1 caption is translated as '上下文窗口' in X, which is correct, same as Y.\"},\"markdown_code_fidelity\":{\"score\":3,\"evidence\":\"Heading level error: X uses '### 上下文工程' for the main title instead of '# 上下文工程' as in source. Code blocks and figures are preserved correctly, e.g., JavaScript code and image links like 'images/fig2-1.svg' are intact.\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All headings, figures, and key claims are preserved. 'Jiayi Weng' is correctly translated as '翁佳怡', and 'LLM' is consistently translated as '大语言模型(LLM)' with parenthetical acronym. No omissions in tool role descriptions (e.g., 'tool_call_id' is retained correctly).\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '产生高质量的工作' (produce high-quality work) is more idiomatic than X's '产出高质量的工作'. '对AI友好的环境' (AI-friendly environment) flows better than X's '对人工智能友好的环境'. Sentence structure in '上下文设定了代理能力的上限' is smoother than X's identical phrase (but Y's overall flow is superior).\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent translation: 'Agent' → '代理', 'Context Engineering' → '上下文工程', 'tool_call_id' → 'tool_call_id' (correctly untranslated). 'Git branching strategy' → 'Git分支策略' (consistent with tech standards). 'API-level context structure' → 'API级上下文结构' (accurate).\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Headings match source levels: '# 上下文工程' (main title), '## 上下文:代理能力的上限' (correct level 2). Code blocks with JavaScript and comments are preserved exactly, e.g., '// ═══ 代理框架构造的请求 ═══' is identical. Figure links like 'images/fig2-2.svg' are unchanged.\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no heading level errors), superior fluency (more idiomatic phrasing), consistent terminology, and perfect markdown fidelity (correct heading hierarchy), whereas X has heading level mistakes and inconsistent terminology.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no heading level errors), superior fluency (more idiomatic phrasing), consistent terminology, and perfect markdown fidelity (correct heading hierarchy), whereas X has heading level mistakes and inconsistent terminology.\"}",
|
||
"usage": {
|
||
"prompt_tokens": 5741,
|
||
"completion_tokens": 786,
|
||
"total_tokens": 6527
|
||
}
|
||
},
|
||
"latency_milliseconds": 21587,
|
||
"validation": {
|
||
"valid": true
|
||
}
|
||
}
|
||
]
|
||
}
|