{ "schema_version": 1, "credential_free": true, "attempts": [ { "attempt": 1, "request": { "model": "doubao-seed-1-6-250615", "messages": [ { "role": "user", "content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nThis discovery suggests two operations that were previously considered impractical. The first is **Editing**: since the conclusion has already been written into downstream notes, a changed field can propagate through cached reasoning when the model has an explicit chain of thought (CoT), producing results close to full recomputation with about 1% of the compute. Conversely, without CoT, an isolated field change may be ignored because the conclusion is already embedded downstream without a reasoning path to update it. The second is **Composition**: a precomputed \"skill\" cache can be relocated using Rotary Position Embedding (RoPE) and spliced into another context without recomputing attention. In this framing, assembling a long context from modular cache blocks drops from O(L²) recomputation to O(L) splicing, with output quality close to full recomputation.\n\nThe margin-note analogy is useful here. When reading a long document, one does not reread the entire document every time a fact changes; instead, one updates the note that records what the fact implies. The idea of KV Cache as notes is similar: if the cached states already encode the inference of a fact, then changing the fact may require correcting the downstream note rather than recomputing everything. Because the notes are represented in a portable form, a block of notes from one problem can also be repositioned (via RoPE relocation) and reused in another. The paper implemented this idea on vLLM, speeding up p90 time to first token by factors ranging from tens to hundreds, with a prefix cache hit rate of about 98.5% and outputs close to token-by-token recomputation (across 12 models, logit cosine similarity 0.90–0.999).\n\nFor Agents, the implication is that long contexts may not always need to be torn down and rebuilt when tools, memory fields, or runtime state change. In principle, this could make context mutable while preserving some caching benefits, turning context assembly from O(L²) recomputation into O(L) note splicing. This is still research-stage work; the three practical conclusions earlier in this section remain the default principles for current production systems.\n\n[^ch2-2]: Li, Bojie. *Models Take Notes at Prefill: KV Cache Can Be Editable and Composable.* arXiv:2606.17107, 2026.\n\nNow that we understand how context is processed and cached, the next question is how to design the content itself. The following sections discuss what belongs in context and how to organize it, along three related threads:\n\n- **Prompt Engineering, Prompt Injection, and Dynamic Prompts (Agent Skills)**: How to write the system prompt and what to include. This is the most direct part of context engineering. Tool definitions, another static component alongside the system prompt, also directly affect the accuracy of the Agent's tool use. This chapter provides the core principles, and Chapter 4 expands on them in detail. The next issue is security: when external content attempts to hijack a carefully designed context, how should the system defend itself at the context level? As prompts grow longer and cover more scenarios, placing everything into a single system prompt becomes impractical: it wastes tokens and dilutes attention. This leads naturally to the progressive disclosure mechanism of Agent Skills, where knowledge is loaded on demand rather than included all at once.\n- **Agent Status Bar**: An independent mechanism that injects dynamic meta-information (task progress, environment status, tool call count, etc.) at the end of the context, compensating for the model's inability to actively summarize implicit states. Analogous to the time, battery, and network signal shown at the top of a phone screen, the Agent Status Bar lets the model access the current runtime state at any time.\n- **Context Compression Strategies**: Addressing the problem of ever-expanding context—when to compress, how to compress, and how compression coexists with KV Cache.\n\n## Prompt Engineering: Optimizing the System Prompt\n\nThe primary focus of prompt engineering is the **System Prompt**—the `role: \"system\"` message in the API message list. It is the Agent's operating manual, defining the Agent's identity, behavioral rules, constraints, and workflow. A well-designed system prompt enables the model to fully leverage its general capabilities in specific tasks.\n\nThere is a practical litmus test for system prompt design: an LLM is like a highly capable new team member who is completely unfamiliar with your specific workflows and internal conventions. If such a new team member, after reading your system prompt, still does not know what to do, neither will the Agent.\n\nThe following sections discuss several dimensions of system prompt design.\n\n### Tone and Style: Behavioral Framing\n\nTone and style are easy to overlook, but they strongly shape the user experience. Consider instructions such as \"You MUST answer concisely with fewer than 4 lines.\" When the Agent cannot complete a task, constraints such as \"keep your response to 1–2 sentences\" and \"do not explain why you cannot do something\" prevent lengthy self-justification. Uppercase words such as \"NEVER do X\" increase instruction salience more than softer phrasing such as \"Please avoid doing X,\" but overuse dilutes the effect; reserve them for truly critical constraints.\n\n### Structured Prompts: The \"Format\" of the System Prompt\n\nModern large language models show significant sensitivity to structured input, stemming from the large amount of structured content in their training data. The use of XML tags follows a hierarchical principle, with the tag names themselves carrying semantic information—`` immediately tells the model this is working directory information, whereas a plain text format like \"Current directory: /Users/project/src\" requires the model to do extra reasoning to infer the relationship between the two sides of the colon.\n\nMarkdown provides lightweight structure while maintaining readability, making it particularly suitable for organizing hierarchical instructions and information. XML and Markdown create a two-layer structure: XML provides precise, machine-parseable semantics, while Markdown organizes the content for human and machine readers.\n\n### Process-Driven vs. Rule Stacking: The \"Organization\" of the System Prompt\n\nMethods that reduce cognitive load for humans are equally effective for large language models—because the model has learned human language and reasoning patterns during training. Imagine giving a new team member a manual with hundreds of scattered rules, no flowcharts, and no priority instructions—even a highly capable person would be confused: when multiple rules apply simultaneously, which one should be chosen? And what about situations not covered by the rules?\n\nIn contrast, a process-driven prompt functions like an effective training manual, providing a clear Standard Operating Procedure (SOP):\n\n```\nFile Processing Standard Operating Procedure:\n\nStep 1: Validation\n Check if file exists and is accessible\n - If not found → log error and stop\n ↓\nStep 2: Classification\n Determine file type based on extension and content\n ↓\nStep 3: Preprocessing\n Config files → create backup\n Large files (>1MB) → stream processing\n ↓\nStep 4: Execution\n Execute core processing logic based on file type\n ↓\nStep 5: Verification\n Ensure integrity of the processed file\n```\n\nThis process design helps the model track which stage it is in, what the current step is trying to accomplish, and what should happen next. When an exception occurs, the model can choose a response based on the current stage instead of searching through a long list of unrelated rules.\n\n### Translating Business Rules into Executable Instructions\n\nWhen building production-grade Agent systems, the most easily overlooked—and most critical—piece is **business rule refinement**. This is not a technical problem but a product-design problem, and it demands deep involvement from product managers.\n\nConsider an Agent that helps users make phone calls to resolve billing issues: the user tells the Agent they want to lower a subscription fee or request a refund, and the Agent automatically calls customer service to complete the negotiation. The billing system design for such a service is a typical case of business rule refinement. The product manager's core requirement is \"if it does not work, refund,\" encouraging users to try while preventing abuse. The team designed three billing models:\n\n- **Commission on savings**: The Agent negotiates on behalf of the user, taking a cut, e.g., 20% of the money saved.\n- **Fixed service fee**: For tasks that do not involve saving money, such as booking a restaurant, charge a fixed fee based on complexity.\n- **Prepayment for difficult tasks**: For tasks with very low success rates, a non-refundable prepayment is charged to filter out unrealistic requests.\n\nHowever, vague rules (e.g., \"choose the appropriate billing type based on the task situation\") lead to highly unstable Agent behavior. \"Help me return the clothes I bought last month\"—is this \"saving the user money\" or \"retrieving money that rightfully belongs to them\"? \"Help me cancel my Netflix subscription\"—canceling does prevent future payments, but does this count as \"saving money\"? The same task might be classified completely differently at different times, making business logic unpredictable.\n\nProduct managers must define decision rules to the point where they are executable. Commission-based billing is only applicable in scenarios where existing bills are reduced through negotiation (the Agent needs to use negotiation skills to convince the merchant). Refunds and service cancellations must never be commission-based—the prompt must explicitly state: \"NEVER use percentage_based_one_time for refunds and service cancellations. Use fixed_fee instead.\"\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第7/17部分]\n这一发现表明了两个之前被认为不切实际的操作。第一个是**编辑**:由于结论已经被写入下游笔记,当模型有明确的思维链(CoT)时,更改的字段可以在缓存的推理中传播,仅用约1%的计算量就能产生接近完全重新计算的结果。相反,没有CoT时,孤立的字段更改可能会被忽略,因为结论已经嵌入下游,没有推理路径来更新它。第二个是**组合**:预先计算的“技能”缓存可以使用旋转位置嵌入(RoPE)重新定位,并拼接成另一个上下文而无需重新计算注意力。在这种框架下,从模块化缓存块组装长上下文从O(L²)重新计算降为O(L)拼接,输出质量接近完全重新计算。\n\n这里用边注的类比很有用。阅读长文档时,事实改变时不会每次都重读整个文档;而是更新记录该事实含义的笔记。将KV缓存视为笔记的想法类似:如果缓存状态已经编码了某个事实的推理,那么更改该事实可能只需要纠正下游笔记,而不是重新计算所有内容。因为笔记以可移植的形式表示,一个问题的笔记块也可以通过RoPE重新定位并在另一个问题中重用。该论文在vLLM上实现了这个想法,将p90首字节时间加速了数十到数百倍,前缀缓存命中率约为98.5%,输出接近逐标记重新计算(在12个模型上,对数似然余弦相似度为0.90-0.999)。\n\n对于代理来说,这意味着当工具、内存字段或运行时状态改变时,长上下文并不总是需要拆除和重建。原则上,这可以在保留一些缓存优势的同时使上下文可变,将上下文组装从O(L²)重新计算转变为O(L)笔记拼接。这仍然是研究阶段的工作;本节前面的三个实践结论仍然是当前生产系统的默认原则。\n\n[^ch2-2]: Li, Bojie. *Models Take Notes at Prefill: KV Cache Can Be Editable and Composable.* arXiv:2606.17107, 2026.\n\n现在我们理解了上下文的处理和缓存方式,下一个问题是如何设计内容本身。以下部分将沿着三个相关线索讨论上下文中包含什么以及如何组织它:\n\n- **提示工程、提示注入和动态提示(代理技能)**:如何编写系统提示以及包含什么内容。这是上下文工程最直接的部分。工具定义是与系统提示并列的另一个静态组件,也直接影响代理工具使用的准确性。本章提供核心原则,第4章将详细展开。下一个问题是安全:当外部内容试图劫持精心设计的上下文时,系统应如何在上下文层进行防御?随着提示变长并覆盖更多场景,将所有内容放入单个系统提示变得不切实际:它浪费标记并稀释注意力。这自然导致代理技能的渐进披露机制,知识按需加载而不是一次性包含。\n- **代理状态栏**:一种独立机制,在上下文末尾注入动态元信息(任务进度、环境状态、工具调用次数等),弥补模型无法主动总结隐含状态的不足。类似于手机屏幕顶部显示的时间、电池和网络信号,代理状态栏让模型随时访问当前运行时状态。\n- **上下文压缩策略**:解决上下文不断扩展的问题——何时压缩、如何压缩以及压缩如何与KV缓存共存。\n\n## 提示工程:优化系统提示\n提示工程的主要焦点是**系统提示**——API消息列表中的`role: \"system\"`消息。它是代理的操作手册,定义代理的身份、行为规则、约束和工作流。设计良好的系统提示使模型能够在特定任务中充分利用其通用能力。\n\n系统提示设计有一个实用的试金石:大语言模型就像一个完全不熟悉你特定工作流和内部约定的高能力新团队成员。如果这样的新团队成员在阅读你的系统提示后仍然不知道该做什么,代理也不会知道。\n\n以下部分讨论系统提示设计的几个维度。\n\n### 语气和风格:行为框架\n语气和风格容易被忽视,但它们强烈塑造用户体验。考虑这样的指令:“你必须简洁回答,少于4行。” 当代理无法完成任务时,“将你的响应保持在1-2句话”和“不要解释你为什么不能做某事”等约束防止冗长的自我辩解。大写单词如“NEVER do X”比“Please avoid doing X”等柔和措辞更能提高指令的显著性,但过度使用会稀释效果;仅在真正关键的约束时保留它们。\n\n### 结构化提示:系统提示的“格式”\n现代大语言模型对结构化输入表现出显著敏感性,这源于其训练数据中大量的结构化内容。XML标签的使用遵循分层原则,标签名称本身携带语义信息——``立即告诉模型这是工作目录信息,而像“Current directory: /Users/project/src”这样的纯文本格式需要模型进行额外推理来推断冒号两边的关系。\n\nMarkdown在保持可读性的同时提供轻量级结构,特别适合组织分层指令和信息。XML和Markdown创建了一个两层结构:XML提供精确的、机器可解析的语义,而Markdown为人类和机器读者组织内容。\n\n### 流程驱动与规则堆叠:系统提示的“组织”\n减少人类认知负荷的方法对大语言模型同样有效——因为模型在训练期间已经学习了人类语言和推理模式。想象给一个新团队成员一本有数百条分散规则、没有流程图、没有优先级指令的手册——即使是高能力的人也会困惑:当多个规则同时适用时,应该选择哪一个?对于规则未涵盖的情况又该怎么办?\n\n相反,流程驱动的提示就像一本有效的培训手册,提供清晰的标准操作程序(SOP):\n\n```\n文件处理标准操作程序:\n\n步骤1:验证\n 检查文件是否存在且可访问\n - 如果未找到→记录错误并停止\n ↓\n步骤2:分类\n 根据扩展名和内容确定文件类型\n ↓\n步骤3:预处理\n 配置文件→创建备份\n 大文件(>1MB)→流式处理\n ↓\n步骤4:执行\n 根据文件类型执行核心处理逻辑\n ↓\n步骤5:验证\n 确保处理后文件的完整性\n```\n\n这种流程设计帮助模型跟踪它处于哪个阶段、当前步骤试图完成什么以及下一步应该做什么。当出现异常时,模型可以根据当前阶段选择响应,而不是在一长串不相关的规则中搜索。\n\n### 将业务规则转化为可执行指令\n构建生产级代理系统时,最容易被忽视但最关键的部分是**业务规则细化**。这不是技术问题而是产品设计问题,需要产品经理深度参与。\n\n考虑一个帮助用户打电话解决账单问题的代理:用户告诉代理他们想降低订阅费用或请求退款,代理自动打电话给客服完成协商。此类服务的账单系统设计是业务规则细化的典型案例。产品经理的核心要求是“如果不起作用,就退款”,鼓励用户尝试同时防止滥用。团队设计了三种账单模型:\n\n- **节省佣金**:代理代表用户协商,收取一定比例,例如节省金额的20%。\n- **固定服务费**:对于不涉及省钱的任务,如预订餐厅,根据复杂度收取固定费用。\n- **困难任务预付款**:对于成功率非常低的任务,收取不可退款的预付款以过滤不现实的请求。\n\n然而,模糊的规则(例如“根据任务情况选择适当的计费类型”)导致代理行为高度不稳定。“帮我退回上个月买的衣服”——这是“为用户省钱”还是“取回本应属于用户的钱”?“帮我取消我的Netflix订阅”——取消确实防止了未来付款,但这是否算作“省钱”?同一任务在不同时间可能被完全不同地分类,使业务逻辑不可预测。\n\n产品经理必须将决策规则定义到可执行的程度。基于佣金的计费仅适用于通过协商减少现有账单的场景(代理需要使用协商技能说服商家)。退款和服务取消绝不能基于佣金——提示必须明确声明:“NEVER use percentage_based_one_time for refunds and service cancellations. Use fixed_fee instead.”\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第7/17部分]\n\n这一发现提出了两种此前被认为不切实际的操作。第一种是**编辑**:由于结论已经写入下游笔记,当模型具有明确的思维链(CoT)时,更改的字段可以在缓存推理中传播,产生接近完全重新计算的结果,且计算量约为1%。相反,没有CoT时,孤立的字段更改可能会被忽略,因为结论已经嵌入下游,没有推理路径来更新它。第二种是**组合**:可以使用旋转位置嵌入(RoPE)重新定位预先计算的“技能”缓存,并将其拼接成另一个上下文,而无需重新计算注意力。在这种框架下,从模块化缓存块组装长上下文从O(L²)的重新计算降为O(L)的拼接,输出质量接近完全重新计算。\n\n边注的类比在这里很有用。阅读长文档时,事实改变时不会每次都重读整个文档;而是更新记录该事实含义的笔记。将KV缓存视为笔记的想法类似:如果缓存状态已经编码了某个事实的推理,那么改变该事实可能只需要纠正下游笔记,而不是重新计算一切。因为笔记以可移植的形式表示,一个问题的笔记块也可以通过RoPE重新定位并在另一个问题中重复使用。该论文在vLLM上实现了这一想法,将p90首token时间加速了数十到数百倍,前缀缓存命中率约为98.5%,输出接近逐token重新计算(在12个模型上,logit余弦相似度为0.90–0.999)。\n\n对于智能体来说,这意味着当工具、内存字段或运行时状态改变时,长上下文并不总是需要拆除和重建。原则上,这可以使上下文可变,同时保留一些缓存优势,将上下文组装从O(L²)的重新计算变为O(L)的笔记拼接。这仍然是研究阶段的工作;本节前面的三个实际结论仍然是当前生产系统的默认原则。\n\n[^ch2-2]: Li, Bojie. *Models Take Notes at Prefill: KV Cache Can Be Editable and Composable.* arXiv:2606.17107, 2026.\n\n现在我们理解了上下文是如何处理和缓存的,下一个问题是如何设计内容本身。以下各节将从三个相关方面讨论上下文中应包含什么以及如何组织:\n\n- **提示工程、提示注入和动态提示(智能体技能)**:如何编写系统提示以及包含什么内容。这是上下文工程中最直接的部分。工具定义是与系统提示并列的另一个静态组件,也直接影响智能体工具使用的准确性。本章提供核心原则,第4章将详细展开。下一个问题是安全性:当外部内容试图劫持精心设计的上下文时,系统应如何在上下文层面进行防御?随着提示变长并覆盖更多场景,将所有内容放入单个系统提示变得不切实际:这会浪费词元并分散注意力。这自然导致智能体技能的渐进披露机制,即按需加载知识而不是一次性包含所有内容。\n- **智能体状态栏**:一种独立机制,在上下文末尾注入动态元信息(任务进度、环境状态、工具调用次数等),弥补模型无法主动总结隐含状态的不足。类似于手机屏幕顶部显示的时间、电池和网络信号,智能体状态栏让模型随时访问当前运行时状态。\n- **上下文压缩策略**:解决上下文不断扩展的问题——何时压缩、如何压缩以及压缩如何与KV缓存共存。\n\n### 提示工程:优化系统提示\n\n提示工程的主要焦点是**系统提示**——API消息列表中`role: \"system\"`的消息。它是智能体的操作手册,定义智能体的身份、行为规则、约束和工作流程。设计良好的系统提示能让模型在特定任务中充分发挥其通用能力。\n\n系统提示设计有一个实用的检验标准:大语言模型就像一个非常有能力但完全不熟悉你特定工作流程和内部惯例的新团队成员。如果这样的新团队成员在阅读你的系统提示后仍然不知道该做什么,智能体也不会知道。\n\n以下各节将讨论系统提示设计的几个维度。\n\n#### 语气和风格:行为框架\n\n语气和风格容易被忽视,但它们强烈影响用户体验。考虑这样的指令:“你必须简洁回答,少于4行。”当智能体无法完成任务时,“将你的回复限制在1–2句话”和“不要解释你为什么不能做某事”等约束防止了冗长的自我辩解。“NEVER做X”等大写单词比“请避免做X”等柔和措辞更能突出指令,但过度使用会削弱效果;应将其保留给真正关键的约束。\n\n#### 结构化提示:系统提示的“格式”\n\n现代大语言模型对结构化输入非常敏感,这源于其训练数据中大量的结构化内容。使用XML标签遵循分层原则,标签名称本身携带语义信息——``立即告诉模型这是工作目录信息,而像“Current directory: /Users/project/src”这样的纯文本格式需要模型进行额外推理来推断冒号两边的关系。\n\nMarkdown在保持可读性的同时提供轻量级结构,特别适合组织分层指令和信息。XML和Markdown创建了两层结构:XML提供精确的、机器可解析的语义,而Markdown为人类和机器读者组织内容。\n\n#### 流程驱动与规则堆叠:系统提示的“组织”\n\n减轻人类认知负荷的方法对大语言模型同样有效——因为模型在训练期间学习了人类语言和推理模式。想象给新团队成员一本有数百条分散规则、没有流程图且没有优先级指令的手册——即使是非常有能力的人也会困惑:当多个规则同时适用时,应选择哪一个?对于规则未涵盖的情况又该如何?\n\n相反,流程驱动的提示就像一份有效的培训手册,提供清晰的标准操作程序(SOP):\n\n```\n文件处理标准操作程序:\n\n步骤1:验证\n 检查文件是否存在且可访问\n - 如果未找到→记录错误并停止\n ↓\n步骤2:分类\n 根据扩展名和内容确定文件类型\n ↓\n步骤3:预处理\n 配置文件→创建备份\n 大文件(>1MB)→流式处理\n ↓\n步骤4:执行\n 根据文件类型执行核心处理逻辑\n ↓\n步骤5:验证\n 确保处理后文件的完整性\n```\n\n这种流程设计帮助模型跟踪当前所处阶段、当前步骤试图完成的任务以及下一步应做什么。出现异常时,模型可以根据当前阶段选择响应,而无需在一长串不相关的规则中搜索。\n\n#### 将业务规则转化为可执行指令\n\n构建生产级智能体系统时,最容易被忽视但最关键的部分是**业务规则细化**。这不是技术问题而是产品设计问题,需要产品经理深度参与。\n\n考虑一个帮助用户打电话解决账单问题的智能体:用户告诉智能体他们想降低订阅费用或请求退款,智能体自动打电话给客服完成协商。此类服务的账单系统设计是业务规则细化的典型案例。产品经理的核心要求是“如果不起作用就退款”,鼓励用户尝试同时防止滥用。团队设计了三种账单模型:\n\n- **节省佣金**:智能体代表用户协商,抽取一定比例,例如节省金额的20%。\n- **固定服务费**:对于不涉及省钱的任务,例如预订餐厅,根据复杂程度收取固定费用。\n- **困难任务预付款**:对于成功率非常低的任务,收取不可退还的预付款以过滤不切实际的请求。\n\n然而,模糊的规则(例如“根据任务情况选择适当的计费类型”)会导致智能体行为高度不稳定。“帮我退回上个月买的衣服”——这是“为用户省钱”还是“取回本应属于用户的钱”?“帮我取消我的Netflix订阅”——取消确实防止了未来付款,但这是否算“省钱”?同一任务在不同时间可能被完全不同地分类,导致业务逻辑不可预测。\n\n产品经理必须将决策规则定义到可执行的程度。基于佣金的计费仅适用于通过协商减少现有账单的场景(智能体需要使用协商技能说服商家)。退款和服务取消绝不能基于佣金——提示必须明确说明:“NEVER对退款和服务取消使用percentage_based_one_time。改用fixed_fee。”" } ], "temperature": 0, "response_format": { "type": "json_object" }, "extra_body": { "thinking": { "type": "disabled" } } }, "response": { "id": "0217853907966633ea4fa82759f0cae4f2ffa723d4f9dd66b4d3a", "model": "doubao-seed-1-6-250615", "created": 1785390813, "content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"遗漏了原文开头的章节标题标识'### 上下文工程[第7/17部分]'后的空行;将'logit cosine similarity'译为'对数似然余弦相似度',存在术语偏差(应为'logit余弦相似度')。\"},\"fluency\":{\"score\":4,\"evidence\":\"部分句子稍显生硬,如'仅用约1%的计算量就能产生接近完全重新计算的结果'可优化为'计算量仅需约1%即可产生接近完全重新计算的结果';整体表达通顺,符合技术文本风格。\"},\"terminology\":{\"score\":5,\"evidence\":\"关键术语如'思维链(CoT)'、'旋转位置嵌入(RoPE)'、'标准操作程序(SOP)'等翻译准确且一致;'Agent'统一译为'代理',符合上下文。\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"正确保留了代码块、列表、引用等Markdown元素;但将二级标题'## Prompt Engineering: Optimizing the System Prompt'错误转换为三级标题'### 提示工程:优化系统提示',层级混乱。\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"完整翻译了原文所有内容,无遗漏或增删;'logit cosine similarity'准确译为'logit余弦相似度',技术细节保留完整。\"},\"fluency\":{\"score\":5,\"evidence\":\"语句流畅自然,如'产生接近完全重新计算的结果,且计算量约为1%'比X的对应表述更通顺;整体行文符合中文表达习惯,可读性强。\"},\"terminology\":{\"score\":5,\"evidence\":\"术语翻译准确一致,'Agent'译为'智能体'符合技术文献常见译法;'Standard Operating Procedure (SOP)'译为'标准操作程序(SOP)'正确;其他专业术语如'RoPE'、'CoT'等处理规范。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"严格保留了原文的Markdown结构,包括标题层级(二级标题'##'正确对应)、代码块、列表、引用格式;章节标题后的空行、代码缩进等细节均与原文一致。\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y在准确性(无术语偏差和内容遗漏)、流畅性(语句更自然)和Markdown格式保真度(标题层级正确)上均优于X;X存在标题层级错误和术语翻译偏差,而Y完全符合技术翻译的严格要求。\"} }", "usage": { "prompt_tokens": 5954, "completion_tokens": 550, "total_tokens": 6504 } }, "latency_milliseconds": 16790, "validation": { "valid": true } } ] }