Files
ai-agent-book/chapter10/book-translation/validation/real_20260730T061500Z_v4/judge_receipts/unit-25.json
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

44 lines
29 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
{
"schema_version": 1,
"credential_free": true,
"attempts": [
{
"attempt": 1,
"request_kind": "quality_judgment",
"request": {
"model": "doubao-seed-1-6-250615",
"messages": [
{
"role": "user",
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\n> **Experiment 2-9 ★★★: Comparison of Context Compression Strategies**\n>\n> We designed a research task: identify and track the employment status of OpenAI co-founders. This task requires multi-step information aggregation, the length of search results varies greatly (from a few thousand to over a hundred thousand characters), and there are clear success criteria. Using Kimi K3 (a reasoning model with a native context of about 1 million tokens; this experiment deliberately limited the context budget to a 128K window to trigger compression), we implemented six strategies:\n>\n> **Strategy 1: No Compression** — All original results from tool calls are kept intact. Multiple searches returned a total of approximately 367,000 characters (7 tool calls, averaging about 52,000 characters each). By the fifth iteration, the cumulative context exceeded the 128K limit (approximately 165,000 tokens), triggering overflow protection and causing task failure. Just a few searches were enough to exhaust the 128K window.\n>\n> **Strategies 2 & 3: Non-Task-Aware Compression** — Individual Summarization generates a 23 paragraph summary for each search result independently, with a compression ratio of 10.9% (in this book, compression ratio refers to \"compressed volume / original volume\"; a smaller number means more aggressive compression). It can complete the task but requires 12 iterations and 276,608 tokens. The main problem is information fragmentation—multiple pages repeatedly describe the same event, wasting context space. Combined Summarization merges all results into a single comprehensive summary, with a compression ratio of 4.3%, requiring 10 iterations and 93,449 tokens. However, when the input is extremely long, it must be truncated, potentially losing information at the end. The common flaw of both is a lack of semantic understanding, making it impossible to distinguish the relevance of information.\n>\n> **Strategy 4: Context-Aware Compression** — The core innovation is incorporating the current query intent and accumulated information into the compression decision process. By specifying \"Given the search query: {query}\" and \"Current context: {context}\" in the compression prompt, the model is guided to generate targeted summaries. The result requires only 7 iterations and 40,157 tokens, with an overall compression ratio of about 3.0%. In one compression instance, compressing 147,877 characters to 1,963 characters (about 1.3%) still retained key information like founder names and position changes; subsequent searches could intelligently extract key information like position changes and new companies, filtering out irrelevant historical background and duplicate content. This success is based on a key insight: in multi-step tasks, the required information density and type vary at different stages—early stages need broad information gathering, middle stages need precise fact verification, and later stages need comprehensive information synthesis. Context-aware compression maximizes information value by dynamically adjusting the focus of compression.\n>\n> **Strategy 5: Context-Aware with Citations** — Adds information provenance to intelligent compression, with each fact accompanied by a source URL citation marker. Token usage increases to 222,992, with a compression ratio of 4.1%, but the citations enable verification. This combines lossy semantic compression with lossless indexing: although the content is compressed, retained source links allow the system to return to the original material.\n>\n> **Strategy 6: Adaptive Windowing** — Based on a key insight: early in the task, context space is abundant, so there is no need to rush compression. The compression mechanism is only activated when approaching the capacity limit, thereby preserving the integrity of the original information as much as possible. The specific implementation includes three core mechanisms:\n>\n> - **Threshold Trigger**: Continuously monitors context usage. Compression is activated only when the prompt token count exceeds 80% of the window (102,400 tokens for a 128K window).\n> - **Batch Compression**: When triggered, compresses all unmarked tool results at once. For example, around the fourth iteration, when the context is detected to exceed the 102,400 token threshold (triggered at approximately 135,600 tokens in practice), all 10 uncompressed tool messages are compressed immediately.\n> - **Duplicate Prevention**: Adds a `[COMPRESSED]` marker to ensure compressed content is never processed again.\n>\n> Although the total token usage is relatively high (174,601), the first few iterations retain the complete original information, providing maximum flexibility for broad initial information gathering.\n>\n>\n> ![Figure 2-17: Processing Flow of Six Compression Strategies](images/fig2-17.svg)\n>\n>\n\n### Production-Grade Hierarchical Compression Mechanism\n\nThe experiment above demonstrates the performance differences among compression strategies. In production, mature Agent systems typically do not rely on a single strategy. Instead, they combine multiple strategies into a hierarchical compression mechanism. Different types of information remain useful for different lengths of time, so the compression strategy should match the expected lifecycle of the information. Using Claude Code's approach as a reference, a mature context management system usually includes five layers:\n\n1. **Tool Result Budget Control**: Large tool outputs are stored on disk; the model only sees a preview summary. Replacement decisions are frozen once made to ensure cache consistency.\n2. **Direct Noise Deletion**: Low-value content (e.g., content from a large set of search results that was only used for a few lines) is removed without summarization—summarizing noise wastes tokens.\n3. **API-Level Micro-Compression**: Leverages the API's context editing capabilities to instruct the server to remove specific tool results from the prefix, while the local message list remains unchanged. The advantage of this layer is zero local implementation cost—the server handles it in one pass. However, according to the prefix invariance principle in this chapter, the cache after the removal point will also become invalid, requiring a cache rebuild. Therefore, it is suitable for use when the context is about to overflow and the cost of rebuilding the cache must be paid anyway, rather than being triggered frequently.\n4. **Archival Summarization**: Performs structured summarization round by round (like `git log`, retaining an independent record for each round, rather than `git squash` which merges them into one), preserving the logical thread of the conversation.\n5. **Full Compression**: LLM-driven complete compression, used as a last resort. Even this is done in two stages: first, try to compress the session memory; if that fails, perform full compression. Full compression is also equipped with a circuit breaker for consecutive failures (a mechanism that automatically stops retrying after a certain number of consecutive failures)—production data shows that many sessions get stuck in loops of repeated compression failures, and the circuit breaker prevents unnecessary spending on these sessions.\n\nThe order of these five layers matters. The first three have the lowest implementation cost and the most controllable impact on the cache, so they should be used first. The last two have higher costs but stronger compression effects, and should serve as fallback methods.\n\n### Design Principles for Compression Strategies\n\nWe have already analyzed the two motivations for compression—controlling length and improving reasoning quality—and the internal mechanism by which “in-context learning is essentially retrieval.” On that basis, we can distill four principles to guide the design of specific compression strategies. The compression discussed here serves the current task; when trajectories from multiple tasks must be consolidated offline into persistent experience, the problem becomes one of continuous evolution, as discussed in Chapter 8.\n\n- **Non-Uniform Distribution of Information Value**: Key decision points, such as personnel lists, have greater value than supporting evidence, such as news details; supporting evidence, in turn, has greater value than redundant noise, such as navigation bars and footer ads.\n- **Semantic Integrity**: \"Sutskever left OpenAI in May 2024\" cannot be compressed to \"Sutskever left\"—the time and company name are critical, non-negotiable information.\n- **Task Relevance**: The same content should yield different compression results for different tasks, such as \"find the list of founders\" versus \"learn about personal background.\"\n- **Compression is Understanding**: Effective compression requires deep semantic understanding—capturing the core meaning of the context with more refined expression. Moreover, the results of explicit compression are reviewable and reusable across sessions.\n\n### Implications for Agent Architecture Design\n\nResearch on context compression strategies points to fundamental issues in Agent system design. **Compression is Understanding**: the module responsible for compression needs language understanding capabilities close to those of the main model, forming a recursive model-call architecture. **Compression Strategy is Coupled with Task Type**: information retrieval tasks need to preserve breadth, analysis tasks need to preserve depth, and creative tasks need to preserve inspiration triggers. Future Agents should be able to select compression strategies adaptively based on the task type.\n\nAlthough compression adds computational overhead because each compression requires an extra LLM call, its return on investment can be extremely high relative to the resulting token-cost savings and improvements in task success. Experiments show that context-aware compression reduces token usage by over 75%.\n\n\n\nANONYMOUS CHINESE X:\n### 上下文工程[第16/17部分]\n\n#### 实验2-9 ★★★:上下文压缩策略比较\n我们设计了一个研究任务:识别并追踪OpenAI联合创始人的任职状态。该任务需要多步骤信息聚合,搜索结果长度差异极大(从几千到超过十万字符),且有明确的成功标准。使用Kimi K3(原生上下文约100万个词元的推理模型;本实验故意将上下文预算限制在128K窗口以触发压缩),我们实施了六种策略:\n\n**策略1:不压缩**——工具调用的所有原始结果完整保留。多次搜索共返回约367,000字符(7次工具调用,每次平均约52,000字符)。到第五次迭代时,累积上下文超过128K限制(约165,000词元),触发溢出保护并导致任务失败。只需几次搜索就耗尽了128K窗口。\n\n**策略2和3:非任务感知压缩**——个体总结为每个搜索结果独立生成2 - 3段摘要,压缩比为10.9%(本书中压缩比指“压缩量/原始量”;数值越小表示压缩越激进)。它可以完成任务,但需要12次迭代和276,608词元。主要问题是信息碎片化——多个页面重复描述同一事件,浪费上下文空间。联合总结将所有结果合并为一个综合摘要,压缩比为4.3%,需要10次迭代和93,449词元。然而,当输入极长时必须截断,可能丢失末尾信息。两者的共同缺陷是缺乏语义理解,无法区分信息的相关性。\n\n**策略4:上下文感知压缩**——核心创新是将当前查询意图和累积信息纳入压缩决策过程。在压缩提示中指定“给定搜索查询:{query}”和“当前上下文:{context}”,引导模型生成针对性摘要。结果仅需7次迭代和40,157词元,总体压缩比约为3.0%。在一次压缩实例中,将147,877字符压缩到1,963字符(约1.3%)仍保留了创始人姓名和职位变动等关键信息;后续搜索可智能提取职位变动和新公司等关键信息,过滤掉不相关的历史背景和重复内容。这一成功基于关键洞察:在多步骤任务中,不同阶段所需信息密度和类型不同——早期需要广泛收集信息,中期需要精确事实验证,后期需要综合信息合成。上下文感知压缩通过动态调整压缩焦点最大化信息价值。\n\n**策略5:带引用的上下文感知**——在智能压缩中添加信息出处,每个事实伴随源URL引用标记。词元使用增加到222,992,压缩比为4.1%,但引用便于验证。这结合了有损语义压缩和无损索引:虽然内容被压缩,但保留的源链接允许系统返回原始材料。\n\n**策略6:自适应窗口**——基于关键洞察:任务早期上下文空间充裕,无需急于压缩。仅在接近容量限制时激活压缩机制,从而尽可能保留原始信息的完整性。具体实现包括三个核心机制:\n- **阈值触发**:持续监控上下文使用情况。仅当提示词元数超过窗口的80%(128K窗口为102,400词元)时激活压缩。\n- **批量压缩**:触发时一次性压缩所有未标记的工具结果。例如,在第四次迭代左右,检测到上下文超过102,400词元阈值(实际在约135,600词元时触发),立即压缩所有10条未压缩的工具消息。\n- **重复预防**:添加`[COMPRESSED]`标记,确保压缩内容不再被处理。\n\n尽管总词元使用量相对较高(174,601),但前几次迭代保留了完整的原始信息,为初始广泛收集信息提供了最大灵活性。\n\n![图2-17:六种压缩策略的处理流程](images/fig2-17.svg)\n\n#### 生产级分层压缩机制\n上述实验展示了压缩策略之间的性能差异。在生产环境中,成熟的Agent系统通常不依赖单一策略,而是将多种策略组合成分层压缩机制。不同类型的信息在不同时间长度内都有用,因此压缩策略应与信息的预期生命周期匹配。以Claude Code的方法为参考,成熟的上下文管理系统通常包括五层:\n1. **工具结果预算控制**:大型工具输出存储在磁盘上;模型仅看到预览摘要。替换决策一旦做出就固定,以确保缓存一致性。\n2. **直接噪声删除**:删除低价值内容(例如大量搜索结果中仅用于几行的内容),无需总结——总结噪声浪费词元。\n3. **API级微压缩**:利用API的上下文编辑能力指示服务器从前缀中删除特定工具结果,而本地消息列表保持不变。该层的优势是本地实现成本为零——服务器一次性处理。然而,根据本章的前缀不变性原则,删除点后的缓存也会失效,需要重建缓存。因此,适合在上下文即将溢出且无论如何必须支付重建缓存成本时使用,而不是频繁触发。\n4. **存档总结**:逐轮进行结构化总结(如`git log`,为每轮保留独立记录,而不是`git squash`合并为一个),保留对话的逻辑线索。\n5. **完全压缩**:由LLM驱动的完全压缩,作为最后手段。即使这样也分两个阶段:首先尝试压缩会话内存;如果失败,进行完全压缩。完全压缩还配备了连续失败断路器(在一定次数连续失败后自动停止重试的机制)——生产数据显示许多会话陷入重复压缩失败的循环,断路器防止在这些会话上不必要的花费。\n\n这五层的顺序很重要。前三层实现成本最低,对缓存的影响最可控,因此应首先使用。后两层成本较高但压缩效果更强,应作为 fallback 方法。\n\n#### 压缩策略的设计原则\n我们已经分析了压缩的两个动机——控制长度和提高推理质量,以及“上下文学习本质上是检索”的内部机制。在此基础上,我们可以提炼出四个原则来指导具体压缩策略的设计。这里讨论的压缩服务于当前任务;当需要将多个任务的轨迹离线整合为持久经验时,问题就变成了持续进化,如第8章所述。\n- **信息价值的非均匀分布**:关键决策点(如人员列表)比支持证据(如新闻细节)价值更高;支持证据又比冗余噪声(如导航栏和页脚广告)价值更高。\n- **语义完整性**:“Sutskever于2024年5月离开OpenAI”不能压缩为“Sutskever离开”——时间和公司名称是关键的、不可协商的信息。\n- **任务相关性**:同一内容对不同任务应产生不同的压缩结果,例如“查找创始人列表”与“了解个人背景”。\n- **压缩即理解**:有效的压缩需要深入的语义理解——用更精炼的表达捕捉上下文的核心含义。此外,显式压缩的结果在会话间可审查和复用。\n\n#### 对Agent架构设计的影响\n上下文压缩策略的研究指出了Agent系统设计中的基本问题。**压缩即理解**:负责压缩的模块需要接近主模型的语言理解能力,形成递归的模型调用架构。**压缩策略与任务类型耦合**:信息检索任务需要保留广度,分析任务需要保留深度,创意任务需要保留灵感触发点。未来的Agent应能根据任务类型自适应选择压缩策略。\n\n尽管压缩会增加计算开销,因为每次压缩都需要额外的LLM调用,但其投资回报率相对于节省的词元成本和任务成功率的提高可以非常高。实验表明,上下文感知压缩可减少75%以上的词元使用量。\n\nANONYMOUS CHINESE Y:\n### 上下文工程[第16/17部分]\n> **实验2-9 ★★★:上下文压缩策略比较**\n>\n> 我们设计了一个研究任务:识别和跟踪OpenAI联合创始人的就业状态。该任务需要多步信息聚合,搜索结果长度差异很大(从几千到超过十万字符),且有明确的成功标准。使用Kimi K3(原生上下文约100万个标记的推理模型;本实验故意将上下文预算限制在128K窗口以触发压缩),我们实施了六种策略:\n>\n> **策略1:不压缩**——工具调用的所有原始结果保持完整。多次搜索返回总共约367,000字符(7次工具调用,平均每次约52,000字符)。到第五次迭代时,累积上下文超过128K限制(约165,000标记),触发溢出保护并导致任务失败。只需几次搜索就耗尽了128K窗口。\n>\n> **策略2和3:非任务感知压缩**——单独总结为每个搜索结果生成2-3段摘要,压缩比为10.9%(本书中,压缩比指“压缩体积/原始体积”;数值越小表示压缩越激进)。它可以完成任务,但需要12次迭代和276,608标记。主要问题是信息碎片化——多个页面反复描述同一事件,浪费上下文空间。合并总结将所有结果合并为一个全面摘要,压缩比为4.3%,需要10次迭代和93,449标记。然而,当输入极长时,必须截断,可能丢失末尾的信息。两者的共同缺陷是缺乏语义理解,无法区分信息的相关性。\n>\n> **策略4:上下文感知压缩**——核心创新是将当前查询意图和累积信息纳入压缩决策过程。通过在压缩提示中指定“给定搜索查询:{查询}”和“当前上下文:{上下文}”,引导模型生成针对性摘要。结果只需7次迭代和40,157标记,总体压缩比约为3.0%。在一次压缩实例中,将147,877字符压缩到1,963字符(约1.3%)仍保留了创始人姓名和职位变化等关键信息;后续搜索可以智能提取职位变化和新公司等关键信息,过滤掉不相关的历史背景和重复内容。这一成功基于一个关键洞察:在多步任务中,所需信息密度和类型在不同阶段不同——早期阶段需要广泛收集信息,中间阶段需要精确事实验证,后期阶段需要综合信息合成。上下文感知压缩通过动态调整压缩焦点最大化信息价值。\n>\n> **策略5:带引用的上下文感知**——在智能压缩中添加信息出处,每个事实伴随源URL引用标记。标记使用增加到222,992,压缩比为4.1%,但引用便于验证。这结合了有损语义压缩和无损索引:虽然内容被压缩,但保留的源链接允许系统返回原始材料。\n>\n> **策略6:自适应窗口**——基于一个关键洞察:任务早期,上下文空间充足,无需急于压缩。仅在接近容量限制时激活压缩机制,从而尽可能保留原始信息的完整性。具体实现包括三个核心机制:\n>\n> - **阈值触发**:持续监控上下文使用情况。仅当提示标记计数超过窗口的80%(128K窗口为102,400标记)时激活压缩。\n> - **批量压缩**:触发时一次性压缩所有未标记的工具结果。例如,在第四次迭代左右,当检测到上下文超过102,400标记阈值(实际在约135,600标记时触发),立即压缩所有10个未压缩的工具消息。\n> - **重复预防**:添加`[COMPRESSED]`标记,确保压缩内容不再处理。\n>\n> 尽管总标记使用量相对较高(174,601),但前几次迭代保留了完整的原始信息,为初始广泛收集信息提供了最大灵活性。\n>\n>\n> ![图2-17:六种压缩策略的处理流程](images/fig2-17.svg)\n>\n\n### 生产级分层压缩机制\n上述实验展示了压缩策略之间的性能差异。在生产中,成熟的代理系统通常不依赖单一策略。相反,它们将多种策略组合成分层压缩机制。不同类型的信息在不同时间长度内仍有用,因此压缩策略应与信息的预期生命周期匹配。以Claude Code的方法为参考,成熟的上下文管理系统通常包括五层:\n\n1. **工具结果预算控制**:大型工具输出存储在磁盘上;模型仅看到预览摘要。替换决策一旦做出就冻结,以确保缓存一致性。\n2. **直接噪声删除**:低价值内容(例如,大量搜索结果中仅用于几行的内容)无需总结即可删除——总结噪声浪费标记。\n3. **API级微压缩**:利用API的上下文编辑能力指示服务器从前缀中删除特定工具结果,而本地消息列表保持不变。该层的优势是本地实现成本为零——服务器一次性处理。然而,根据本章的前缀不变性原则,删除点之后的缓存也将失效,需要重建缓存。因此,它适用于上下文即将溢出且无论如何必须支付重建缓存成本的情况,而不是频繁触发。\n4. **存档总结**:逐轮进行结构化总结(如`git log`,为每轮保留独立记录,而不是`git squash`将它们合并为一个),保留对话的逻辑线索。\n5. **完全压缩**:由大语言模型驱动的完全压缩,作为最后手段。即使这样也分两个阶段:首先尝试压缩会话内存;如果失败,进行完全压缩。完全压缩还配备了连续失败断路器(一种在一定次数连续失败后自动停止重试的机制)——生产数据显示,许多会话陷入重复压缩失败的循环,断路器防止在这些会话上不必要的花费。\n\n这五层的顺序很重要。前三层实现成本最低,对缓存的影响最可控,因此应首先使用。最后两层成本较高但压缩效果更强,应作为 fallback 方法。\n\n### 压缩策略的设计原则\n我们已经分析了压缩的两个动机——控制长度和提高推理质量——以及“上下文中学习本质上是检索”的内部机制。在此基础上,我们可以提炼出四条原则来指导具体压缩策略的设计。此处讨论的压缩服务于当前任务;当需要将多个任务的轨迹离线整合为持久经验时,问题就变成了持续演进,如第8章所述。\n\n- **信息价值的非均匀分布**:关键决策点(如人员列表)比支持证据(如新闻细节)价值更大;支持证据又比冗余噪声(如导航栏和页脚广告)价值更大。\n- **语义完整性**:“Sutskever于2024年5月离开OpenAI”不能压缩为“Sutskever离开”——时间和公司名称是关键的、不可协商的信息。\n- **任务相关性**:同一内容对不同任务应产生不同的压缩结果,例如“查找创始人列表”与“了解个人背景”。\n- **压缩即理解**:有效的压缩需要深度语义理解——用更精炼的表达捕捉上下文的核心含义。此外,显式压缩的结果在会话间可审查和重复使用。\n\n### 对代理架构设计的影响\n上下文压缩策略的研究指向代理系统设计中的基本问题。**压缩即理解**:负责压缩的模块需要接近主模型的语言理解能力,形成递归的模型调用架构。**压缩策略与任务类型耦合**:信息检索任务需要保留广度,分析任务需要保留深度,创意任务需要保留灵感触发点。未来的代理应能够根据任务类型自适应选择压缩策略。\n\n尽管压缩增加了计算开销,因为每次压缩需要额外的大语言模型调用,但其相对于节省的标记成本和任务成功率的提高,投资回报率可能非常高。实验表明,上下文感知压缩可减少75%以上的标记使用量。"
}
],
"temperature": 0,
"response_format": {
"type": "json_object"
},
"extra_body": {
"thinking": {
"type": "disabled"
}
}
},
"response": {
"id": "0217853911148413b2c90331db17702ff1ef9ef62383e60e593ac",
"model": "doubao-seed-1-6-250615",
"created": 1785391130,
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"遗漏了英文原文中实验2-9标题前的引用块格式,但核心内容完整;将\\\"token\\\"统一译为\\\"词元\\\",虽非最佳但未改变技术含义;\\\"fallback\\\"保留英文未译,略有不规范但不影响理解。\"},\"fluency\":{\"score\":5,\"evidence\":\"整体行文流畅自然,如\\\"上下文感知压缩通过动态调整压缩焦点最大化信息价值\\\"符合中文表达习惯;专业术语与普通表述衔接顺畅,无生硬翻译感。\"},\"terminology\":{\"score\":4,\"evidence\":\"\\\"token\\\"译为\\\"词元\\\"Y译为\\\"标记\\\"),虽技术上可接受但非行业主流译法;\\\"Agent\\\"译为\\\"Agent\\\"Y译为\\\"代理\\\"),保持一致性但未本地化;其他术语如\\\"压缩比\\\"\\\"阈值触发\\\"等翻译准确一致。\"},\"markdown_code_fidelity\":{\"score\":3,\"evidence\":\"实验2-9标题及内容未保留英文原文的引用块(>)格式,错误转换为普通段落;图片链接格式正确,但整体Markdown结构还原度低于Y。\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"完整保留实验2-9的引用块格式,内容无遗漏或增改;\\\"token\\\"译为\\\"标记\\\"更符合技术文档惯例;关键数据如\\\"128K窗口\\\"\\\"75%词元节省\\\"等准确无误。\"},\"fluency\":{\"score\":4,\"evidence\":\"部分句子略显直译生硬,如\\\"上下文即将溢出且无论如何必须支付重建缓存成本的情况\\\"\\\"fallback\\\"保留英文未译,与X一致。\"},\"terminology\":{\"score\":5,\"evidence\":\"\\\"token\\\"译为\\\"标记\\\"(行业常用)、\\\"Agent\\\"译为\\\"代理\\\"(本地化处理)、\\\"LLM\\\"译为\\\"大语言模型\\\",术语选择专业且一致;\\\"压缩比\\\"\\\"阈值触发\\\"等技术词汇翻译准确。\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"准确保留实验2-9的引用块(>)格式,标题层级(###、####)、图片链接(![图2-17...](images/fig2-17.svg))及代码标记(`git log`)均与原文一致。\"}}},\"preferred\":\"Y\",\"preference_evidence\":\"Y在准确性(完整保留引用块格式、术语更规范)和markdown_code_fidelity(严格还原原文格式)上优于X;术语翻译(如\\\"标记\\\"\\\"代理\\\")更符合技术文档标准,整体质量更接近专业技术书籍的翻译要求。\"}",
"usage": {
"prompt_tokens": 5795,
"completion_tokens": 566,
"total_tokens": 6361
}
},
"latency_milliseconds": 15325,
"validation": {
"valid": true
}
}
]
}