Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
189 lines
122 KiB
JSON
189 lines
122 KiB
JSON
{
|
||
"schema_version": 1,
|
||
"credential_free": true,
|
||
"attempts": [
|
||
{
|
||
"attempt": 1,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nIs every component truly indispensable? The most direct way to find out is an **ablation study**—the diagnostic method of ruling out causes one at a time: remove component A and see whether the system still works, then component B, and so on, until each component’s contribution is clear. Experiment 1-1 applies exactly this method to the five components above. The results are direct: without tool definitions, the Agent is completely incapable of action; without tool results, it does not receive feedback from the previous step, so it calls the same tool repeatedly, becoming stuck in an infinite loop; without the reasoning in assistant messages, consecutive decisions start contradicting each other; without message history, the Agent loses task continuity and restarts the whole task from the beginning, repeating steps already done. Each component’s role rests on experimental evidence, not just theoretical inference.\n\n### Experiment 1-1 ★★: The Critical Role of Context\n\nWe probed how each context component shapes Agent behavior with a systematic **ablation study**. Of the five components above, four were tested—the system prompt, as the Agent’s basic identity definition, was exempt: without it the Agent has no role awareness at all, and the test would be meaningless. As Figure 1-2 shows, the experiment ran five controlled groups: a complete baseline retaining every component, plus four groups each missing one, to observe each component’s effect on Agent performance.\n\n\n\nThe experimental results revealed the irreplaceable role of each context component. **Tool Definitions** (part of the static prefix) are the foundation of the Agent’s action capability; without them, the Agent cannot recognize or call any tools. **Tool Results** are key to closed-loop control; their absence deprives the Agent of execution feedback and causes it to fall into an infinite loop. The **reasoning process** (the reasoning part of assistant messages) preserves the reasons for the Agent’s previous decisions, making the overall reasoning more coherent and preventing contradictory decisions. **Message history** (user messages, assistant messages, and tool results from previous rounds) prevents redundant operations, maintains task execution coherence, and avoids repeating the same mistakes.\n\nThe experiment's core insight: **context determines what information the Agent has at decision time, and the Agent can only decide based on that information**. Just as a person missing crucial documents cannot make sound judgments, an Agent missing any context component suffers a severe loss of decision-making ability—without tool definitions it does not know what tools exist; without previous execution results it does not know what has already been done.\n\n### The ReAct Loop\n\nWith the three components in hand, a natural question follows: how do they work together? The ReAct loop is the core mechanism that connects LLM, context, and tools into one system. We can examine it step by step.\n\nThe core pattern by which an Agent executes a task is called **ReAct** (Reasoning + Acting). The name mentions only reasoning and acting, but the actual loop has three stages: the model first **reasons** about what to do next, then calls a tool to **act**, then **observes** the tool’s result and reasons about the subsequent step. This “reason → act → observe → reason → act → observe” loop repeats until the task is done.\n\nConsider a concrete example—aggregating revenue across multiple currencies—to understand an Agent’s **trajectory**: the message history that accumulates as the Agent works, comprising user messages, assistant messages (with their reasoning and tool calls), and tool results. On every LLM call, the complete context the model receives is the **static prefix** (system prompt + tool definitions) plus the **trajectory** (dynamic message history) (Figure 1-3). This shows a key fact: **Agent context = static prefix + trajectory**. Concretely, the static prefix is the first two of the five components above (system prompt + tool definitions); the trajectory is the last three (user messages + assistant messages + tool results, growing with each interaction). From this complete context the LLM generates its next response, which is then appended to the trajectory for the subsequent call.\n\n\n\nHere is the structure of a trajectory, in pseudocode:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # First iteration - LLM receives the above trajectory and generates a response\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # No direct reply to the user\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # Agent framework executes tools, adds results to trajectory\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # Second iteration - LLM receives the complete trajectory, including tool results\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # Third iteration - LLM receives the complete trajectory and generates the final answer\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\nNote that the system prompt and tool definitions are not shown in the trajectory—they serve as the static prefix and are automatically prepended to the trajectory before each LLM call.\n\nIn our experiment, this loop was clearly visible. In the first round, the Agent analyzed the task and called three currency conversion tools in parallel; in the second, it fed the conversion results to a code interpreter for the more computationally intensive calculation; in the third, having confirmed all calculations were complete, it produced the final answer. A complex multi-step task was completed in 3 iterations and 4 tool calls.\n\nThe elegance of this design lies in the **cumulative nature of the context**. Every LLM call receives the complete trajectory, so the model knows which stage of the task it is in, what was tried before, and what the outcome was. Just as people keep reviewing and summarizing while solving a problem, the Agent maintains a global view of the task through its trajectory. And because the trajectory is structured—user messages, assistant messages (reasoning + tool calls), and tool results all separated cleanly—the system is highly interpretable and debuggable.\n\nThe trajectory is more than an execution record; it is evidence of the Agent’s capability. Analyzing trajectories at scale reveals behavior patterns, better decision paths, and better tool designs. Trajectory data can even be distilled into a knowledge base, or used to train stronger Agent models via reinforcement learning—closing the loop of learning from experience.\n\nNow that we understand the Agent's operating loop, we examine two experiments to see how different models drive it.\n\n#### Experiment 1-2 ★: Kimi K3 Native Agent Capability\n\nThis experiment demonstrates the native Agent capability of **Kimi K3**, an example of the “Model as Agent” paradigm. Released by Moonshot AI in 2026, Kimi K3 is a Mixture of Experts (MoE) model with approximately 2.8 trillion parameters. MoE can be viewed as a team of experts: for each kind of problem, the system activates only the few experts best suited to it rather than the entire model, preserving capability without paying the full efficiency cost. Kimi K3 has a 1 million token context window, native visual understanding, and an always-on “thinking mode.” Through reinforcement learning, it has internalized the tool-calling **decision policy** as a native capability: when to call a tool, which tool to call, and what arguments to pass are all decided by the model, allowing it to carry out tasks such as web searches autonomously. To be precise, what is internalized is the *when and how to call* decision; the tools themselves, such as `web_search` and `code_runner`, still execute server-side as API-level built-in tools. Kimi runs these official tools through a server-side script engine called Formula.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第4/9部分]\n每个组件真的都不可或缺吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看系统是否还能工作,然后是组件B,依此类推,直到每个组件的贡献清晰可见。实验1-1正是对上述五个组件应用了这种方法。结果直接明了:没有工具定义,代理完全无法行动;没有工具结果,它无法从之前的步骤获得反馈,所以会反复调用同一个工具,陷入无限循环;没有助手消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理失去任务连续性,从头重新开始整个任务,重复已做的步骤。每个组件的作用都有实验证据支持,而不仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的消融研究探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义,被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:保留所有组件的完整基线组,以及四组各缺失一个组件的组,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺失它们会让代理失去执行反馈,导致陷入无限循环。**推理过程**(助手消息中的推理部分)保留了代理先前决策的理由,使整体推理更连贯,防止矛盾决策。**消息历史**(之前轮次的用户消息、助手消息和工具结果)防止重复操作,保持任务执行连贯性,避免重复同样的错误。\n\n实验的核心洞察是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。就像一个人缺少关键文件无法做出明智判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义它不知道存在哪些工具;没有之前的执行结果它不知道已经做了什么。\n\n### ReAct循环\n有了这三个组件,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到推理和行动,但实际循环有三个阶段:模型首先**推理**下一步做什么,然后调用工具来**行动**,然后**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环重复直到任务完成。\n\n以聚合多种货币的收入为例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助手消息(带有其推理和工具调用)和工具结果。在每次大语言模型调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助手消息+工具结果,随每次交互增长)。大语言模型从这个完整上下文中生成其下一个响应,然后该响应附加到轨迹中用于后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # 没有直接回复用户\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\n请注意,系统提示和工具定义未显示在轨迹中——它们作为静态前缀,在每次大语言模型调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量更大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的优雅之处在于**上下文的累积性**。每次大语言模型调用都接收完整轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。就像人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且因为轨迹是结构化的——用户消息、助手消息(推理+工具调用)和工具结果都清晰分离——系统具有高度可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证据。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的循环。\n\n现在我们理解了代理的操作循环,接下来审视两个实验,看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3的本地代理能力\n这个实验展示了**Kimi K3**的本地代理能力,它是“模型即代理”范式的一个示例。由月之暗面科技于2026年发布的Kimi K3是一个混合专家(MoE)模型,约有2.8万亿参数。MoE可以看作是一组专家:对于每种问题,系统只激活最适合它的少数专家,而不是整个模型,在保持能力的同时不支付全部效率成本。Kimi K3有100万个标记的上下文窗口、本地视觉理解和始终在线的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化成本地能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过名为Formula的服务器端脚本引擎运行这些官方工具。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第4/9部分]\n每个组件真的都是不可或缺的吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看看系统是否仍然有效,然后是组件B,依此类推,直到每个组件的贡献清晰明了。实验1-1正是将这种方法应用于上述五个组件。结果一目了然:没有工具定义,代理完全无法行动;没有工具结果,它不会收到上一步的反馈,因此会反复调用同一个工具,陷入无限循环;没有助理消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理会失去任务连续性,从头开始重新执行整个任务,重复已经完成的步骤。每个组件的作用都有实验证据支持,而非仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的**消融研究**探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:一组保留所有组件的完整基线,另外四组每组缺少一个组件,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺少它们会剥夺代理的执行反馈,导致其陷入无限循环。**推理过程**(助理消息中的推理部分)保留了代理先前决策的原因,使整体推理更连贯,防止矛盾决策。**消息历史**(用户消息、助理消息和之前轮次的工具结果)防止冗余操作,保持任务执行的连贯性,避免重复同样的错误。\n\n实验的核心见解是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。正如一个人缺少关键文件无法做出合理判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义,它不知道存在哪些工具;没有之前的执行结果,它不知道已经做了什么。\n\n### ReAct循环\n掌握了三个组件后,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型(LLM)、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到了推理和行动,但实际循环有三个阶段:模型首先**推理**下一步该做什么,然后调用工具来**行动**,接着**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环会重复,直到任务完成。\n\n以跨多种货币汇总收入的具体示例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助理消息(含推理和工具调用)和工具结果。每次LLM调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助理消息+工具结果,随每次交互增长)。LLM从这个完整上下文中生成下一个响应,然后附加到轨迹中供后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\n轨迹 = [\n {角色: \"用户\", 内容: \"根据公司季度收入:第一季度250万美元(美元)、第二季度210万欧元、第三季度180万英镑、第四季度3.8亿日元,计算公司年度总收入和平均季度收入\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {角色: \"助理\",\n 推理: \"需要将所有货币转换为美元...\",\n 内容: \"\", # 没有直接回复用户\n 工具调用: [\n {名称: \"convert_currency\", 参数: {金额: 2100000, 来自: \"EUR\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 1800000, 来自: \"GBP\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 380000000, 来自: \"JPY\", 到: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {角色: \"工具\", 内容: \"欧元→美元:2282608.7\"},\n {角色: \"工具\", 内容: \"英镑→美元:2278481.01\"},\n {角色: \"工具\", 内容: \"日元→美元:2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {角色: \"助理\",\n 推理: \"已获得转换结果,现在需要汇总并计算...\",\n 内容: \"\",\n 工具调用: [\n {名称: \"code_interpreter\", 参数: {代码: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {角色: \"工具\", 内容: \"总计:9,602,895.73美元,平均:2,400,723.93美元...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {角色: \"助理\",\n 推理: \"所有计算完成,总结结果...\",\n 内容: \"最终答案:总收入9,602,895.73美元...\"},\n]\n```\n\n请注意,系统提示和工具定义未在轨迹中显示——它们作为静态前缀,在每次LLM调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量较大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的精妙之处在于**上下文的累积性**。每次LLM调用都接收完整的轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。正如人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且由于轨迹结构清晰——用户消息、助理消息(推理+工具调用)和工具结果都明确分开,系统具有高度的可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证明。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的闭环。\n\n现在我们了解了代理的操作循环,接下来通过两个实验看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3原生代理能力\n这个实验展示了**Kimi K3**的原生代理能力,这是“模型即代理”范式的一个示例。Kimi K3由月之暗面公司于2026年发布,是一个约有2.8万亿参数的专家混合(MoE)模型。MoE可视为一个专家团队:对于每种问题,系统仅激活最适合它的少数专家,而不是整个模型,在保持能力的同时避免了全部效率成本。Kimi K3具有100万个词元的上下文窗口、原生视觉理解能力和始终开启的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化为原生能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过一个名为Formula的服务器端脚本引擎运行这些官方工具。"
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "021785390307069e4feb8d226730241afb8b5c4c4c35b326f714a",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390322,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}",
|
||
"usage": {
|
||
"prompt_tokens": 6206,
|
||
"completion_tokens": 528,
|
||
"total_tokens": 6734
|
||
}
|
||
},
|
||
"latency_milliseconds": 15737,
|
||
"validation": {
|
||
"valid": false,
|
||
"error_type": "ValueError",
|
||
"error": "judge variants must contain exactly X and Y"
|
||
}
|
||
},
|
||
{
|
||
"attempt": 2,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nIs every component truly indispensable? The most direct way to find out is an **ablation study**—the diagnostic method of ruling out causes one at a time: remove component A and see whether the system still works, then component B, and so on, until each component’s contribution is clear. Experiment 1-1 applies exactly this method to the five components above. The results are direct: without tool definitions, the Agent is completely incapable of action; without tool results, it does not receive feedback from the previous step, so it calls the same tool repeatedly, becoming stuck in an infinite loop; without the reasoning in assistant messages, consecutive decisions start contradicting each other; without message history, the Agent loses task continuity and restarts the whole task from the beginning, repeating steps already done. Each component’s role rests on experimental evidence, not just theoretical inference.\n\n### Experiment 1-1 ★★: The Critical Role of Context\n\nWe probed how each context component shapes Agent behavior with a systematic **ablation study**. Of the five components above, four were tested—the system prompt, as the Agent’s basic identity definition, was exempt: without it the Agent has no role awareness at all, and the test would be meaningless. As Figure 1-2 shows, the experiment ran five controlled groups: a complete baseline retaining every component, plus four groups each missing one, to observe each component’s effect on Agent performance.\n\n\n\nThe experimental results revealed the irreplaceable role of each context component. **Tool Definitions** (part of the static prefix) are the foundation of the Agent’s action capability; without them, the Agent cannot recognize or call any tools. **Tool Results** are key to closed-loop control; their absence deprives the Agent of execution feedback and causes it to fall into an infinite loop. The **reasoning process** (the reasoning part of assistant messages) preserves the reasons for the Agent’s previous decisions, making the overall reasoning more coherent and preventing contradictory decisions. **Message history** (user messages, assistant messages, and tool results from previous rounds) prevents redundant operations, maintains task execution coherence, and avoids repeating the same mistakes.\n\nThe experiment's core insight: **context determines what information the Agent has at decision time, and the Agent can only decide based on that information**. Just as a person missing crucial documents cannot make sound judgments, an Agent missing any context component suffers a severe loss of decision-making ability—without tool definitions it does not know what tools exist; without previous execution results it does not know what has already been done.\n\n### The ReAct Loop\n\nWith the three components in hand, a natural question follows: how do they work together? The ReAct loop is the core mechanism that connects LLM, context, and tools into one system. We can examine it step by step.\n\nThe core pattern by which an Agent executes a task is called **ReAct** (Reasoning + Acting). The name mentions only reasoning and acting, but the actual loop has three stages: the model first **reasons** about what to do next, then calls a tool to **act**, then **observes** the tool’s result and reasons about the subsequent step. This “reason → act → observe → reason → act → observe” loop repeats until the task is done.\n\nConsider a concrete example—aggregating revenue across multiple currencies—to understand an Agent’s **trajectory**: the message history that accumulates as the Agent works, comprising user messages, assistant messages (with their reasoning and tool calls), and tool results. On every LLM call, the complete context the model receives is the **static prefix** (system prompt + tool definitions) plus the **trajectory** (dynamic message history) (Figure 1-3). This shows a key fact: **Agent context = static prefix + trajectory**. Concretely, the static prefix is the first two of the five components above (system prompt + tool definitions); the trajectory is the last three (user messages + assistant messages + tool results, growing with each interaction). From this complete context the LLM generates its next response, which is then appended to the trajectory for the subsequent call.\n\n\n\nHere is the structure of a trajectory, in pseudocode:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # First iteration - LLM receives the above trajectory and generates a response\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # No direct reply to the user\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # Agent framework executes tools, adds results to trajectory\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # Second iteration - LLM receives the complete trajectory, including tool results\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # Third iteration - LLM receives the complete trajectory and generates the final answer\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\nNote that the system prompt and tool definitions are not shown in the trajectory—they serve as the static prefix and are automatically prepended to the trajectory before each LLM call.\n\nIn our experiment, this loop was clearly visible. In the first round, the Agent analyzed the task and called three currency conversion tools in parallel; in the second, it fed the conversion results to a code interpreter for the more computationally intensive calculation; in the third, having confirmed all calculations were complete, it produced the final answer. A complex multi-step task was completed in 3 iterations and 4 tool calls.\n\nThe elegance of this design lies in the **cumulative nature of the context**. Every LLM call receives the complete trajectory, so the model knows which stage of the task it is in, what was tried before, and what the outcome was. Just as people keep reviewing and summarizing while solving a problem, the Agent maintains a global view of the task through its trajectory. And because the trajectory is structured—user messages, assistant messages (reasoning + tool calls), and tool results all separated cleanly—the system is highly interpretable and debuggable.\n\nThe trajectory is more than an execution record; it is evidence of the Agent’s capability. Analyzing trajectories at scale reveals behavior patterns, better decision paths, and better tool designs. Trajectory data can even be distilled into a knowledge base, or used to train stronger Agent models via reinforcement learning—closing the loop of learning from experience.\n\nNow that we understand the Agent's operating loop, we examine two experiments to see how different models drive it.\n\n#### Experiment 1-2 ★: Kimi K3 Native Agent Capability\n\nThis experiment demonstrates the native Agent capability of **Kimi K3**, an example of the “Model as Agent” paradigm. Released by Moonshot AI in 2026, Kimi K3 is a Mixture of Experts (MoE) model with approximately 2.8 trillion parameters. MoE can be viewed as a team of experts: for each kind of problem, the system activates only the few experts best suited to it rather than the entire model, preserving capability without paying the full efficiency cost. Kimi K3 has a 1 million token context window, native visual understanding, and an always-on “thinking mode.” Through reinforcement learning, it has internalized the tool-calling **decision policy** as a native capability: when to call a tool, which tool to call, and what arguments to pass are all decided by the model, allowing it to carry out tasks such as web searches autonomously. To be precise, what is internalized is the *when and how to call* decision; the tools themselves, such as `web_search` and `code_runner`, still execute server-side as API-level built-in tools. Kimi runs these official tools through a server-side script engine called Formula.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第4/9部分]\n每个组件真的都不可或缺吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看系统是否还能工作,然后是组件B,依此类推,直到每个组件的贡献清晰可见。实验1-1正是对上述五个组件应用了这种方法。结果直接明了:没有工具定义,代理完全无法行动;没有工具结果,它无法从之前的步骤获得反馈,所以会反复调用同一个工具,陷入无限循环;没有助手消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理失去任务连续性,从头重新开始整个任务,重复已做的步骤。每个组件的作用都有实验证据支持,而不仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的消融研究探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义,被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:保留所有组件的完整基线组,以及四组各缺失一个组件的组,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺失它们会让代理失去执行反馈,导致陷入无限循环。**推理过程**(助手消息中的推理部分)保留了代理先前决策的理由,使整体推理更连贯,防止矛盾决策。**消息历史**(之前轮次的用户消息、助手消息和工具结果)防止重复操作,保持任务执行连贯性,避免重复同样的错误。\n\n实验的核心洞察是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。就像一个人缺少关键文件无法做出明智判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义它不知道存在哪些工具;没有之前的执行结果它不知道已经做了什么。\n\n### ReAct循环\n有了这三个组件,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到推理和行动,但实际循环有三个阶段:模型首先**推理**下一步做什么,然后调用工具来**行动**,然后**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环重复直到任务完成。\n\n以聚合多种货币的收入为例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助手消息(带有其推理和工具调用)和工具结果。在每次大语言模型调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助手消息+工具结果,随每次交互增长)。大语言模型从这个完整上下文中生成其下一个响应,然后该响应附加到轨迹中用于后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # 没有直接回复用户\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\n请注意,系统提示和工具定义未显示在轨迹中——它们作为静态前缀,在每次大语言模型调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量更大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的优雅之处在于**上下文的累积性**。每次大语言模型调用都接收完整轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。就像人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且因为轨迹是结构化的——用户消息、助手消息(推理+工具调用)和工具结果都清晰分离——系统具有高度可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证据。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的循环。\n\n现在我们理解了代理的操作循环,接下来审视两个实验,看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3的本地代理能力\n这个实验展示了**Kimi K3**的本地代理能力,它是“模型即代理”范式的一个示例。由月之暗面科技于2026年发布的Kimi K3是一个混合专家(MoE)模型,约有2.8万亿参数。MoE可以看作是一组专家:对于每种问题,系统只激活最适合它的少数专家,而不是整个模型,在保持能力的同时不支付全部效率成本。Kimi K3有100万个标记的上下文窗口、本地视觉理解和始终在线的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化成本地能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过名为Formula的服务器端脚本引擎运行这些官方工具。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第4/9部分]\n每个组件真的都是不可或缺的吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看看系统是否仍然有效,然后是组件B,依此类推,直到每个组件的贡献清晰明了。实验1-1正是将这种方法应用于上述五个组件。结果一目了然:没有工具定义,代理完全无法行动;没有工具结果,它不会收到上一步的反馈,因此会反复调用同一个工具,陷入无限循环;没有助理消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理会失去任务连续性,从头开始重新执行整个任务,重复已经完成的步骤。每个组件的作用都有实验证据支持,而非仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的**消融研究**探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:一组保留所有组件的完整基线,另外四组每组缺少一个组件,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺少它们会剥夺代理的执行反馈,导致其陷入无限循环。**推理过程**(助理消息中的推理部分)保留了代理先前决策的原因,使整体推理更连贯,防止矛盾决策。**消息历史**(用户消息、助理消息和之前轮次的工具结果)防止冗余操作,保持任务执行的连贯性,避免重复同样的错误。\n\n实验的核心见解是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。正如一个人缺少关键文件无法做出合理判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义,它不知道存在哪些工具;没有之前的执行结果,它不知道已经做了什么。\n\n### ReAct循环\n掌握了三个组件后,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型(LLM)、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到了推理和行动,但实际循环有三个阶段:模型首先**推理**下一步该做什么,然后调用工具来**行动**,接着**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环会重复,直到任务完成。\n\n以跨多种货币汇总收入的具体示例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助理消息(含推理和工具调用)和工具结果。每次LLM调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助理消息+工具结果,随每次交互增长)。LLM从这个完整上下文中生成下一个响应,然后附加到轨迹中供后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\n轨迹 = [\n {角色: \"用户\", 内容: \"根据公司季度收入:第一季度250万美元(美元)、第二季度210万欧元、第三季度180万英镑、第四季度3.8亿日元,计算公司年度总收入和平均季度收入\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {角色: \"助理\",\n 推理: \"需要将所有货币转换为美元...\",\n 内容: \"\", # 没有直接回复用户\n 工具调用: [\n {名称: \"convert_currency\", 参数: {金额: 2100000, 来自: \"EUR\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 1800000, 来自: \"GBP\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 380000000, 来自: \"JPY\", 到: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {角色: \"工具\", 内容: \"欧元→美元:2282608.7\"},\n {角色: \"工具\", 内容: \"英镑→美元:2278481.01\"},\n {角色: \"工具\", 内容: \"日元→美元:2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {角色: \"助理\",\n 推理: \"已获得转换结果,现在需要汇总并计算...\",\n 内容: \"\",\n 工具调用: [\n {名称: \"code_interpreter\", 参数: {代码: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {角色: \"工具\", 内容: \"总计:9,602,895.73美元,平均:2,400,723.93美元...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {角色: \"助理\",\n 推理: \"所有计算完成,总结结果...\",\n 内容: \"最终答案:总收入9,602,895.73美元...\"},\n]\n```\n\n请注意,系统提示和工具定义未在轨迹中显示——它们作为静态前缀,在每次LLM调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量较大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的精妙之处在于**上下文的累积性**。每次LLM调用都接收完整的轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。正如人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且由于轨迹结构清晰——用户消息、助理消息(推理+工具调用)和工具结果都明确分开,系统具有高度的可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证明。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的闭环。\n\n现在我们了解了代理的操作循环,接下来通过两个实验看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3原生代理能力\n这个实验展示了**Kimi K3**的原生代理能力,这是“模型即代理”范式的一个示例。Kimi K3由月之暗面公司于2026年发布,是一个约有2.8万亿参数的专家混合(MoE)模型。MoE可视为一个专家团队:对于每种问题,系统仅激活最适合它的少数专家,而不是整个模型,在保持能力的同时避免了全部效率成本。Kimi K3具有100万个词元的上下文窗口、原生视觉理解能力和始终开启的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化为原生能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过一个名为Formula的服务器端脚本引擎运行这些官方工具。"
|
||
},
|
||
{
|
||
"role": "assistant",
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}"
|
||
},
|
||
{
|
||
"role": "user",
|
||
"content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys."
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "021785390322807e4feb8d226730241afb8b5c4c4c35b32e5d542",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390335,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}",
|
||
"usage": {
|
||
"prompt_tokens": 6804,
|
||
"completion_tokens": 528,
|
||
"total_tokens": 7332
|
||
}
|
||
},
|
||
"latency_milliseconds": 12585,
|
||
"validation": {
|
||
"valid": false,
|
||
"error_type": "ValueError",
|
||
"error": "judge variants must contain exactly X and Y"
|
||
}
|
||
},
|
||
{
|
||
"attempt": 3,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nIs every component truly indispensable? The most direct way to find out is an **ablation study**—the diagnostic method of ruling out causes one at a time: remove component A and see whether the system still works, then component B, and so on, until each component’s contribution is clear. Experiment 1-1 applies exactly this method to the five components above. The results are direct: without tool definitions, the Agent is completely incapable of action; without tool results, it does not receive feedback from the previous step, so it calls the same tool repeatedly, becoming stuck in an infinite loop; without the reasoning in assistant messages, consecutive decisions start contradicting each other; without message history, the Agent loses task continuity and restarts the whole task from the beginning, repeating steps already done. Each component’s role rests on experimental evidence, not just theoretical inference.\n\n### Experiment 1-1 ★★: The Critical Role of Context\n\nWe probed how each context component shapes Agent behavior with a systematic **ablation study**. Of the five components above, four were tested—the system prompt, as the Agent’s basic identity definition, was exempt: without it the Agent has no role awareness at all, and the test would be meaningless. As Figure 1-2 shows, the experiment ran five controlled groups: a complete baseline retaining every component, plus four groups each missing one, to observe each component’s effect on Agent performance.\n\n\n\nThe experimental results revealed the irreplaceable role of each context component. **Tool Definitions** (part of the static prefix) are the foundation of the Agent’s action capability; without them, the Agent cannot recognize or call any tools. **Tool Results** are key to closed-loop control; their absence deprives the Agent of execution feedback and causes it to fall into an infinite loop. The **reasoning process** (the reasoning part of assistant messages) preserves the reasons for the Agent’s previous decisions, making the overall reasoning more coherent and preventing contradictory decisions. **Message history** (user messages, assistant messages, and tool results from previous rounds) prevents redundant operations, maintains task execution coherence, and avoids repeating the same mistakes.\n\nThe experiment's core insight: **context determines what information the Agent has at decision time, and the Agent can only decide based on that information**. Just as a person missing crucial documents cannot make sound judgments, an Agent missing any context component suffers a severe loss of decision-making ability—without tool definitions it does not know what tools exist; without previous execution results it does not know what has already been done.\n\n### The ReAct Loop\n\nWith the three components in hand, a natural question follows: how do they work together? The ReAct loop is the core mechanism that connects LLM, context, and tools into one system. We can examine it step by step.\n\nThe core pattern by which an Agent executes a task is called **ReAct** (Reasoning + Acting). The name mentions only reasoning and acting, but the actual loop has three stages: the model first **reasons** about what to do next, then calls a tool to **act**, then **observes** the tool’s result and reasons about the subsequent step. This “reason → act → observe → reason → act → observe” loop repeats until the task is done.\n\nConsider a concrete example—aggregating revenue across multiple currencies—to understand an Agent’s **trajectory**: the message history that accumulates as the Agent works, comprising user messages, assistant messages (with their reasoning and tool calls), and tool results. On every LLM call, the complete context the model receives is the **static prefix** (system prompt + tool definitions) plus the **trajectory** (dynamic message history) (Figure 1-3). This shows a key fact: **Agent context = static prefix + trajectory**. Concretely, the static prefix is the first two of the five components above (system prompt + tool definitions); the trajectory is the last three (user messages + assistant messages + tool results, growing with each interaction). From this complete context the LLM generates its next response, which is then appended to the trajectory for the subsequent call.\n\n\n\nHere is the structure of a trajectory, in pseudocode:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # First iteration - LLM receives the above trajectory and generates a response\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # No direct reply to the user\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # Agent framework executes tools, adds results to trajectory\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # Second iteration - LLM receives the complete trajectory, including tool results\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # Third iteration - LLM receives the complete trajectory and generates the final answer\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\nNote that the system prompt and tool definitions are not shown in the trajectory—they serve as the static prefix and are automatically prepended to the trajectory before each LLM call.\n\nIn our experiment, this loop was clearly visible. In the first round, the Agent analyzed the task and called three currency conversion tools in parallel; in the second, it fed the conversion results to a code interpreter for the more computationally intensive calculation; in the third, having confirmed all calculations were complete, it produced the final answer. A complex multi-step task was completed in 3 iterations and 4 tool calls.\n\nThe elegance of this design lies in the **cumulative nature of the context**. Every LLM call receives the complete trajectory, so the model knows which stage of the task it is in, what was tried before, and what the outcome was. Just as people keep reviewing and summarizing while solving a problem, the Agent maintains a global view of the task through its trajectory. And because the trajectory is structured—user messages, assistant messages (reasoning + tool calls), and tool results all separated cleanly—the system is highly interpretable and debuggable.\n\nThe trajectory is more than an execution record; it is evidence of the Agent’s capability. Analyzing trajectories at scale reveals behavior patterns, better decision paths, and better tool designs. Trajectory data can even be distilled into a knowledge base, or used to train stronger Agent models via reinforcement learning—closing the loop of learning from experience.\n\nNow that we understand the Agent's operating loop, we examine two experiments to see how different models drive it.\n\n#### Experiment 1-2 ★: Kimi K3 Native Agent Capability\n\nThis experiment demonstrates the native Agent capability of **Kimi K3**, an example of the “Model as Agent” paradigm. Released by Moonshot AI in 2026, Kimi K3 is a Mixture of Experts (MoE) model with approximately 2.8 trillion parameters. MoE can be viewed as a team of experts: for each kind of problem, the system activates only the few experts best suited to it rather than the entire model, preserving capability without paying the full efficiency cost. Kimi K3 has a 1 million token context window, native visual understanding, and an always-on “thinking mode.” Through reinforcement learning, it has internalized the tool-calling **decision policy** as a native capability: when to call a tool, which tool to call, and what arguments to pass are all decided by the model, allowing it to carry out tasks such as web searches autonomously. To be precise, what is internalized is the *when and how to call* decision; the tools themselves, such as `web_search` and `code_runner`, still execute server-side as API-level built-in tools. Kimi runs these official tools through a server-side script engine called Formula.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第4/9部分]\n每个组件真的都不可或缺吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看系统是否还能工作,然后是组件B,依此类推,直到每个组件的贡献清晰可见。实验1-1正是对上述五个组件应用了这种方法。结果直接明了:没有工具定义,代理完全无法行动;没有工具结果,它无法从之前的步骤获得反馈,所以会反复调用同一个工具,陷入无限循环;没有助手消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理失去任务连续性,从头重新开始整个任务,重复已做的步骤。每个组件的作用都有实验证据支持,而不仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的消融研究探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义,被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:保留所有组件的完整基线组,以及四组各缺失一个组件的组,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺失它们会让代理失去执行反馈,导致陷入无限循环。**推理过程**(助手消息中的推理部分)保留了代理先前决策的理由,使整体推理更连贯,防止矛盾决策。**消息历史**(之前轮次的用户消息、助手消息和工具结果)防止重复操作,保持任务执行连贯性,避免重复同样的错误。\n\n实验的核心洞察是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。就像一个人缺少关键文件无法做出明智判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义它不知道存在哪些工具;没有之前的执行结果它不知道已经做了什么。\n\n### ReAct循环\n有了这三个组件,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到推理和行动,但实际循环有三个阶段:模型首先**推理**下一步做什么,然后调用工具来**行动**,然后**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环重复直到任务完成。\n\n以聚合多种货币的收入为例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助手消息(带有其推理和工具调用)和工具结果。在每次大语言模型调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助手消息+工具结果,随每次交互增长)。大语言模型从这个完整上下文中生成其下一个响应,然后该响应附加到轨迹中用于后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # 没有直接回复用户\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\n请注意,系统提示和工具定义未显示在轨迹中——它们作为静态前缀,在每次大语言模型调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量更大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的优雅之处在于**上下文的累积性**。每次大语言模型调用都接收完整轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。就像人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且因为轨迹是结构化的——用户消息、助手消息(推理+工具调用)和工具结果都清晰分离——系统具有高度可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证据。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的循环。\n\n现在我们理解了代理的操作循环,接下来审视两个实验,看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3的本地代理能力\n这个实验展示了**Kimi K3**的本地代理能力,它是“模型即代理”范式的一个示例。由月之暗面科技于2026年发布的Kimi K3是一个混合专家(MoE)模型,约有2.8万亿参数。MoE可以看作是一组专家:对于每种问题,系统只激活最适合它的少数专家,而不是整个模型,在保持能力的同时不支付全部效率成本。Kimi K3有100万个标记的上下文窗口、本地视觉理解和始终在线的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化成本地能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过名为Formula的服务器端脚本引擎运行这些官方工具。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第4/9部分]\n每个组件真的都是不可或缺的吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看看系统是否仍然有效,然后是组件B,依此类推,直到每个组件的贡献清晰明了。实验1-1正是将这种方法应用于上述五个组件。结果一目了然:没有工具定义,代理完全无法行动;没有工具结果,它不会收到上一步的反馈,因此会反复调用同一个工具,陷入无限循环;没有助理消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理会失去任务连续性,从头开始重新执行整个任务,重复已经完成的步骤。每个组件的作用都有实验证据支持,而非仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的**消融研究**探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:一组保留所有组件的完整基线,另外四组每组缺少一个组件,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺少它们会剥夺代理的执行反馈,导致其陷入无限循环。**推理过程**(助理消息中的推理部分)保留了代理先前决策的原因,使整体推理更连贯,防止矛盾决策。**消息历史**(用户消息、助理消息和之前轮次的工具结果)防止冗余操作,保持任务执行的连贯性,避免重复同样的错误。\n\n实验的核心见解是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。正如一个人缺少关键文件无法做出合理判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义,它不知道存在哪些工具;没有之前的执行结果,它不知道已经做了什么。\n\n### ReAct循环\n掌握了三个组件后,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型(LLM)、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到了推理和行动,但实际循环有三个阶段:模型首先**推理**下一步该做什么,然后调用工具来**行动**,接着**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环会重复,直到任务完成。\n\n以跨多种货币汇总收入的具体示例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助理消息(含推理和工具调用)和工具结果。每次LLM调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助理消息+工具结果,随每次交互增长)。LLM从这个完整上下文中生成下一个响应,然后附加到轨迹中供后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\n轨迹 = [\n {角色: \"用户\", 内容: \"根据公司季度收入:第一季度250万美元(美元)、第二季度210万欧元、第三季度180万英镑、第四季度3.8亿日元,计算公司年度总收入和平均季度收入\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {角色: \"助理\",\n 推理: \"需要将所有货币转换为美元...\",\n 内容: \"\", # 没有直接回复用户\n 工具调用: [\n {名称: \"convert_currency\", 参数: {金额: 2100000, 来自: \"EUR\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 1800000, 来自: \"GBP\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 380000000, 来自: \"JPY\", 到: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {角色: \"工具\", 内容: \"欧元→美元:2282608.7\"},\n {角色: \"工具\", 内容: \"英镑→美元:2278481.01\"},\n {角色: \"工具\", 内容: \"日元→美元:2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {角色: \"助理\",\n 推理: \"已获得转换结果,现在需要汇总并计算...\",\n 内容: \"\",\n 工具调用: [\n {名称: \"code_interpreter\", 参数: {代码: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {角色: \"工具\", 内容: \"总计:9,602,895.73美元,平均:2,400,723.93美元...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {角色: \"助理\",\n 推理: \"所有计算完成,总结结果...\",\n 内容: \"最终答案:总收入9,602,895.73美元...\"},\n]\n```\n\n请注意,系统提示和工具定义未在轨迹中显示——它们作为静态前缀,在每次LLM调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量较大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的精妙之处在于**上下文的累积性**。每次LLM调用都接收完整的轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。正如人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且由于轨迹结构清晰——用户消息、助理消息(推理+工具调用)和工具结果都明确分开,系统具有高度的可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证明。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的闭环。\n\n现在我们了解了代理的操作循环,接下来通过两个实验看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3原生代理能力\n这个实验展示了**Kimi K3**的原生代理能力,这是“模型即代理”范式的一个示例。Kimi K3由月之暗面公司于2026年发布,是一个约有2.8万亿参数的专家混合(MoE)模型。MoE可视为一个专家团队:对于每种问题,系统仅激活最适合它的少数专家,而不是整个模型,在保持能力的同时避免了全部效率成本。Kimi K3具有100万个词元的上下文窗口、原生视觉理解能力和始终开启的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化为原生能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过一个名为Formula的服务器端脚本引擎运行这些官方工具。"
|
||
},
|
||
{
|
||
"role": "assistant",
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}"
|
||
},
|
||
{
|
||
"role": "user",
|
||
"content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys."
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "021785390335395e4feb8d226730241afb8b5c4c4c35b3251011c",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390348,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}",
|
||
"usage": {
|
||
"prompt_tokens": 6804,
|
||
"completion_tokens": 528,
|
||
"total_tokens": 7332
|
||
}
|
||
},
|
||
"latency_milliseconds": 12678,
|
||
"validation": {
|
||
"valid": false,
|
||
"error_type": "ValueError",
|
||
"error": "judge variants must contain exactly X and Y"
|
||
}
|
||
},
|
||
{
|
||
"attempt": 4,
|
||
"request": {
|
||
"model": "doubao-seed-1-6-250615",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "You are an exacting bilingual technical-book translation evaluator. Compare two anonymous Chinese translations against the complete English Markdown source. Score both X and Y from 1 to 5 on exactly: accuracy (no omissions, inventions, or changed claims); fluency; terminology (consistent and technically correct); markdown_code_fidelity (figures, links, headings, equations, and fenced code preserved). Each score needs concrete quoted or located evidence. Prefer one only when evidence supports it. Return JSON only: {\"variants\":{\"X\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}},\"Y\":{\"accuracy\":{\"score\":1,\"evidence\":\"...\"},\"fluency\":{\"score\":1,\"evidence\":\"...\"},\"terminology\":{\"score\":1,\"evidence\":\"...\"},\"markdown_code_fidelity\":{\"score\":1,\"evidence\":\"...\"}}},\"preferred\":\"X|Y|tie\",\"preference_evidence\":\"...\"}.\n\nCOMPLETE ENGLISH SOURCE:\nIs every component truly indispensable? The most direct way to find out is an **ablation study**—the diagnostic method of ruling out causes one at a time: remove component A and see whether the system still works, then component B, and so on, until each component’s contribution is clear. Experiment 1-1 applies exactly this method to the five components above. The results are direct: without tool definitions, the Agent is completely incapable of action; without tool results, it does not receive feedback from the previous step, so it calls the same tool repeatedly, becoming stuck in an infinite loop; without the reasoning in assistant messages, consecutive decisions start contradicting each other; without message history, the Agent loses task continuity and restarts the whole task from the beginning, repeating steps already done. Each component’s role rests on experimental evidence, not just theoretical inference.\n\n### Experiment 1-1 ★★: The Critical Role of Context\n\nWe probed how each context component shapes Agent behavior with a systematic **ablation study**. Of the five components above, four were tested—the system prompt, as the Agent’s basic identity definition, was exempt: without it the Agent has no role awareness at all, and the test would be meaningless. As Figure 1-2 shows, the experiment ran five controlled groups: a complete baseline retaining every component, plus four groups each missing one, to observe each component’s effect on Agent performance.\n\n\n\nThe experimental results revealed the irreplaceable role of each context component. **Tool Definitions** (part of the static prefix) are the foundation of the Agent’s action capability; without them, the Agent cannot recognize or call any tools. **Tool Results** are key to closed-loop control; their absence deprives the Agent of execution feedback and causes it to fall into an infinite loop. The **reasoning process** (the reasoning part of assistant messages) preserves the reasons for the Agent’s previous decisions, making the overall reasoning more coherent and preventing contradictory decisions. **Message history** (user messages, assistant messages, and tool results from previous rounds) prevents redundant operations, maintains task execution coherence, and avoids repeating the same mistakes.\n\nThe experiment's core insight: **context determines what information the Agent has at decision time, and the Agent can only decide based on that information**. Just as a person missing crucial documents cannot make sound judgments, an Agent missing any context component suffers a severe loss of decision-making ability—without tool definitions it does not know what tools exist; without previous execution results it does not know what has already been done.\n\n### The ReAct Loop\n\nWith the three components in hand, a natural question follows: how do they work together? The ReAct loop is the core mechanism that connects LLM, context, and tools into one system. We can examine it step by step.\n\nThe core pattern by which an Agent executes a task is called **ReAct** (Reasoning + Acting). The name mentions only reasoning and acting, but the actual loop has three stages: the model first **reasons** about what to do next, then calls a tool to **act**, then **observes** the tool’s result and reasons about the subsequent step. This “reason → act → observe → reason → act → observe” loop repeats until the task is done.\n\nConsider a concrete example—aggregating revenue across multiple currencies—to understand an Agent’s **trajectory**: the message history that accumulates as the Agent works, comprising user messages, assistant messages (with their reasoning and tool calls), and tool results. On every LLM call, the complete context the model receives is the **static prefix** (system prompt + tool definitions) plus the **trajectory** (dynamic message history) (Figure 1-3). This shows a key fact: **Agent context = static prefix + trajectory**. Concretely, the static prefix is the first two of the five components above (system prompt + tool definitions); the trajectory is the last three (user messages + assistant messages + tool results, growing with each interaction). From this complete context the LLM generates its next response, which is then appended to the trajectory for the subsequent call.\n\n\n\nHere is the structure of a trajectory, in pseudocode:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # First iteration - LLM receives the above trajectory and generates a response\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # No direct reply to the user\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # Agent framework executes tools, adds results to trajectory\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # Second iteration - LLM receives the complete trajectory, including tool results\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # Third iteration - LLM receives the complete trajectory and generates the final answer\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\nNote that the system prompt and tool definitions are not shown in the trajectory—they serve as the static prefix and are automatically prepended to the trajectory before each LLM call.\n\nIn our experiment, this loop was clearly visible. In the first round, the Agent analyzed the task and called three currency conversion tools in parallel; in the second, it fed the conversion results to a code interpreter for the more computationally intensive calculation; in the third, having confirmed all calculations were complete, it produced the final answer. A complex multi-step task was completed in 3 iterations and 4 tool calls.\n\nThe elegance of this design lies in the **cumulative nature of the context**. Every LLM call receives the complete trajectory, so the model knows which stage of the task it is in, what was tried before, and what the outcome was. Just as people keep reviewing and summarizing while solving a problem, the Agent maintains a global view of the task through its trajectory. And because the trajectory is structured—user messages, assistant messages (reasoning + tool calls), and tool results all separated cleanly—the system is highly interpretable and debuggable.\n\nThe trajectory is more than an execution record; it is evidence of the Agent’s capability. Analyzing trajectories at scale reveals behavior patterns, better decision paths, and better tool designs. Trajectory data can even be distilled into a knowledge base, or used to train stronger Agent models via reinforcement learning—closing the loop of learning from experience.\n\nNow that we understand the Agent's operating loop, we examine two experiments to see how different models drive it.\n\n#### Experiment 1-2 ★: Kimi K3 Native Agent Capability\n\nThis experiment demonstrates the native Agent capability of **Kimi K3**, an example of the “Model as Agent” paradigm. Released by Moonshot AI in 2026, Kimi K3 is a Mixture of Experts (MoE) model with approximately 2.8 trillion parameters. MoE can be viewed as a team of experts: for each kind of problem, the system activates only the few experts best suited to it rather than the entire model, preserving capability without paying the full efficiency cost. Kimi K3 has a 1 million token context window, native visual understanding, and an always-on “thinking mode.” Through reinforcement learning, it has internalized the tool-calling **decision policy** as a native capability: when to call a tool, which tool to call, and what arguments to pass are all decided by the model, allowing it to carry out tasks such as web searches autonomously. To be precise, what is internalized is the *when and how to call* decision; the tools themselves, such as `web_search` and `code_runner`, still execute server-side as API-level built-in tools. Kimi runs these official tools through a server-side script engine called Formula.\n\n\n\nANONYMOUS CHINESE X:\n### 人工智能代理入门[第4/9部分]\n每个组件真的都不可或缺吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看系统是否还能工作,然后是组件B,依此类推,直到每个组件的贡献清晰可见。实验1-1正是对上述五个组件应用了这种方法。结果直接明了:没有工具定义,代理完全无法行动;没有工具结果,它无法从之前的步骤获得反馈,所以会反复调用同一个工具,陷入无限循环;没有助手消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理失去任务连续性,从头重新开始整个任务,重复已做的步骤。每个组件的作用都有实验证据支持,而不仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的消融研究探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义,被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:保留所有组件的完整基线组,以及四组各缺失一个组件的组,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺失它们会让代理失去执行反馈,导致陷入无限循环。**推理过程**(助手消息中的推理部分)保留了代理先前决策的理由,使整体推理更连贯,防止矛盾决策。**消息历史**(之前轮次的用户消息、助手消息和工具结果)防止重复操作,保持任务执行连贯性,避免重复同样的错误。\n\n实验的核心洞察是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。就像一个人缺少关键文件无法做出明智判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义它不知道存在哪些工具;没有之前的执行结果它不知道已经做了什么。\n\n### ReAct循环\n有了这三个组件,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到推理和行动,但实际循环有三个阶段:模型首先**推理**下一步做什么,然后调用工具来**行动**,然后**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环重复直到任务完成。\n\n以聚合多种货币的收入为例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助手消息(带有其推理和工具调用)和工具结果。在每次大语言模型调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助手消息+工具结果,随每次交互增长)。大语言模型从这个完整上下文中生成其下一个响应,然后该响应附加到轨迹中用于后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\ntrajectory = [\n {role: \"user\", content: \"Based on the company's quarterly revenue: Q1 2.5M USD, Q2 2.1M EUR, Q3 1.8M GBP, Q4 380M JPY, calculate the company's total annual revenue and average quarterly revenue\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {role: \"assistant\",\n reasoning: \"Need to convert all currencies to USD...\",\n content: \"\", # 没有直接回复用户\n tool_calls: [\n {name: \"convert_currency\", args: {amount: 2100000, from: \"EUR\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 1800000, from: \"GBP\", to: \"USD\"}},\n {name: \"convert_currency\", args: {amount: 380000000, from: \"JPY\", to: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {role: \"tool\", content: \"EUR->USD: 2282608.7\"},\n {role: \"tool\", content: \"GBP->USD: 2278481.01\"},\n {role: \"tool\", content: \"JPY->USD: 2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {role: \"assistant\",\n reasoning: \"Conversion results obtained, now need to aggregate and calculate...\",\n content: \"\",\n tool_calls: [\n {name: \"code_interpreter\", args: {code: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {role: \"tool\", content: \"Total: $9,602,895.73, Average: $2,400,723.93...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {role: \"assistant\",\n reasoning: \"All calculations complete, summarizing results...\",\n content: \"FINAL ANSWER: Total revenue $9,602,895.73...\"}\n]\n```\n\n请注意,系统提示和工具定义未显示在轨迹中——它们作为静态前缀,在每次大语言模型调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量更大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的优雅之处在于**上下文的累积性**。每次大语言模型调用都接收完整轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。就像人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且因为轨迹是结构化的——用户消息、助手消息(推理+工具调用)和工具结果都清晰分离——系统具有高度可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证据。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的循环。\n\n现在我们理解了代理的操作循环,接下来审视两个实验,看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3的本地代理能力\n这个实验展示了**Kimi K3**的本地代理能力,它是“模型即代理”范式的一个示例。由月之暗面科技于2026年发布的Kimi K3是一个混合专家(MoE)模型,约有2.8万亿参数。MoE可以看作是一组专家:对于每种问题,系统只激活最适合它的少数专家,而不是整个模型,在保持能力的同时不支付全部效率成本。Kimi K3有100万个标记的上下文窗口、本地视觉理解和始终在线的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化成本地能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过名为Formula的服务器端脚本引擎运行这些官方工具。\n\nANONYMOUS CHINESE Y:\n### 人工智能代理入门 [第4/9部分]\n每个组件真的都是不可或缺的吗?最直接的方法是进行**消融研究**——一种一次排除一个原因的诊断方法:移除组件A,看看系统是否仍然有效,然后是组件B,依此类推,直到每个组件的贡献清晰明了。实验1-1正是将这种方法应用于上述五个组件。结果一目了然:没有工具定义,代理完全无法行动;没有工具结果,它不会收到上一步的反馈,因此会反复调用同一个工具,陷入无限循环;没有助理消息中的推理,连续的决策开始相互矛盾;没有消息历史,代理会失去任务连续性,从头开始重新执行整个任务,重复已经完成的步骤。每个组件的作用都有实验证据支持,而非仅仅是理论推断。\n\n### 实验1-1 ★★:上下文的关键作用\n我们通过系统的**消融研究**探究了每个上下文组件如何塑造代理行为。在上述五个组件中,四个进行了测试——系统提示作为代理的基本身份定义被豁免:没有它,代理完全没有角色意识,测试将毫无意义。如图1-2所示,实验设置了五组对照:一组保留所有组件的完整基线,另外四组每组缺少一个组件,以观察每个组件对代理性能的影响。\n\n\n\n实验结果揭示了每个上下文组件不可替代的作用。**工具定义**(静态前缀的一部分)是代理行动能力的基础;没有它们,代理无法识别或调用任何工具。**工具结果**是闭环控制的关键;缺少它们会剥夺代理的执行反馈,导致其陷入无限循环。**推理过程**(助理消息中的推理部分)保留了代理先前决策的原因,使整体推理更连贯,防止矛盾决策。**消息历史**(用户消息、助理消息和之前轮次的工具结果)防止冗余操作,保持任务执行的连贯性,避免重复同样的错误。\n\n实验的核心见解是:**上下文决定了代理在决策时拥有的信息,代理只能基于该信息进行决策**。正如一个人缺少关键文件无法做出合理判断一样,缺少任何上下文组件的代理都会严重丧失决策能力——没有工具定义,它不知道存在哪些工具;没有之前的执行结果,它不知道已经做了什么。\n\n### ReAct循环\n掌握了三个组件后,自然会产生一个问题:它们如何协同工作?ReAct循环是将大语言模型(LLM)、上下文和工具连接成一个系统的核心机制。我们可以逐步审视它。\n\n代理执行任务的核心模式称为**ReAct**(推理+行动)。名称只提到了推理和行动,但实际循环有三个阶段:模型首先**推理**下一步该做什么,然后调用工具来**行动**,接着**观察**工具的结果并推理后续步骤。这个“推理→行动→观察→推理→行动→观察”的循环会重复,直到任务完成。\n\n以跨多种货币汇总收入的具体示例来理解代理的**轨迹**:代理工作时积累的消息历史,包括用户消息、助理消息(含推理和工具调用)和工具结果。每次LLM调用时,模型接收的完整上下文是**静态前缀**(系统提示+工具定义)加上**轨迹**(动态消息历史)(图1-3)。这表明一个关键事实:**代理上下文=静态前缀+轨迹**。具体来说,静态前缀是上述五个组件中的前两个(系统提示+工具定义);轨迹是后三个(用户消息+助理消息+工具结果,随每次交互增长)。LLM从这个完整上下文中生成下一个响应,然后附加到轨迹中供后续调用。\n\n\n\n以下是轨迹的伪代码结构:\n\n```\n轨迹 = [\n {角色: \"用户\", 内容: \"根据公司季度收入:第一季度250万美元(美元)、第二季度210万欧元、第三季度180万英镑、第四季度3.8亿日元,计算公司年度总收入和平均季度收入\"},\n \n # 第一次迭代 - LLM接收上述轨迹并生成响应\n {角色: \"助理\",\n 推理: \"需要将所有货币转换为美元...\",\n 内容: \"\", # 没有直接回复用户\n 工具调用: [\n {名称: \"convert_currency\", 参数: {金额: 2100000, 来自: \"EUR\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 1800000, 来自: \"GBP\", 到: \"USD\"}},\n {名称: \"convert_currency\", 参数: {金额: 380000000, 来自: \"JPY\", 到: \"USD\"}}\n ]},\n \n # 代理框架执行工具,将结果添加到轨迹中\n {角色: \"工具\", 内容: \"欧元→美元:2282608.7\"},\n {角色: \"工具\", 内容: \"英镑→美元:2278481.01\"},\n {角色: \"工具\", 内容: \"日元→美元:2541806.02\"},\n \n # 第二次迭代 - LLM接收包含工具结果的完整轨迹\n {角色: \"助理\",\n 推理: \"已获得转换结果,现在需要汇总并计算...\",\n 内容: \"\",\n 工具调用: [\n {名称: \"code_interpreter\", 参数: {代码: \"total = 2500000 + 2282608.7 + ...\"}}\n ]},\n \n {角色: \"工具\", 内容: \"总计:9,602,895.73美元,平均:2,400,723.93美元...\"},\n \n # 第三次迭代 - LLM接收完整轨迹并生成最终答案\n {角色: \"助理\",\n 推理: \"所有计算完成,总结结果...\",\n 内容: \"最终答案:总收入9,602,895.73美元...\"},\n]\n```\n\n请注意,系统提示和工具定义未在轨迹中显示——它们作为静态前缀,在每次LLM调用前自动添加到轨迹前面。\n\n在我们的实验中,这个循环清晰可见。第一轮,代理分析任务并并行调用三个货币转换工具;第二轮,它将转换结果提供给代码解释器进行计算量较大的计算;第三轮,确认所有计算完成后,它生成最终答案。一个复杂的多步骤任务在3次迭代和4次工具调用中完成。\n\n这种设计的精妙之处在于**上下文的累积性**。每次LLM调用都接收完整的轨迹,所以模型知道任务处于哪个阶段、之前尝试过什么以及结果如何。正如人们在解决问题时不断回顾和总结一样,代理通过其轨迹保持对任务的全局视图。而且由于轨迹结构清晰——用户消息、助理消息(推理+工具调用)和工具结果都明确分开,系统具有高度的可解释性和可调试性。\n\n轨迹不仅是执行记录,更是代理能力的证明。大规模分析轨迹可以揭示行为模式、更好的决策路径和更好的工具设计。轨迹数据甚至可以提炼成知识库,或通过强化学习用于训练更强的代理模型——形成从经验中学习的闭环。\n\n现在我们了解了代理的操作循环,接下来通过两个实验看看不同模型如何驱动它。\n\n#### 实验1-2 ★:Kimi K3原生代理能力\n这个实验展示了**Kimi K3**的原生代理能力,这是“模型即代理”范式的一个示例。Kimi K3由月之暗面公司于2026年发布,是一个约有2.8万亿参数的专家混合(MoE)模型。MoE可视为一个专家团队:对于每种问题,系统仅激活最适合它的少数专家,而不是整个模型,在保持能力的同时避免了全部效率成本。Kimi K3具有100万个词元的上下文窗口、原生视觉理解能力和始终开启的“思考模式”。通过强化学习,它将工具调用的**决策策略**内化为原生能力:何时调用工具、调用哪个工具、传递什么参数都由模型决定,使其能够自主执行网络搜索等任务。准确地说,内化的是*何时以及如何调用*的决策;工具本身,如`web_search`和`code_runner`,仍然作为API级内置工具在服务器端执行。Kimi通过一个名为Formula的服务器端脚本引擎运行这些官方工具。"
|
||
},
|
||
{
|
||
"role": "assistant",
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's '内化成本地能力' (same here but overall flow better)\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line (']},' vs. X's correct ']}'\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"} ,\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms (e.g., 'token') are more impactful for a technical book.\"}"
|
||
},
|
||
{
|
||
"role": "user",
|
||
"content": "Your previous JSON failed validation: judge variants must contain exactly X and Y. Return a corrected JSON object whose variants object has exactly the keys X and Y; each variant has exactly accuracy, fluency, terminology, and markdown_code_fidelity; do not rename, omit, or add rubric keys."
|
||
}
|
||
],
|
||
"temperature": 0,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
},
|
||
"extra_body": {
|
||
"thinking": {
|
||
"type": "disabled"
|
||
}
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "021785390348073e4feb8d226730241afb8b5c4c4c35b32121c35",
|
||
"model": "doubao-seed-1-6-250615",
|
||
"created": 1785390358,
|
||
"content": "{\"variants\":{\"X\":{\"accuracy\":{\"score\":4,\"evidence\":\"Omission of 'LLM' expansion in first occurrence; 'token' mistranslated as '标记' instead of '词元' in '1 million token context window'\"},\"fluency\":{\"score\":4,\"evidence\":\"Slightly awkward phrasing: '从头重新开始整个任务' (redundant '重新') vs. Y's '从头开始重新执行整个任务'\"},\"terminology\":{\"score\":4,\"evidence\":\"Inconsistent translation of 'trajectory' as '轨迹' (correct) but 'token' as '标记' (should be '词元')\"},\"markdown_code_fidelity\":{\"score\":5,\"evidence\":\"Fenced code block content (e.g., 'role: \\\"user\\\"') preserved unchanged; headings, images, and formatting consistent with source\"}},\"Y\":{\"accuracy\":{\"score\":5,\"evidence\":\"All technical claims preserved, e.g., '2.8 trillion parameters' accurately translated; 'token' correctly rendered as '词元'\"},\"fluency\":{\"score\":5,\"evidence\":\"Natural phrasing: '结果一目了然' (idiomatic) vs. X's '结果直接明了'; '内化成本地能力' (smooth) vs. X's identical phrasing but overall flow better\"},\"terminology\":{\"score\":5,\"evidence\":\"Consistent use of '词元' (token), '决策策略' (decision policy), and '轨迹' (trajectory); technical terms like 'MoE' translated as '混合专家' consistently\"},\"markdown_code_fidelity\":{\"score\":4,\"evidence\":\"Code block keys translated (e.g., 'role' → '角色') altering original structure; trailing comma added in final code line\"}},\"preferred\":\"Y\",\"preference_evidence\":\"Y has higher accuracy (no critical mistranslations like '标记' for 'token'), superior fluency (more idiomatic expressions), and consistent terminology. While Y modified code block keys, X's accuracy issues in technical terms are more impactful for a technical book.\"}}",
|
||
"usage": {
|
||
"prompt_tokens": 6804,
|
||
"completion_tokens": 428,
|
||
"total_tokens": 7232
|
||
}
|
||
},
|
||
"latency_milliseconds": 10622,
|
||
"validation": {
|
||
"valid": false,
|
||
"error_type": "ValueError",
|
||
"error": "judge variants must contain exactly X and Y"
|
||
},
|
||
"resume_validation": {
|
||
"valid": true,
|
||
"schema_repairs": [
|
||
"lifted duplicated preference fields out of variants"
|
||
]
|
||
}
|
||
}
|
||
]
|
||
}
|