Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
68 lines
7.1 KiB
JSON
68 lines
7.1 KiB
JSON
{
|
||
"signature": "5db468b21d65ce64103fec1ce935870b4152125c2e637403437f8514521269e6",
|
||
"parsed": {
|
||
"narration": "前面几页我们把架构和训练细节都讲完了,那效果到底行不行?这一页就是论文交出的成绩单。请大家看屏幕上这张机器翻译结果表,重点看加粗的第三行:Transformer big 在英德任务上拿到28.4 BLEU,比此前最好的集成模型高出约2个点;英法方向达到41.8,同样是新的最优纪录。更值得关注的是最右边训练成本这一列,它只用了2.3乘10的19次方的浮点运算量,连对手的三分之一都不到。也就是说,纯注意力机制不但质量更高,还更省钱。这样的架构在其他任务上还能打吗?下一页接着看。",
|
||
"visual_references": [
|
||
"标题 Machine Translation Results",
|
||
"结果表格中加粗的 Transformer (big) 行",
|
||
"EN-DE 列的 28.4 与 EN-FR 列的 41.8",
|
||
"Training Cost (FLOPs) 列的 2.3×10^19 与上方两行数值对比"
|
||
]
|
||
},
|
||
"receipt": {
|
||
"provider": "Moonshot",
|
||
"request": {
|
||
"model": "kimi-k3",
|
||
"messages": [
|
||
{
|
||
"role": "user",
|
||
"content": "你是一位严谨而自然的中文技术讲师,正在讲解论文 Attention Is All You Need。\n这是第 9/12 页。下面给出该页真实 Slidev 源码:\n<slide>\n## Machine Translation Results\n\n| Model | EN-DE BLEU | EN-FR BLEU | Training Cost (FLOPs) |\n|-------|------------|------------|-----------------------|\n| GNMT + RL Ensemble | 26.30 | 41.16 | $1.8 \\cdot 10^{20}$ |\n| ConvS2S Ensemble | 26.36 | 41.29 | $7.7 \\cdot 10^{19}$ |\n| **Transformer (big)** | **28.4** | **41.8** | **$2.3 \\cdot 10^{19}$** |\n\n- Transformer outperforms all previous state-of-the-art models\n- Achieves better results with significantly lower training cost\n- 28.4 BLEU on EN-DE (↑2 BLEU over previous best)\n- 41.8 BLEU on EN-FR (new state-of-the-art)\n</slide>\n请写一段 180–240 个中文字符左右的口语讲解。必须满足:\n1. 用引导性叙事解释“为什么”和“它与前后页的关系”,不能逐条照读幻灯片;\n2. 明确指向当前屏幕上的可见元素(标题、公式、表格、数值或图片);有图片/表格/公式时必须说出观众该看哪里;\n3. 不虚构源码中没有的数值或结论;开头/结尾与页序匹配;\n4. 只返回 JSON:{\"narration\":\"...\",\"visual_references\":[\"...\"]}。\n"
|
||
}
|
||
],
|
||
"temperature": 1,
|
||
"max_tokens": 1200,
|
||
"response_format": {
|
||
"type": "json_object"
|
||
}
|
||
},
|
||
"response": {
|
||
"id": "chatcmpl-6a6a7f6c89b4c0360c645bb5",
|
||
"choices": [
|
||
{
|
||
"finish_reason": "stop",
|
||
"index": 0,
|
||
"logprobs": null,
|
||
"message": {
|
||
"content": "{\"narration\":\"前面几页我们把架构和训练细节都讲完了,那效果到底行不行?这一页就是论文交出的成绩单。请大家看屏幕上这张机器翻译结果表,重点看加粗的第三行:Transformer big 在英德任务上拿到28.4 BLEU,比此前最好的集成模型高出约2个点;英法方向达到41.8,同样是新的最优纪录。更值得关注的是最右边训练成本这一列,它只用了2.3乘10的19次方的浮点运算量,连对手的三分之一都不到。也就是说,纯注意力机制不但质量更高,还更省钱。这样的架构在其他任务上还能打吗?下一页接着看。\",\"visual_references\":[\"标题 Machine Translation Results\",\"结果表格中加粗的 Transformer (big) 行\",\"EN-DE 列的 28.4 与 EN-FR 列的 41.8\",\"Training Cost (FLOPs) 列的 2.3×10^19 与上方两行数值对比\"]}",
|
||
"refusal": null,
|
||
"role": "assistant",
|
||
"annotations": null,
|
||
"audio": null,
|
||
"function_call": null,
|
||
"tool_calls": null,
|
||
"reasoning_content": "The task: write a Chinese oral narration for slide 9/12 of Attention Is All You Need, about Machine Translation Results. Need 180-240 Chinese characters, guiding narrative explaining \"why\" and relation to previous/next pages, pointing to visible elements (table, numbers), no fabrication, JSON only output.\n\nSlide 9 is about translation results — previous pages presumably covered architecture/training details, so this page shows the payoff. Point to the table: EN-DE BLEU 28.4 vs 26.30/26.36, EN-FR 41.8 vs 41.16/41.29, training cost FLOPs 2.3e19 vs 1.8e20/7.7e19. Emphasize better quality at lower cost.\n\nWrite ~200 characters. Include visual_references list: \"Machine Translation Results 标题\", \"BLEU 对比表格\", \"28.4 / 41.8 加粗数值\", \"Training Cost (FLOPs) 列 2.3·10^19\".\n\nCount characters carefully. Let me draft:\n\n\"讲了这么多结构设计,效果到底怎么样?现在看屏幕这张翻译结果表格。这是第9页,也是整篇论文最有说服力的证据。请大家看表格里加粗的那一行:Transformer big 在英德翻译上拿到28.4 BLEU,比此前最好的集成模型高出约2个点;英法方向达到41.8,同样是新纪录。更关键的是最右边一列——训练成本,它只用了2.3×10^19 次浮点运算,不到对手的三分之一。也就是说,纯注意力架构不仅质量更高,还更便宜。那它在其他任务上表现如何?下一页继续看。\"\n\nCount: roughly let me estimate. That's around 200+ characters. Let me count more precisely.\n\n\"讲了这么多结构设计,效果到底怎么样?\" = 19\n\"现在看屏幕这张翻译结果表格。\" = 14\n\"这是第9页,也是整篇论文最有说服力的证据。\" = 21\n\"请大家看表格里加粗的那一行:\" = 14\n\"Transformer big 在英德翻译上拿到28.4 BLEU,比此前最好的集成模型高出约2个点;\" — count Chinese chars plus some Latin. The requirement says 180–240 个中文字符 — likely total characters. Let me just aim ~210 total.\n\n\"英法方向达到41.8,同样是新纪录。\" = ~17\n\"更关键的是最右边一列——训练成本,它只用了2.3×10的19次方次浮点运算,不到对手的三分之一。\" \n\"也就是说,纯注意力架构不仅质量更高,还更便宜。\"\n\"那它在其他任务上表现如何?下一页继续看。\"\n\nTotal likely ~215. Good.\n\nActually the slide says \"↑2 BLEU over previous best\" so \"高出约2个点\" is fine. Cost comparison: 2.3e19 vs 7.7e19 (ConvS2S) — that's less than a third; vs 1.8e20 GNMT it's about 1/8. Saying \"不到对手的三分之一\" is safe for both.\n\nJSON output only."
|
||
}
|
||
}
|
||
],
|
||
"created": 1785364333,
|
||
"model": "kimi-k3",
|
||
"object": "chat.completion",
|
||
"service_tier": null,
|
||
"system_fingerprint": null,
|
||
"usage": {
|
||
"completion_tokens": 892,
|
||
"prompt_tokens": 484,
|
||
"total_tokens": 1376,
|
||
"completion_tokens_details": {
|
||
"accepted_prediction_tokens": null,
|
||
"audio_tokens": null,
|
||
"reasoning_tokens": 666,
|
||
"rejected_prediction_tokens": null
|
||
},
|
||
"prompt_tokens_details": null
|
||
}
|
||
},
|
||
"elapsed_ms": 33233.501
|
||
}
|
||
}
|