Files
ai-agent-book/chapter8/cot-distillation/test_load_verified_messages_colon.py
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

33 lines
1.1 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Regression: load_verified_messages must accept fullwidth colon and case-insensitive Final Answer."""
import json
from pathlib import Path
from train_student import load_verified_messages
def test_load_verified_messages_fullwidth_colon(tmp_path: Path):
sample = {
"messages": [
{"role": "user", "content": "What is 2 + 2?"},
{"role": "assistant", "content": "<think>\n2+2=4\n</think>\n\nFinal Answer4"},
]
}
dataset = tmp_path / "dataset.jsonl"
dataset.write_text(json.dumps(sample) + "\n", encoding="utf-8")
rows = load_verified_messages(dataset)
assert len(rows) == 1
assert rows[0][1]["content"].endswith("Final Answer4")
def test_load_verified_messages_lowercase_colon(tmp_path: Path):
sample = {
"messages": [
{"role": "user", "content": "What is 2 + 2?"},
{"role": "assistant", "content": "<think>\n2+2=4\n</think>\n\nfinal answer: 4"},
]
}
dataset = tmp_path / "dataset.jsonl"
dataset.write_text(json.dumps(sample) + "\n", encoding="utf-8")
rows = load_verified_messages(dataset)
assert len(rows) == 1