Files
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

56 lines
3.2 KiB
Markdown

# 实验 9-1:客服 Agent 的三层轨迹验证器
本实验对应正文的“从运行轨迹中获得学习信号”。它不把用户满意度或一个总分当作学习信号,而是依次核对环境结果、执行过程与语言质量,并在每个失败维度中保留证据轮次。
`verifier.py` 实现三层结构:结果层读取最终订单状态;过程层检查业务规则、隐私、事实依据和承诺—行动一致性;质量层按“表达质量、合规变通”Rubric 评价开放性指标。示例默认使用确定性的 `HeuristicQualityJudge`,所以不需要 API Key;项目也提供遵循同一 `QualityJudge` 接口的真实 LLM 实现,下两层仍坚持使用环境真值和程序规则。
`sample_trajectories.json` 包含正常退款、虚假承诺、违规泄露和过度拒绝四类轨迹,并带有专家标签。`calibration.py` 按维度报告违规识别的精确率、召回率与标签一致率。`demo.py` 还对比了只有一个总分的输出与带证据的多维诊断。
## Code map
- **Run first:** python demo.py (deterministic HeuristicQualityJudge, no API key).
- **Start here:** verifier.py composes the result, process and quality layers.
- **Core behavior:** customer_service_env.py::run_case supplies environment truth; calibration.py compares dimensions with expert labels.
- **State / protocol:** sample_trajectories.json, structured verdict schema and evidence turns.
- **Verifier:** test_verifier.py plus calibration precision/recall; LLM quality judging never replaces the first two code gates.
- **Experiment variable:** single scalar score versus dimensioned verdict with evidence/confidence.
- **Skip on first pass:** provider client and demo formatting.
运行方法:
```bash
python demo.py
python -m unittest -v test_verifier.py
```
以上是确定性校准路径。若要真实调用 LLM 评价表达质量与合规变通:
```bash
# 从仓库根目录开始:使用共享的第 8 章环境
uv sync --locked --python 3.12 --extra ch8
# Apple Silicon macOS 需要 macOS 14+(锁文件中的 bitsandbytes wheel 要求);
# 更早的 macOS 请使用下方单项目兼容路径。
# 切换目录前先激活环境:
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# Windows cmd: .venv\Scripts\activate.bat
# 未安装 uv 时可用 pip 兜底:
# python -m pip install -e ".[ch8]"
cd chapter8/trajectory-verifier
# 迁移期间仍支持单项目兼容路径:
# python -m pip install -r requirements.txt
cp env.example .env
export OPENAI_API_KEY=your_api_key_here
python demo.py --judge llm --model gpt-5.6
```
真实模式使用 OpenAI Responses API,并要求模型按相同 schema 返回逐维结论、证据轮次和置信度;环境结果与过程规则两层仍由代码判断。该命令会产生真实 API 费用,输出可能随模型版本变化,应继续用专家标签检查每个维度,而不能只观察总分。
真实系统应扩大专家校准集,并把低置信度或高风险轨迹交给第二个验证器或人工复核。样例中的 `quality_facts` 是离线实验对 LLM 判读结果的显式表示,并不意味着生产系统可以预先获得这些字段。