Files
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

85 lines
2.8 KiB
Python

"""
Configuration for Log Sanitization with Local LLM
"""
import os
from pathlib import Path
# Ollama Configuration
# 默认使用 0.6B 超小模型,呼应本章“小模型也能胜任结构化任务”的论点,
# 且可在 CPU / 消费级设备上运行;可用 --model 覆盖为 qwen3:1.7b、qwen3:4b 等。
OLLAMA_MODEL = "qwen3:0.6b"
OLLAMA_TEMPERATURE = 0.1 # Low temperature for consistent detection
# Paths
PROJECT_ROOT = Path(__file__).parent
OUTPUT_DIR = PROJECT_ROOT / "output"
OUTPUT_DIR.mkdir(exist_ok=True)
# Performance Metrics Configuration
METRICS_FILE = OUTPUT_DIR / "performance_metrics.json"
# Evaluation Framework Path
# The user-memory-evaluation framework lives in chapter3, not chapter2.
# PROJECT_ROOT = chapter3/log-sanitization, so go up two levels to the repo root.
EVAL_FRAMEWORK_PATH = PROJECT_ROOT.parent.parent / "chapter3" / "user-memory-evaluation"
# System Prompt for PII Detection
SYSTEM_PROMPT = """You are a privacy protection agent that detects Level 3 PII.
Level 3 PII includes:
- Social Security Numbers (SSN) - format: XXX-XX-XXXX or XXXXXXXXX
- Credit Card Numbers - format: XXXX XXXX XXXX XXXX or 16 digits
- Credit Card Expiry Date and CVV
- Bank Account Numbers
- Full Residential Addresses
- Medical Record Numbers
- Medical Diagnoses and Treatment Details
- Prescription Information
- Driver's License Numbers
- Passport Numbers
- Financial PINs
- Tax ID Numbers
- Health Insurance IDs
- Biometric Data
- Usernames for Financial Accounts
- Passwords
Analyze the conversation and return JSON with a pii_items array. Each item must include:
- type: the PII category label
- value: the exact sensitive substring copied verbatim from the input text
Do not include labels or explanations inside value. NEVER use placeholders."""
USER_PROMPT_TEMPLATE = """Analyze the following conversation for Level 3 PII:
{conversation_text}"""
# JSON Schema for structured output
PII_DETECTION_SCHEMA = {
"type": "object",
"properties": {
"pii_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"type": {
"type": "string",
"description": "PII category label, e.g. ssn, credit_card_number, email"
},
"value": {
"type": "string",
"description": "Exact sensitive substring copied verbatim from the input text"
}
},
"required": ["type", "value"],
"additionalProperties": False
},
"description": "Array of structured PII items. The value field must be copied verbatim from the input text."
}
},
"required": ["pii_items"],
"additionalProperties": False
}