Files
ai-agent-book/chapter10/book-translation/sample_book/chapter2.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

25 lines
1.2 KiB
Markdown

# Chapter 2: The Transformer and Attention
Modern language models are built on the **transformer** architecture. Its
central idea is **attention**: instead of reading a sequence strictly left to
right, the model lets every token look at every other token and decide which
ones matter. This is why a transformer can connect a pronoun to a noun that
appeared many tokens earlier.
Attention works on the **embedding** of each token. For every token the model
computes three vectors — a query, a key, and a value — and uses them to weigh
how much each token should attend to the others.
```python
def attention(query, key, value):
scores = query @ key.T # similarity between tokens
weights = softmax(scores) # attention weights
return weights @ value # weighted embedding
```
Because attention compares every token with every other token, its cost grows
quickly as the prompt gets longer. This is the root cause of the latency
problems we will attack in the next chapter. Still, attention is what gives the
transformer its power: during inference, it lets the model route information
flexibly across the whole prompt rather than through a fixed pipeline.