Files
ai-agent-book/chapter10/book-translation/sample_book/chapter2.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

1.2 KiB

Chapter 2: The Transformer and Attention

Modern language models are built on the transformer architecture. Its central idea is attention: instead of reading a sequence strictly left to right, the model lets every token look at every other token and decide which ones matter. This is why a transformer can connect a pronoun to a noun that appeared many tokens earlier.

Attention works on the embedding of each token. For every token the model computes three vectors — a query, a key, and a value — and uses them to weigh how much each token should attend to the others.

def attention(query, key, value):
    scores = query @ key.T          # similarity between tokens
    weights = softmax(scores)       # attention weights
    return weights @ value          # weighted embedding

Because attention compares every token with every other token, its cost grows quickly as the prompt gets longer. This is the root cause of the latency problems we will attack in the next chapter. Still, attention is what gives the transformer its power: during inference, it lets the model route information flexibly across the whole prompt rather than through a fixed pipeline.