Files
ai-agent-book/chapter5/paper-to-ppt/validation/runs/exp5-4-real-pdf-both-20260730-v5/dual_round4_slides.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

5.2 KiB
Raw Blame History

theme
theme
default

Attention Is All You Need

A Revolutionary Architecture for Sequence Transduction

Ashish Vaswani et al.
NIPS 2017


Abstract

  • First sequence transduction model based solely on attention
  • Dispenses with recurrence and convolutions entirely
  • More parallelizable and faster to train than traditional models
  • New state-of-the-art: 28.4 BLEU (EN-DE) and 41.8 BLEU (EN-FR)
  • Generalizes well to English constituency parsing

Background: Limitations of Traditional Models

Recurrent Models (RNN/LSTM/GRU)

  • Inherently sequential computation → limited parallelization
  • O(n) sequential operations for long-range dependencies

Convolutional Models

  • Fixed kernel size restricts context → requires multiple layers
  • Logarithmic path length for distant connections

Key Innovation: Self-Attention Mechanism

  • Connects all positions with constant operations
  • Enables parallel computation across sequence
  • Directly models long-range dependencies
  • More efficient than RNN/CNN for typical sequence lengths

Transformer Model Architecture

Encoder-decoder structure with stacked self-attention and feed-forward layers.


Encoder Architecture

  • Stack of 6 identical layers
  • Each layer has two sub-layers:
    1. Multi-head self-attention mechanism
    2. Position-wise feed-forward network
  • Residual connections + layer normalization
  • Output dimension: dmodel = 512

Decoder Architecture

  • Stack of 6 identical layers
  • Three sub-layers per layer:
    1. Masked multi-head self-attention (prevents future positions)
    2. Encoder-decoder attention (queries from decoder, keys/values from encoder)
    3. Position-wise feed-forward network
  • Residual connections + layer normalization

Attention Mechanisms

Scaled Dot-Product Attention

\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V
  • Scaling avoids gradient vanishing for large dk

Multi-Head Attention

\text{MultiHead}(Q, K, V) = \text{Concat}(\text{head}_1,...,\text{head}_h)W^O
  • h=8 parallel heads, dk=dv=64 → captures diverse patterns

Positional Encoding

  • Injects sequence order information (no recurrence/convolution)
  • Added to input embeddings (same dmodel dimension)
  • Uses sine/cosine functions with varying frequencies: PE_{(pos,2i)} = \sin(pos/10000^{2i/d_{\text{model}}}) PE_{(pos,2i+1)} = \cos(pos/10000^{2i/d_{\text{model}}})
  • Enables learning of relative position relationships

Why Self-Attention?

Aspect Self-Attention Recurrent Convolutional
Complexity O(n²·d) O(n·d²) O(k·n·d²)
Parallelization O(1) O(n) O(1)
Long-range path length O(1) O(n) O(logk(n))
  • Superior parallelization and dependency modeling

Training Setup

Data & Hardware

  • WMT 2014 EN-DE (4.5M) and EN-FR (36M) sentence pairs
  • 8 NVIDIA P100 GPUs, 12h (base)/3.5d (big model) training

Optimization

  • Adam (β1=0.9, β2=0.98, ϵ=10⁻⁹) with linear warmup (4000 steps) + inverse square root decay
  • Regularization: residual dropout (0.1), label smoothing (0.1)

Machine Translation Results

Model EN-DE BLEU EN-FR BLEU Training Cost (FLOPs)
GNMT + RL Ensemble 26.30 41.16 1.8×10²⁰ / 1.1×10²¹
ConvS2S Ensemble 26.36 41.29 7.7×10¹⁹ / 1.2×10²¹
Transformer (big) 28.4 41.8 2.3×10¹⁹
  • New state-of-the-art with 4× lower training cost

Model Ablations & Generalization

Key Ablations (EN-DE Dev Set)

Variation Dev BLEU Insight
Single head 24.9 Multi-head critical
No dropout 25.3 Regularization needed
Learned pos encoding 25.7 Sinusoidal ≈ learned

Constituency Parsing

  • 91.3 F1 (WSJ only) vs. 91.7 (state-of-the-art)
  • 92.7 F1 (semi-supervised) → strong generalization

Attention Visualization: Long-Distance Dependencies

Encoder self-attention (layer 5) showing how "making" attends to distant words to complete the phrase "making...more difficult".


Attention Visualization: Anaphora Resolution

Attention heads resolving the pronoun "its" to its referent "The Law" with sharp attention focusing.


Limitations

  • Quadratic complexity in sequence length
  • Less efficient for very long sequences
  • Still sequential in generation process

Future Work

  • Local/restricted attention mechanisms
  • Extension to other modalities (images, audio)
  • Non-sequential generation approaches
  • Efficient handling of large inputs/outputs

Conclusion

  • Transformer replaces recurrence/convolution with self-attention
  • Sets new state-of-the-art in machine translation
  • Faster training via parallelization
  • Generalizes well to diverse sequence tasks
  • Foundation for modern attention-based NLP models