Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
5.0 KiB
5.0 KiB
theme
| theme |
|---|
| default |
Attention Is All You Need
A Revolutionary Architecture for Sequence Transduction
Ashish Vaswani et al.
NIPS 2017
Background
- Dominant sequence models rely on RNNs/LSTMs/GRUs
- Encoder-decoder architectures with auxiliary attention
- Sequential computation limits parallelization
- Long-range dependencies are challenging to model
Motivation
- Recurrent networks: Inherently sequential, poor parallelization
- Convolutional networks: Limited receptive field, layered dependencies
- Both: Computation grows with distance between positions
- Need for architecture with parallelization and global dependencies
Transformer: Key Innovation
- First sequence transduction model based entirely on attention
- Dispenses with recurrence and convolutions entirely
- Significantly more parallelizable than RNN/CNN models
- Achieves state-of-the-art results with lower training cost
Transformer Architecture (Figure 1)
Encoder (left) and decoder (right) stacks with self-attention and feed-forward layers.
Encoder Structure
- Stack of 6 identical layers
- Each layer: Multi-head self-attention + feed-forward network
- Residual connections around each sub-layer
- Layer normalization; output dimension dmodel=512
Decoder Structure
- Stack of 6 identical layers
- Three sub-layers: Masked self-attention, encoder-decoder attention, feed-forward
- Residual connections and layer normalization
- Masking prevents attending to future positions
Attention Mechanism
- Maps query (Q) and key-value (K,V) pairs to output
- Output: Weighted sum of values, weights from query-key compatibility
- Two primary types: Additive attention and dot-product attention
- Transformer uses scaled dot-product attention with multi-head extension
Scaled Dot-Product Attention
- Compute dot products of Q with all K
- Scale by 1/√dk to prevent softmax gradient vanishing
- Apply softmax to get weights on values
- Formula: Attention(Q,K,V) = softmax(QKT/√dk)V
Multi-Head Attention
- Project Q, K, V h times with learned linear projections
- Perform attention in parallel on each projection (heads)
- Concatenate results and project to final output
- h=8 heads, dk=dv=dmodel/h=64 (total cost similar to single-head)
Self-Attention Advantages
- Computational complexity: O(n²·d) vs O(n·d²) for RNNs
- Parallelization: O(1) sequential operations vs O(n) for RNNs
- Long-range dependencies: Constant path length vs O(n) for RNNs
- Interpretability: Attention distributions reveal dependency patterns
Positional Encoding
- Inject sequence order information (no recurrence/convolution)
- Added to input embeddings (same dimension dmodel=512)
- Uses sine/cosine functions with varying frequencies
- Supports learning of relative position relationships
Training Setup
- Datasets: WMT 2014 EN-DE (4.5M) and EN-FR (36M) sentence pairs
- Hardware: 8 NVIDIA P100 GPUs; base model trained 12h, big model 3.5 days
- Optimizer: Adam (β1=0.9, β2=0.98, ϵ=10⁻⁹) with linear warmup learning rate
- Regularization: Residual dropout (Pdrop=0.1), label smoothing (ϵls=0.1)
Translation Performance
- EN-DE: 28.4 BLEU (+2+ over previous SOTA including ensembles)
- EN-FR: 41.8 BLEU (new single-model state-of-the-art)
- Training cost: Significantly lower than competitors (e.g., 1/4 of prior SOTA)
- Base model outperforms most previous models at fraction of training time
Long-Distance Attention (Figure 3)
Encoder self-attention linking "making" to distant "more difficult" (layer 5).
Anaphora Attention (Figure 4)
Attention heads resolving "its" to "Law" and "application" (layer 5).
Generalization: Constituency Parsing
- 4-layer Transformer with dmodel=1024
- WSJ only (40K sentences): 91.3 F1 (comparable to SOTA)
- Semi-supervised (17M sentences): 92.7 F1 (outperforms most prior models)
- Demonstrates transferability to non-translation tasks
Limitations
- O(n²) complexity for sequence length n (challenging for very long sequences)
- Requires explicit positional encoding for sequence order
- Generation remains auto-regressive (sequential)
- Less explored for non-text modalities (images, audio, video)
Conclusion
- Transformer establishes new SOTA in machine translation
- Significantly faster training due to parallelization
- Self-attention effectively captures global dependencies
- Generalizes well to other sequence tasks beyond translation
Future Work
- Extend to multi-modal inputs/outputs (images, audio, video)
- Develop local restricted attention for long sequences
- Reduce sequential constraints in generation process
- Explore more interpretable attention patterns