---
theme: default
---
# Attention Is All You Need
## A Revolutionary Architecture for Sequence Transduction
Ashish Vaswani et al.
NIPS 2017
---
## Background
- Dominant sequence models rely on RNNs/LSTMs/GRUs
- Encoder-decoder architectures with auxiliary attention
- Sequential computation limits parallelization
- Long-range dependencies are challenging to model
---
## Motivation
- Recurrent networks: Inherently sequential, poor parallelization
- Convolutional networks: Limited receptive field, layered dependencies
- Both: Computation grows with distance between positions
- Need for architecture with parallelization and global dependencies
---
## Transformer: Key Innovation
- First sequence transduction model based entirely on attention
- Dispenses with recurrence and convolutions entirely
- Significantly more parallelizable than RNN/CNN models
- Achieves state-of-the-art results with lower training cost
---
## Transformer Architecture (Figure 1)
Encoder (left) and decoder (right) stacks with self-attention and feed-forward layers.
---
## Encoder Structure
- Stack of 6 identical layers
- Each layer: Multi-head self-attention + feed-forward network
- Residual connections around each sub-layer
- Layer normalization; output dimension dmodel=512
---
## Decoder Structure
- Stack of 6 identical layers
- Three sub-layers: Masked self-attention, encoder-decoder attention, feed-forward
- Residual connections and layer normalization
- Masking prevents attending to future positions
---
## Attention Mechanism
- Maps query (Q) and key-value (K,V) pairs to output
- Output: Weighted sum of values, weights from query-key compatibility
- Two primary types: Additive attention and dot-product attention
- Transformer uses scaled dot-product attention with multi-head extension
---
## Scaled Dot-Product Attention
- Compute dot products of Q with all K
- Scale by 1/√dk to prevent softmax gradient vanishing
- Apply softmax to get weights on values
- Formula: Attention(Q,K,V) = softmax(QKT/√dk)V
---
## Multi-Head Attention
- Project Q, K, V h times with learned linear projections
- Perform attention in parallel on each projection (heads)
- Concatenate results and project to final output
- h=8 heads, dk=dv=dmodel/h=64 (total cost similar to single-head)
---
## Self-Attention Advantages
- Computational complexity: O(n²·d) vs O(n·d²) for RNNs
- Parallelization: O(1) sequential operations vs O(n) for RNNs
- Long-range dependencies: Constant path length vs O(n) for RNNs
- Interpretability: Attention distributions reveal dependency patterns
---
## Positional Encoding
- Inject sequence order information (no recurrence/convolution)
- Added to input embeddings (same dimension dmodel=512)
- Uses sine/cosine functions with varying frequencies
- Supports learning of relative position relationships
---
## Training Setup
- Datasets: WMT 2014 EN-DE (4.5M) and EN-FR (36M) sentence pairs
- Hardware: 8 NVIDIA P100 GPUs; base model trained 12h, big model 3.5 days
- Optimizer: Adam (β1=0.9, β2=0.98, ϵ=10⁻⁹) with linear warmup learning rate
- Regularization: Residual dropout (Pdrop=0.1), label smoothing (ϵls=0.1)
---
## Translation Performance
- EN-DE: 28.4 BLEU (+2+ over previous SOTA including ensembles)
- EN-FR: 41.8 BLEU (new single-model state-of-the-art)
- Training cost: Significantly lower than competitors (e.g., 1/4 of prior SOTA)
- Base model outperforms most previous models at fraction of training time
---
## Long-Distance Attention (Figure 3)
Encoder self-attention linking "making" to distant "more difficult" (layer 5).
---
## Anaphora Attention (Figure 4)
Attention heads resolving "its" to "Law" and "application" (layer 5).
---
## Generalization: Constituency Parsing
- 4-layer Transformer with dmodel=1024
- WSJ only (40K sentences): 91.3 F1 (comparable to SOTA)
- Semi-supervised (17M sentences): 92.7 F1 (outperforms most prior models)
- Demonstrates transferability to non-translation tasks
---
## Limitations
- O(n²) complexity for sequence length n (challenging for very long sequences)
- Requires explicit positional encoding for sequence order
- Generation remains auto-regressive (sequential)
- Less explored for non-text modalities (images, audio, video)
---
## Conclusion
- Transformer establishes new SOTA in machine translation
- Significantly faster training due to parallelization
- Self-attention effectively captures global dependencies
- Generalizes well to other sequence tasks beyond translation
---
## Future Work
- Extend to multi-modal inputs/outputs (images, audio, video)
- Develop local restricted attention for long sequences
- Reduce sequential constraints in generation process
- Explore more interpretable attention patterns