--- theme: default --- # Attention Is All You Need ## A Revolutionary Architecture for Sequence Transduction Ashish Vaswani et al. NIPS 2017 --- ## Problem Statement - Traditional sequence models rely on RNNs or CNNs - RNNs have inherent sequential computation (hard to parallelize) - CNNs require multiple layers to capture long-range dependencies - Both have increasing path lengths between distant positions **Key Insight**: Replace recurrence and convolution with attention mechanisms --- ## The Transformer: Model Architecture - Encoder-decoder structure with stacked self-attention layers - No recurrence or convolution - Significantly more parallelizable --- ## Encoder Stack - 6 identical layers - Each layer has two sub-layers: 1. Multi-head self-attention mechanism 2. Position-wise fully connected feed-forward network - Residual connections + layer normalization around each sub-layer - Output dimension: dmodel = 512 --- ## Decoder Stack - 6 identical layers - Three sub-layers per layer: 1. Masked multi-head self-attention 2. Multi-head attention over encoder output 3. Position-wise fully connected feed-forward network - Residual connections + layer normalization - Masking prevents attending to future positions --- ## Attention Mechanism ### Scaled Dot-Product Attention $$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V$$ - Queries (Q), Keys (K), Values (V) are matrices - Scaling by $\sqrt{d_k}$ prevents gradient vanishing - More efficient than additive attention --- ## Multi-Head Attention - Project Q, K, V h times with different linear projections - Perform attention in parallel on each projection - Concatenate results and project again $$\text{MultiHead}(Q, K, V) = \text{Concat}(\text{head}_1, ..., \text{head}_h)W^O$$ --- ## Applications of Attention 1. **Encoder-decoder attention**: - Queries from decoder, keys/values from encoder - Allows decoder to attend to all input positions 2. **Encoder self-attention**: - All Q, K, V from previous encoder layer - Each position attends to all positions 3. **Decoder self-attention**: - All Q, K, V from previous decoder layer - Masked to prevent attending to future positions --- ## Positional Encoding Since model has no recurrence/convolution, we add positional information: $$PE_{(pos, 2i)} = \sin\left(pos / 10000^{2i/d_{\text{model}}}\right)$$ $$PE_{(pos, 2i+1)} = \cos\left(pos / 10000^{2i/d_{\text{model}}}\right)$$ - Sinusoidal functions with varying frequencies - Allows model to learn relative position information - Performed as well as learned positional embeddings --- ## Why Self-Attention? | Layer Type | Complexity | Sequential Ops | Max Path Length | |------------|------------|----------------|-----------------| | Self-Attention | O(n²·d) | O(1) | O(1) | | Recurrent | O(n·d²) | O(n) | O(n) | | Convolutional | O(k·n·d²) | O(1) | O(logₖn) | - Better parallelization than RNNs - Shorter path lengths than CNNs/RNNs - More interpretable attention patterns --- ## Training Details - **Datasets**: WMT 2014 EN-DE (4.5M) and EN-FR (36M) - **Hardware**: 8 NVIDIA P100 GPUs - **Optimizer**: Adam (β₁=0.9, β₂=0.98, ϵ=10⁻⁹) - **Learning Rate**: $lrate = d_{\text{model}}^{-0.5} \cdot \min(\text{step_num}^{-0.5}, \text{step_num} \cdot \text{warmup_steps}^{-1.5})$ - **Regularization**: Dropout (P=0.1), Label Smoothing (ϵ=0.1) --- ## Machine Translation Results | Model | EN-DE BLEU | EN-FR BLEU | Training Cost (FLOPs) | |-------|------------|------------|------------------------| | GNMT + RL | 24.6 | 39.92 | 2.3·10¹⁹ | | ConvS2S | 25.16 | 40.46 | 9.6·10¹⁸ | | Transformer (base) | 27.3 | 38.1 | 3.3·10¹⁸ | | **Transformer (big)** | **28.4** | **41.8** | **2.3·10¹⁹** | - Transformer (big) outperforms all previous state-of-the-art - Achieves new SOTA with significantly less training cost --- ## Model Variations | Configuration | Dev PPL | Dev BLEU | |---------------|---------|----------| | Base | 4.92 | 25.8 | | 1 attention head | 5.29 | 24.9 | | 32 attention heads | 5.01 | 25.4 | | No dropout | 5.77 | 24.6 | | Learned positional embeddings | 4.92 | 25.7 | | Big model | 4.33 | 26.4 | - Multiple attention heads improve performance - Dropout is crucial for generalization - Sinusoidal and learned positional encodings perform similarly --- ## English Constituency Parsing | Parser | Training | WSJ 23 F1 | |--------|----------|-----------| | Petrov et al. (2006) | WSJ only | 90.4 | | Dyer et al. (2016) | WSJ only | 91.7 | | **Transformer (4 layers)** | **WSJ only** | **91.3** | | Vinyals & Kaiser et al. | Semi-supervised | 92.1 | | **Transformer (4 layers)** | **Semi-supervised** | **92.7** | - Transformer generalizes well to other sequence tasks - Achieves competitive results with minimal task-specific modifications --- ## Attention Visualization: Long-Distance Dependencies - Attention heads follow distant dependencies - Example: Verb "making" attends to "more difficult" - Different colors represent different attention heads --- ## Attention Visualization: Anaphora Resolution - Some heads specialize in resolving pronouns - Example: "its" clearly attends to "The Law" - Sharp attention patterns for coreference resolution --- ## Limitations and Future Work ### Limitations: - High memory consumption with long sequences (O(n²)) - Still generates output sequentially ### Future Work: - Local, restricted attention mechanisms for large inputs - Extend to other modalities (images, audio, video) - Make generation less sequential - Explore more efficient attention variants --- ## Conclusion - Transformer is the first transduction model based entirely on attention - Replaces RNNs/CNNs with multi-headed self-attention - Significantly faster training with better parallelization - Achieves new state-of-the-art in machine translation - Generalizes well to other sequence tasks like constituency parsing **Code**: https://github.com/tensorflow/tensor2tensor --- ## Thank You!