02 · KEY IDEA
No Recurrence
Sequential RNN hidden states are removed entirely — positions are processed in parallel.
No Convolution
No fixed-width kernels; every position can connect to every other position directly.
Self-Attention Only
Multi-headed self-attention computes all representations, in both encoder and decoder.
Source: Vaswani et al., Abstract and §1 (paper p. 1–2).