02 · KEY IDEA

The Transformer: Attention Is All You Need

No Recurrence

Sequential RNN hidden states are removed entirely — positions are processed in parallel.

No Convolution

No fixed-width kernels; every position can connect to every other position directly.

Self-Attention Only

Multi-headed self-attention computes all representations, in both encoder and decoder.

Source: Vaswani et al., Abstract and §1 (paper p. 1–2).