NEURIPS 2017 · ARXIV:1706.03762
The Transformer: sequence transduction based entirely on attention — no recurrence, no convolution.
Ashish Vaswani · Noam Shazeer · Niki Parmar · Jakob UszkoreitLlion Jones · Aidan N. Gomez · Łukasz Kaiser · Illia Polosukhin
Google Brain · Google Research · University of Toronto
The Transformer — model architectureFigure 1, paper p. 3