YANK!
PDFSEP 27, 2026

Attention Is All You Need: The Transformer Paper's Five Consequential Claims

The 2017 Google Brain paper that scrapped recurrence and convolutions entirely, proposing the Transformer architecture that now underlies nearly every large language model.

The five all come from pages 2 to 8 of 15.

Create from thisView source

Yank’s 5 are Yank’s picks from this source. Want something else?

Choose your own
May find some of the same lines.
1 of 5

We propose the Transformer, the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

Speaker not establishedPAGE 2
Why Yank kept it

Positions the Transformer as a genuine architectural departure, not merely an incremental improvement on RNN or convolutional sequence models.

2 of 5

The Transformer (big) achieves 28.4 BLEU on WMT 2014 English-to-German, outperforming all previously reported models including ensembles by more than 2 BLEU, trained in 3.5 days on 8 P100 GPUs.

Speaker not establishedPAGE 8
Why Yank kept it

Anchors the paper's empirical case with a concrete state-of-the-art margin and a surprisingly modest training cost.

3 of 5

A self-attention layer connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires O(n) sequential operations.

Speaker not establishedPAGE 6
Why Yank kept it

Explains the architectural reason self-attention is better suited than recurrence for learning long-range dependencies.

4 of 5

Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions. With a single attention head, averaging inhibits this.

Speaker not establishedPAGE 5
Why Yank kept it

Justifies the multi-head design choice by showing what a single averaged head structurally cannot do.

5 of 5

Label smoothing hurts perplexity, as the model learns to be more unsure, but improves accuracy and BLEU score.

Speaker not establishedPAGE 8
Why Yank kept it

Resolves an apparent contradiction, showing a training trick that hurts one metric while genuinely improving the ones that matter for translation.

Didn’t see the line you wanted? Choose your own

Ask this source

Ask about anything in this document. Yank answers only from what the source says, shows you where, and tells you when the source doesn’t say.

Answers are written by Yank from the source, for you. Nothing you ask is saved or published.