YANK!
PDFSEP 28, 2026

BERT: Bidirectional Pre-training Breaks the NLP Leaderboard

The landmark Google AI paper that introduced BERT, demonstrating how masking tokens and conditioning on both left and right context simultaneously produced a single pre-trained model that shattered eleven NLP benchmarks at once.

View original

AI summary

Generated by Yank from this source. A summary, not quoted language.

BERT (Bidirectional Encoder Representations from Transformers) is a new language representation model introduced by researchers at Google AI Language. Unlike prior approaches such as OpenAI GPT, which processes text left-to-right, BERT is pre-trained to condition on both left and right context simultaneously in all layers, making it deeply bidirectional. It achieves this through two unsupervised pre-training tasks: a Masked Language Model (MLM), which randomly masks 15% of input tokens and trains the model to predict them from surrounding context, and a Next Sentence Prediction (NSP) task, which trains the model to understand relationships between sentence pairs.

BERT comes in two sizes: BERT BASE (12 layers, 768 hidden size, 110M parameters) and BERT LARGE (24 layers, 1024 hidden size, 340M parameters). It is pre-trained on the BooksCorpus (800M words) and English Wikipedia (2,500M words). Once pre-trained, the model can be fine-tuned for a wide range of downstream tasks by adding just one additional output layer, with all parameters updated end-to-end. Fine-tuning is described as relatively inexpensive, with results replicable in at most one hour on a single Cloud TPU.

Yank summary

What matters, in Yank’s words, each linked to where the source says it.

  1. BERT obtains new state-of-the-art results on eleven NLP tasks, pushing the GLUE score to 80.5% (a 7.7-point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (a 5.1-point improvement). p. 1

  2. BERT uses two pre-training tasks: a Masked LM that randomly masks 15% of tokens for prediction, and a Next Sentence Prediction task trained on 50% real and 50% random sentence pairs. p. 4

  3. BERT is pre-trained on the BooksCorpus (800M words) and English Wikipedia (2,500M words), using a document-level corpus to extract long contiguous sequences. p. 5

  4. Ablation studies show that removing the NSP task significantly hurts performance on QNLI, MNLI, and SQuAD 1.1, confirming its importance. p. 8

  5. For the feature-based approach on NER, concatenating the last four hidden layers of BERT as features yields a Dev F1 of 96.1, only 0.3 F1 behind full fine-tuning. p. 9

The five worth keeping

May find some of the same lines.
01 / 05

BERT is the first fine-tuning based representation model that achieves state-of-the-art performance on a large suite of sentence-level and token-level tasks, outperforming many task-specific architectures.

PAGE 2
In context

Unlike Radford et al. (2018), which uses unidirectional language models for pre-training, BERT uses masked language models to enable pre-trained deep bidirectional representations. This is also in contrast to Peters et al. (2018a), which uses a shallow concatenation of independently trained left-to-right and right-to-left LMs. • We show that pre-trained representations reduce the need for many heavily-engineered task-specific architectures. BERT is the first fine-tuning based representation model that achieves state-of-the-art performance on a large suite of sentence-level and token-level tasks, outper-forming many task-specific architectures. • BERT advances the state of the art for eleven…

What was said, against the card

BERT is the first fine-tuning based representation model that achieves state-of-the-art performance on a large suite of sentence-level and token-level tasks, outperforming many task-specific architectures. • BERT

Distilled · paraphrased from this passageRead in original ↗
Receipt
Source
PDF · arxiv.org
Read from
Document text
Located
Page 2
Speaker
Not established
Checked
Sep 28, 2026
02 / 05

Bidirectional conditioning would allow each word to indirectly 'see itself', and the model could trivially predict the target word in a multi-layered context — so BERT instead masks 15% of tokens at random and predicts only those.

PAGE 4
In context

Instead, we pre-train BERT using two unsupervised tasks, described in this section. This step is presented in the left part of Figure 1. Task #1: Masked LM Intuitively, it is reason-able to believe that a deep bidirectional model is strictly more powerful than either a left-to-right model or the shallow concatenation of a left-to-right and a right-to-left model. Unfortunately, standard conditional language models can only be trained left-to-right or right-to-left, since bidirectional conditioning would allow each word to in-directly “see itself”, and the model could trivially predict the target word in a multi-layered context. former is often referred to as a “Transformer encoder” while the…

What was said, against the card

only be trained left-to-right or right-to-left, since Bidirectional conditioning would allow each word to indirectly 'see itself', and the model could trivially predict the target word in a multi-layered context — so BERT instead masks 15% of tokens at random and predicts only those.

Distilled · paraphrased from this passageRead in original ↗
Receipt
Source
PDF · arxiv.org
Read from
Document text
Located
Page 4
Speaker
Not established
Checked
Sep 28, 2026
03 / 05

This is the first work to demonstrate convincingly that scaling to extreme model sizes also leads to large improvements on very small scale tasks, provided that the model has been sufficiently pre-trained.

PAGE 8
In context

…. For example, the largest Transformer explored in Vaswani et al. (2017) is (L=6, H=1024, A=16) with 100M parameters for the encoder, and the largest Transformer we have found in the literature is (L=64, H=512, A=2) with 235M parameters (Al-Rfou et al., 2018). By contrast, BERT BASE contains 110M parameters and BERT LARGE contains 340M parameters. It has long been known that increasing the model size will lead to continual improvements on large-scale tasks such as machine translation and language modeling, which is demonstrated by the LM perplexity of held-out training data shown in Table 6. However, we believe that this is the first work to demonstrate convinc-ingly that scaling to extreme…

What was said, against the card

LARGE contains 340M parameters. It has long been known that increasing the model size will lead to continual improvements on large-scale tasks such as machine translation and language modeling, which is demonstrated by the LM perplexity of held-out training data shown in Table 6. However, we believe that This is the first work to demonstrate convincingly that scaling to extreme model sizes also leads to large improvements on very small scale tasks, provided that the model has been sufficiently pre-trained.

Distilled · paraphrased from this passageRead in original ↗
Receipt
Source
PDF · arxiv.org
Read from
Document text
Located
Page 8
Speaker
Not established
Checked
Sep 28, 2026
04 / 05

BERT LARGE obtains a GLUE score of 80.5 versus OpenAI GPT's 72.8 — an architecture otherwise nearly identical except for attention direction.

PAGE 6
In context

…we use the same pre-trained checkpoint but perform different fine-tuning data shuffling and classifier layer initialization. 9 Results are presented in Table 1. Both BERT BASE and BERT LARGE outperform all systems on all tasks by a substantial margin, obtaining 4.5% and 7.0% respective average accuracy improvement over the prior state of the art. Note that BERT BASE and OpenAI GPT are nearly identical in terms of model architecture apart from the attention masking. For the largest and most widely reported GLUE task, MNLI, BERT obtains a 4.6% absolute accuracy improvement. On the official GLUE leaderboard 10 , BERT LARGE obtains a score of 80.5, compared to OpenAI GPT, which obtains 72.8 as…

What was said, against the card

OpenAI GPT are nearly identical in terms of model architecture apart from the attention masking. For the largest and most widely reported GLUE task, MNLI, BERT LARGE obtains a 4.6% absolute accuracy improvement. On the official GLUE leaderboard 10 score of 80.5 versus OpenAI GPT's 72.8 — BERT LARGE an architecture otherwise nearly identical except for attention direction.

Distilled · paraphrased from this passageRead in original ↗
Receipt
Source
PDF · arxiv.org
Read from
Document text
Located
Page 6
Speaker
Not established
Checked
Sep 28, 2026
05 / 05

The Transformer encoder does not know which words it will be asked to predict or which have been replaced by random words, so it is forced to keep a distributional contextual representation of every input token.

PAGE 12
In context

…ing to hairy ), our masking procedure can be further illustrated by • 80% of the time: Replace the word with the [MASK] token, e.g., my dog is hairy → my dog is [MASK] • 10% of the time: Replace the word with a random word, e.g., my dog is hairy → my dog is apple • 10% of the time: Keep the word unchanged, e.g., my dog is hairy → my dog is hairy . The purpose of this is to bias the representation towards the actual observed word. The advantage of this procedure is that the Transformer encoder does not know which words it will be asked to predict or which have been re-placed by random words, so it is forced to keep a distributional contextual representation of every input token. Additionally,…

What was said, against the card

The Transformer encoder does not know which words it will be asked to predict or which have been replaced by random words, so it is forced to keep a distributional contextual representation of every input token.

Distilled · paraphrased from this passageRead in original ↗
Receipt
Source
PDF · arxiv.org
Read from
Document text
Located
Page 12
Speaker
Not established
Checked
Sep 28, 2026