Back to library

AI / Technology

Attention Is All You Need

Three key questions about this paper

What problem does Attention Is All You Need address?

Authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin Published: arXiv:1706.03762, first submitted 12 June 2017; source version v7 dated 2 August 2023 Primary source: arXiv abstract · Full paper

What evidence supports the main claim in Attention Is All You Need?

The strongest evidence is machine translation on WMT 2014. The “big” Transformer reached 28.4 BLEU on English→German, versus 26.36 for the best prior ensemble listed in Table 2, and 41.8 BLEU on English→French, versus 41.29 for the best listed ensemble. The big models trained for 300,000 steps over 3.5 days on eight NVIDIA P100 GPUs. The smaller base models trained for 100,000 steps in about 12 hours on the same number of GPUs. [Source: Abstract; §§5.2, 6.1; Table 2]

What limitation should readers know about Attention Is All You Need?

These results come from two translation benchmarks and one parsing task, with 2017-era tokenization, hardware, baselines, and tuning. BLEU is an automatic overlap metric, not a direct measure of factuality, safety, or human preference. The paper does not test very long sequences, memory use at modern context lengths, multilingual coverage beyond the reported pairs, robustness, or deployment latency.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro