Back to library

AI / Technology

An Image is Worth 16x16 WordsTransformers for Image Recognition at Scale

Three key questions about this paper

What problem does An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale address?

Authors: Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby Published: ICLR 2021; arXiv version 2, 3 June 2021 Primary source: arXiv:2010.11929 · Full paper

What evidence supports the main claim in An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale?

The headline JFT-pretrained ViT-H/14 reached 88.55% ImageNet accuracy, 90.72% on ImageNet-ReaL, 94.55% on CIFAR-100, and 77.63% average accuracy across VTAB. In the paper's comparison, it used 2,500 TPUv3-core-days of pretraining, versus 9,900 for the BiT-L convolutional baseline. These figures are averages over three fine-tuning runs where the paper reports them, and they combine architecture choices with training choices. [Paper: Table 2; §4.2]

What limitation should readers know about An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale?

The study pretrained models on ImageNet (1.3 million images), ImageNet-21k (14 million), and the private JFT dataset (303 million). It then evaluated transfer on ImageNet, ImageNet-ReaL, CIFAR-10/100, Oxford-IIIT Pets, Oxford Flowers-102, and VTAB's 19 tasks; each VTAB task used 1,000 training examples. [Paper: §4.1]

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro