Back to library

AI / Technology

Learning Transferable Visual Models From Natural Language Supervision

Three key questions about this paper

What problem does Learning Transferable Visual Models From Natural Language Supervision address?

Alec Radford et al. · OpenAI · arXiv:2103.00020 · submitted 26 February 2021 Primary source: arXiv record · paper PDF

What evidence supports the main claim in Learning Transferable Visual Models From Natural Language Supervision?

These comparisons show that language-supervised pre-training can produce a reusable task interface. They do not show that CLIP beats the best specialist on every dataset: the main zero-shot baseline was a linear classifier on ResNet-50 features, which the authors note was already below the state of the art on most datasets. Source: §6, p. 19.

What limitation should readers know about Learning Transferable Visual Models From Natural Language Supervision?

CLIP performed poorly on satellite classification, lymph-node tumor detection, synthetic counting, traffic-sign recognition, and estimating distance to the nearest car. The paper says performance on some novel tasks can be near random. It also gives an instructive distribution-shift failure: zero-shot CLIP reached 88% on handwritten MNIST digits, yet a logistic regression model on raw pixels did better.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro