Back to library

AI / Technology

Scaling Laws for Neural Language Models

Three key questions about this paper

What problem does Scaling Laws for Neural Language Models address?

Jared Kaplan and colleagues, 2020 · arXiv:2001.08361

What evidence supports the main claim in Scaling Laws for Neural Language Models?

The strongest result is consistency across a wide sweep of experiments: the combined fits covered eight orders of magnitude in optimally allocated compute, six in model size, and more than two in dataset size. The overfitting fit says data should grow approximately as N^0.74; the authors translate that into about 5× more data for an 8× larger model.

What limitation should readers know about Scaling Laws for Neural Language Models?

Visual 3. The arrow stops at the evidence boundary: the experiments support claims about measured loss, while capability and cross-domain claims remain questions.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro