Back to library

AI / Technology

Training Compute-Optimal Large Language Models

Three key questions about this paper

What problem does Training Compute-Optimal Large Language Models address?

Jordan Hoffmann and colleagues · arXiv:2203.15556 · submitted 29 March 2022

What evidence supports the main claim in Training Compute-Optimal Large Language Models?

Chinchilla reached 67.6% five-shot accuracy on MMLU, versus 60.0% for Gopher, and performed better on 51 of the 57 MMLU tasks, tied on two, and was worse on four. On the 62 reported BIG-bench tasks, its mean accuracy was 65.1% versus 54.4%, with worse performance on four tasks. It also reduced WikiText-103 perplexity from 7.75 to 7.16 and improved over Gopher across every reported Pile subset.

What limitation should readers know about Training Compute-Optimal Large Language Models?

The authors had only two comparable runs at the largest scale—Gopher and Chinchilla—and no intermediate large-scale checks. Their projection assumes a power-law frontier, yet they observed curvature at high compute, which may mean even smaller models are optimal there. The scaling experiments used less than one epoch of data, so repeated-data training was left for future work.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro