Back to library

AI / Technology

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Three key questions about this paper

What problem does Chain-of-Thought Prompting Elicits Reasoning in Large Language Models address?

Larger language models had improved many language tasks, yet standard prompting still struggled on arithmetic word problems, commonsense questions, and symbolic manipulation. A standard few-shot prompt demonstrates only input–output pairs. That format tells the model what answer shape is wanted, but gives no example of how to break a multi-step problem apart. Source: Introduction, paragraphs 1–6

What evidence supports the main claim in Chain-of-Thought Prompting Elicits Reasoning in Large Language Models?

On GSM8K, PaLM 540B rose from 17.9% with standard prompting to 56.9% with chain-of-thought prompting, a gain of 39.0 percentage points. The same model improved from 69.4% to 79.0% on SVAMP and from 79.2% to 93.3% on MAWPS. But the effect was uneven: on ASDiv the increase was only 72.1% to 73.9%. Source: Appendix Table 2

What limitation should readers know about Chain-of-Thought Prompting Elicits Reasoning in Large Language Models?

On GSM8K, PaLM 540B rose from 17.9% with standard prompting to 56.9% with chain-of-thought prompting, a gain of 39.0 percentage points. The same model improved from 69.4% to 79.0% on SVAMP and from 79.2% to 93.3% on MAWPS. But the effect was uneven: on ASDiv the increase was only 72.1% to 73.9%. Source: Appendix Table 2

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro