Back to library

AI / Technology

Training language models to follow instructions with human feedback

Three key questions about this paper

What problem does Training language models to follow instructions with human feedback address?

Long Ouyang et al. · OpenAI · submitted 4 March 2022 · arXiv:2203.02155 · PDF

What evidence supports the main claim in Training language models to follow instructions with human feedback?

The paper also reports that on closed-domain tasks, hallucinations fell from 41% for GPT-3 to 21% for InstructGPT. On TruthfulQA, truthful and informative answers appeared about twice as often. With a respectful prompt, InstructGPT produced about 25% fewer toxic outputs than GPT-3. Bias results on Winogender and CrowS-Pairs did not improve significantly. Source: Sections 1 and 4.2, pp. 3 and 13–15.

What limitation should readers know about Training language models to follow instructions with human feedback?

Visual 3. The learned preference signal passes through a narrow human sample. The paper's own limitation is representativeness, not a failure to use human judgment at all.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro