Back to library

AI / Technology

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Three key questions about this paper

What problem does Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks address?

Patrick Lewis et al. · NeurIPS 2020 · arXiv:2005.11401 · version 4, 12 April 2021

What evidence supports the main claim in Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks?

For Jeopardy question generation, human evaluators judged RAG-Token more factual than BART in 42.7% of 452 comparisons, while judging BART more factual in 7.1%; the remaining cases were ties, both poor, or had no majority. RAG was also judged more specific in 37.4% versus 16.8% for BART. These are pairwise judgments on one task, not a universal factuality rate. Paper, §4.3 and Table 4

What limitation should readers know about Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks?

The authors queried 82 world-leader positions using either a December 2016 or December 2018 Wikipedia index. With the matching index, accuracy was 70% for 2016 leaders and 68% for 2018 leaders. With mismatched index and target year, it fell to 12% and 4%. This is direct evidence that replacing external memory can change time-specific answers without retraining the model; it does not prove every updated fact will be followed correctly. Paper, §4.5

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro