Back to library

AI / Technology

Visual Instruction Tuning

Three key questions about this paper

What problem does Visual Instruction Tuning address?

Authors: Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee Published: NeurIPS 2023 (Oral); arXiv v2, 11 December 2023 Primary source: Paper abstract · Full paper (PDF) Reading note: Results and limitations below refer to this 2023 paper and its LLaVA setup, not later LLaVA versions.

What evidence supports the main claim in Visual Instruction Tuning?

The strongest causal evidence is the data ablation on LLaVA-Bench (COCO). On 90 generated questions over 30 COCO images, the model without instruction tuning scored 21.5 relative to the GPT-4 reference; the full instruction dataset scored 85.1. Conversation-only training reached 73.8, while adding description and reasoning examples improved the overall result.

What limitation should readers know about Visual Instruction Tuning?

The authors explicitly warn about hallucinated details, inherited bias from CLIP and Vicuna/LLaMA, harmful inputs, energy costs at larger scale, and the difficulty of evaluating fine-grained visual grounding. They describe the work as an initial step focused mainly on real-life tasks, not proof of safe general-purpose visual understanding. Paper, Broader Impact Paper, Conclusion

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro