Back to library

AI / Technology

High-Resolution Image Synthesis with Latent Diffusion Models

Three key questions about this paper

What problem does High-Resolution Image Synthesis with Latent Diffusion Models address?

Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · CVPR 2022 · arXiv:2112.10752 · version 2, 13 April 2022

What evidence supports the main claim in High-Resolution Image Synthesis with Latent Diffusion Models?

On class-conditional ImageNet at 256×256, the reported best guided LDM reaches FID 3.60, compared with 4.59 for the cited guided ADM baseline, while the paper reports substantially lower training compute for the LDM. On 256×256 image inpainting, the fine-tuned large LDM reports FID 1.50, which the authors identify as a new state of the art in their comparison. These are task-specific benchmark results, not proof that every latent model is cheaper or better in every setting.

What limitation should readers know about High-Resolution Image Synthesis with Latent Diffusion Models?

On class-conditional ImageNet at 256×256, the reported best guided LDM reaches FID 3.60, compared with 4.59 for the cited guided ADM baseline, while the paper reports substantially lower training compute for the LDM. On 256×256 image inpainting, the fine-tuned large LDM reports FID 1.50, which the authors identify as a new state of the art in their comparison. These are task-specific benchmark results, not proof that every latent model is cheaper or better in every setting.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro