Latent diffusion models run the diffusion process in compressed latent space for efficiency, as used in Stable Diffusion. Motivation: Running diffusion in pixel space is computationally expensive (high-dimensional). Compress to latent space first. Architecture: VAE encoder compresses images to latent representation, diffusion U-Net operates in latent space, VAE decoder reconstructs image from generated latents. Efficiency gains: 4-8× spatial compression (256×256 image → 32×32 latents), dramatically faster training and inference, lower memory requirements. Training stages: Train VAE (encoder-decoder) separately, train diffusion model on encoded latents. Components: VAE with KL regularization, U-Net with cross-attention for conditioning, CLIP text encoder for text-to-image. Stable Diffusion specifics: Trained by Stability AI, open-source weights, 4× latent compression, efficient enough for consumer GPUs. Advantages: Faster iteration in research, accessible to broader community, enables real-time applications. Trade-offs: VAE reconstruction can lose details, two-stage training complexity. Impact: Democratized high-quality image generation, foundation for most current open-source image generation.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.