diffusion

**Diffusion models** generate data by **learning to reverse a gradual noising process** — progressively adding Gaussian noise to data during training, then learning to denoise step-by-step during generation, producing high-quality images, audio, and video that rival or exceed GANs. **What Are Diffusion Models?** - **Definition**: Generative models based on denoising process. - **Training**: Learn to reverse gradual corruption by noise. - **Generation**: Start from pure noise, iteratively denoise. - **Examples**: Stable Diffusion, DALL-E, Midjourney, Sora. **Why Diffusion Works** - **Stable Training**: No adversarial dynamics (unlike GANs). - **Quality**: State-of-the-art image generation. - **Flexibility**: Conditional generation, inpainting, editing. - **Theory**: Strong mathematical foundation. **Forward Process (Noising)** **Gradual Corruption**: ``` x_0 → x_1 → x_2 → ... → x_T (data) (pure noise) At each step: x_t = √(α_t) × x_{t-1} + √(1-α_t) × ε Where ε ~ N(0, I) is Gaussian noise α_t follows a schedule (typically 0.9999 to 0.0001) ``` **Closed Form to Any Step**: ``` x_t = √(ᾱ_t) × x_0 + √(1-ᾱ_t) × ε Where ᾱ_t = Π_{s=1}^t α_s (cumulative product) ``` **Visual**: ```svg t=0 t=250 t=500 t=750 t=1000┌────┐ ┌────┐ ┌────┐ ┌────┐ ┌────┐🐱 🐱+ε ≈≈≈ ░░░ ▓▓▓clear noise noisy v.noisy noise└────┘ └────┘ └────┘ └────┘ └────┘ ``` **Reverse Process (Denoising)** **Learning to Denoise**: ``` Train neural network ε_θ to predict noise: Loss = ||ε - ε_θ(x_t, t)||² Given noisy image x_t and timestep t, predict the noise ε that was added. ``` **Generation** (Sampling): ``` Start: x_T ~ N(0, I) (pure noise) For t = T, T-1, ..., 1: Predict noise: ε̂ = ε_θ(x_t, t) Compute x_{t-1} using ε̂ Return: x_0 (generated sample) ``` **Implementation Sketch** **Training Loop**: ```python import torch import torch.nn.functional as F def train_step(model, x_0, noise_scheduler): # Sample random timesteps t = torch.randint(0, T, (batch_size,)) # Sample noise noise = torch.randn_like(x_0) # Add noise to get x_t x_t = noise_scheduler.add_noise(x_0, noise, t) # Predict noise predicted_noise = model(x_t, t) # MSE loss loss = F.mse_loss(predicted_noise, noise) return loss ``` **Sampling Loop**: ```python @torch.no_grad() def sample(model, noise_scheduler, shape): # Start from pure noise x = torch.randn(shape) # Iteratively denoise for t in reversed(range(T)): # Predict noise predicted_noise = model(x, t) # Compute previous step x = noise_scheduler.step(predicted_noise, t, x) return x ``` **Key Architectures** **U-Net (Standard)**: ```svg ┌─────────────────────────────────────────────────────────┐ Noisy Image + t └─────────────────────────────────────────────────────────┘ ┌───▼───┐ Encoder (downsampling) Conv └───┬───┘ │──────────────────┐ ┌───▼───┐ Conv └───┬───┘ │─────────────┐ ┌───▼───┐ Bottom └───┬───┘ ┌───▼───┐←────────┘ Decoder (upsampling) Conv skip └───┬───┘ ┌───▼───┐←─────────────┘ Conv skip └───┬───┘ Predicted Noise ``` **DiT (Diffusion Transformer)**: ``` Modern alternative using transformers instead of U-Net: - Used in Sora, recent SOTA models - Better scaling properties - Patch-based processing ``` **Conditional Generation** **Text-to-Image**: ```python # Classifier-free guidance def guided_sample(model, prompt, guidance_scale=7.5): text_embeddings = encode_text(prompt) for t in reversed(range(T)): # Conditional prediction noise_cond = model(x, t, text_embeddings) # Unconditional prediction noise_uncond = model(x, t, null_embedding) # Guided prediction noise = noise_uncond + guidance_scale * (noise_cond - noise_uncond) x = denoise_step(x, noise, t) return x ``` **Popular Models** ``` Model | Type | Open Source -------------------|----------------|------------ Stable Diffusion | Text-to-image | Yes DALL-E 3 | Text-to-image | No Midjourney | Text-to-image | No Sora | Text-to-video | No Runway Gen-2 | Text-to-video | No AudioLDM | Text-to-audio | Yes ``` Diffusion models are **the dominant paradigm for generative AI** — their stable training, high quality outputs, and flexibility for conditioning have made them the foundation of modern image, video, and audio generation systems.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account