flamingo

**Flamingo** is a **visual language model (VLM) developed by DeepMind** — enabling few-shot learning for vision tasks by fusing a frozen pre-trained vision encoder and a frozen large language model (LLM) with novel gated cross-attention layers. **What Is Flamingo?** - **Definition**: A family of VLM models (up to 80B parameters). - **Key Capability**: In-context few-shot learning (e.g., show it 2 examples of a task, and it does the 3rd). - **Input**: Interleaved images and text (e.g., a webpage with text and pictures). - **Output**: Free-form text generation. **Why Flamingo Matters** - **Frozen Components**: Keeps the "smart" LLM (Chinchilla) and Vision (NFNet) weights frozen, training only connecting layers. - **Perceiver Resampler**: Compresses variable visual features into a fixed number of tokens. - **Gated Cross-Attention**: Inject visual information into the LLM without disrupting its text capabilities. - **Benchmark Smasher**: Beat state-of-the-art fine-tuned models using only few-shot prompts. **Flamingo** is **the blueprint for modern VLMs** — establishing the standard architecture (Frozen ViT + Projector + Frozen LLM) used by LLaVA, IDEFICS, and others.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account