Flamingo is a visual language model (VLM) developed by DeepMind — enabling few-shot learning for vision tasks by fusing a frozen pre-trained vision encoder and a frozen large language model (LLM) with novel gated cross-attention layers.
What Is Flamingo?
- Definition: A family of VLM models (up to 80B parameters).
- Key Capability: In-context few-shot learning (e.g., show it 2 examples of a task, and it does the 3rd).
- Input: Interleaved images and text (e.g., a webpage with text and pictures).
- Output: Free-form text generation.
Why Flamingo Matters
- Frozen Components: Keeps the "smart" LLM (Chinchilla) and Vision (NFNet) weights frozen, training only connecting layers.
- Perceiver Resampler: Compresses variable visual features into a fixed number of tokens.
- Gated Cross-Attention: Inject visual information into the LLM without disrupting its text capabilities.
- Benchmark Smasher: Beat state-of-the-art fine-tuned models using only few-shot prompts.
Flamingo is the blueprint for modern VLMs — establishing the standard architecture (Frozen ViT + Projector + Frozen LLM) used by LLaVA, IDEFICS, and others.
flamingomultimodal ai
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.