Home Knowledge Base LLaVA (Large Language-and-Vision Assistant)

LLaVA (Large Language-and-Vision Assistant) is the pioneering open-source vision-language model that introduced visual instruction tuning — connecting a CLIP vision encoder to a LLaMA/Vicuna language model and training on GPT-4-generated visual conversation data to create a multimodal assistant that can describe images, answer visual questions, reason about visual content, and follow complex instructions involving both text and images.

What Is LLaVA?

LLaVA Model Versions

VersionVision EncoderLLMProjectionTraining DataKey Improvement
LLaVA 1.0CLIP ViT-L/14Vicuna-13BLinear158KFirst visual instruction tuning
LLaVA 1.5CLIP ViT-L/14@336Vicuna-7B/13B2-layer MLP665KBetter projection, higher res
LLaVA 1.6 (NeXT)CLIP ViT-L/14@672Mistral-7B/Vicuna-13BMLP1M+Dynamic high resolution
LLaVA-OneVisionSigLIPQwen2-7B/72BMLP3M+Video understanding

Why LLaVA Matters

LLaVA is the open-source vision-language model that established visual instruction tuning as the standard approach for building multimodal AI assistants — demonstrating that connecting a CLIP vision encoder to an LLM through a simple projection layer, trained on GPT-4-generated visual conversation data, produces powerful multimodal capabilities that rival proprietary systems.

llavavisual instructiontuning

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.