Home Knowledge Base LLaVA

LLaVA (Large Language and Vision Assistant) is an open-source multimodal model — that combines a vision encoder (CLIP ViT-L) with an LLM (Vicuna/LLaMA) to creating a "visual chatbot" with capabilities similar to GPT-4 Vision.

What Is LLaVA?

Why LLaVA Matters

Training Stages

1. Feature Alignment: Pre-training to align image features to word embeddings. 2. Visual Instruction Tuning: Fine-tuning on the GPT-4 generated instruction data (conversations, reasoning).

LLaVA is the "Hello World" of modern VLMs — its simple, effective recipe became the standard basline for nearly all subsequent open-source multimodal research.

llava (large language and vision assistant)llavalarge language and vision assistantmultimodal ai

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.