Home Knowledge Base Company

Mistral is an efficient open-source language model family featuring innovations like sliding window attention. Company: Mistral AI (French startup, founded by ex-DeepMind/Meta researchers). Mistral 7B (Sept 2023): Outperformed LLaMA 2 13B despite being half the size. Best 7B model at release. Key innovations: Sliding window attention: Attend to only recent W tokens (4096), reducing memory, enabling long sequences. Grouped Query Attention: Efficient KV cache like LLaMA 2 70B. Rolling buffer cache: Fixed memory for KV cache regardless of sequence length. Architecture: 32 layers, 4096 hidden dim, 32 heads, 8 KV heads. Training: Undisclosed data and process, focused on quality and efficiency. License: Apache 2.0 (fully open, commercial OK). Mixtral 8x7B: Mixture of Experts version, 46.7B total but 12.9B active per token. Matches GPT-3.5 quality. Ecosystem: Widely adopted for fine-tuning, local deployment, and production use. Impact: Proved smaller, well-trained models can exceed larger ones. Efficiency-focused approach influential.

mistralfoundation model

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.