language model pretraining

**Language Model Pretraining** is the **foundational training phase where a large neural network (transformer) learns general language understanding and generation capabilities from vast text corpora (hundreds of billions to trillions of tokens) — using self-supervised objectives (masked language modeling for BERT-style models, next-token prediction for GPT-style models) that capture grammar, facts, reasoning patterns, and world knowledge in the model's parameters, creating a versatile foundation that is then adapted to specific tasks through fine-tuning or prompting**. **Pretraining Objectives** **Causal Language Modeling (CLM) — GPT-style**: - Predict the next token given all previous tokens: P(x_t | x_1, ..., x_{t-1}). - Unidirectional attention mask — each token attends only to previous tokens (no future leakage). - Training loss: negative log-likelihood of the training corpus. Maximize the probability of each actual next token. - Used by: GPT-1/2/3/4, LLaMA, Mistral, Claude. The dominant paradigm for generative models. **Masked Language Modeling (MLM) — BERT-style**: - Randomly mask 15% of input tokens. Predict the masked tokens from context (both left and right). - Bidirectional attention — each token sees the full context. Better for understanding tasks. - Used by: BERT, RoBERTa, DeBERTa. Dominant for classification, NER, and extractive tasks. **Prefix Language Modeling — T5/UL2**: - Encoder-decoder architecture. Encoder processes the input (prefix) bidirectionally. Decoder generates the output (continuation/answer) autoregressively. - Flexible: handles both understanding (encode passage → decode answer) and generation (encode prompt → decode text). **Scaling Laws** Compute-optimal training (Chinchilla, Hoffmann et al.): - Loss ∝ N^{-0.076} × D^{-0.095}, where N = parameters, D = training tokens. - Optimal allocation: tokens ≈ 20 × parameters. A 70B parameter model should train on ~1.4T tokens. - Undertrained models (too few tokens per parameter) waste compute — better to train a smaller model on more data. **Training Data** - **Common Crawl**: Web-scraped text. Largest source (petabytes). Requires heavy filtering (deduplication, quality filtering, toxic content removal). - **Books**: BookCorpus, Pile-of-Law, etc. High quality, long-form text. - **Code**: GitHub, Stack Overflow. Improves reasoning and structured output generation. - **Curated Datasets**: Wikipedia, academic papers, instruction-following data. - **Data Quality > Quantity**: LLaMA trained on 1.4T tokens of curated data matches GPT-3 (trained on 300B lower-quality tokens) at 1/10th the size. Filtering, deduplication, and domain balancing are critical. **Training Infrastructure** Training a frontier LLM: - GPT-4 scale: ~25,000 GPUs × 90-120 days = ~$100M compute cost. - LLaMA 70B: 2,048 A100 GPUs × 21 days. Uses FSDP (Fully Sharded Data Parallel) + tensor parallelism. - Stability: checkpoint every 1-2 hours. Hardware failures are frequent at scale — training must be resumable. Loss spikes require manual intervention (rollback, adjust learning rate). Language Model Pretraining is **the self-supervised foundation that transforms raw text into general-purpose language intelligence** — the compute-intensive phase that extracts the statistical patterns of human language and world knowledge into neural network parameters, creating the foundation models that power modern NLP.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account