Transformer-XL (Extra Long) is a transformer architecture designed for modeling long-range dependencies by introducing segment-level recurrence and relative positional encoding, enabling the model to capture dependencies beyond the fixed context window of standard transformers. Transformer-XL caches and reuses hidden states from previous segments during both training and inference, effectively extending the receptive field without proportionally increasing computation.
Why Transformer-XL Matters in AI/ML: Transformer-XL addresses the context fragmentation problem of standard transformers, where fixed-length segments break long-range dependencies at segment boundaries, by introducing recurrent connections between segments.
• Segment-level recurrence — Hidden states from the previous segment are cached and concatenated with the current segment's states during self-attention computation, allowing information to flow across segment boundaries; the effective context length grows linearly with the number of layers (L × segment_length) • Relative positional encoding — Standard absolute positional embeddings fail when states from different segments are mixed; Transformer-XL introduces relative position biases in the attention score computation that depend only on the distance between query and key positions, naturally handling cross-segment attention • Extended context during evaluation — At inference time, Transformer-XL can use much longer cached history than the training segment length, enabling context lengths of thousands of tokens with models trained on 512-token segments • No context fragmentation — Standard transformers trained on fixed chunks lose all information at segment boundaries; Transformer-XL's recurrence ensures information flows across boundaries, capturing dependencies that span multiple segments • State reuse efficiency — Cached hidden states from the previous segment do not require gradient computation, reducing the additional training cost of recurrence; only the forward pass through cached states is needed
| Property | Transformer-XL | Standard Transformer |
|---|---|---|
| Context Window | L × segment_length | Fixed segment_length |
| Cross-Segment Info Flow | Yes (recurrence) | No (independent segments) |
| Positional Encoding | Relative | Absolute |
| Cached States | Previous segment hidden states | None |
| Evaluation Context | Extensible (>> training) | Fixed (= training) |
| Training Overhead | ~20-30% (cache forward pass) | Baseline |
| Dependencies Captured | Long-range (thousands of tokens) | Within-segment only |
Transformer-XL fundamentally solved the context fragmentation problem in autoregressive language modeling by introducing segment-level recurrence with relative positional encoding, enabling transformers to capture dependencies spanning thousands of tokens and establishing the architectural foundation for subsequent long-context models including XLNet and Compressive Transformer.
Related Topics
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.