Home Knowledge Base Transformer-XL (Extra Long)

Transformer-XL (Extra Long) is a transformer architecture designed for modeling long-range dependencies by introducing segment-level recurrence and relative positional encoding, enabling the model to capture dependencies beyond the fixed context window of standard transformers. Transformer-XL caches and reuses hidden states from previous segments during both training and inference, effectively extending the receptive field without proportionally increasing computation.

Why Transformer-XL Matters in AI/ML: Transformer-XL addresses the context fragmentation problem of standard transformers, where fixed-length segments break long-range dependencies at segment boundaries, by introducing recurrent connections between segments.

Segment-level recurrence — Hidden states from the previous segment are cached and concatenated with the current segment's states during self-attention computation, allowing information to flow across segment boundaries; the effective context length grows linearly with the number of layers (L × segment_length) • Relative positional encoding — Standard absolute positional embeddings fail when states from different segments are mixed; Transformer-XL introduces relative position biases in the attention score computation that depend only on the distance between query and key positions, naturally handling cross-segment attention • Extended context during evaluation — At inference time, Transformer-XL can use much longer cached history than the training segment length, enabling context lengths of thousands of tokens with models trained on 512-token segments • No context fragmentation — Standard transformers trained on fixed chunks lose all information at segment boundaries; Transformer-XL's recurrence ensures information flows across boundaries, capturing dependencies that span multiple segments • State reuse efficiency — Cached hidden states from the previous segment do not require gradient computation, reducing the additional training cost of recurrence; only the forward pass through cached states is needed

PropertyTransformer-XLStandard Transformer
Context WindowL × segment_lengthFixed segment_length
Cross-Segment Info FlowYes (recurrence)No (independent segments)
Positional EncodingRelativeAbsolute
Cached StatesPrevious segment hidden statesNone
Evaluation ContextExtensible (>> training)Fixed (= training)
Training Overhead~20-30% (cache forward pass)Baseline
Dependencies CapturedLong-range (thousands of tokens)Within-segment only

Transformer-XL fundamentally solved the context fragmentation problem in autoregressive language modeling by introducing segment-level recurrence with relative positional encoding, enabling transformers to capture dependencies spanning thousands of tokens and establishing the architectural foundation for subsequent long-context models including XLNet and Compressive Transformer.

memory transformer-xlllm architecture

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.