Home Knowledge Base Recurrent LLM Architectures (RWKV, Mamba)

Recurrent LLM Architectures (RWKV, Mamba) are models that achieve linear-time sequence processing by replacing quadratic self-attention with recurrent or state-space mechanisms, enabling efficient processing of very long sequences while maintaining competitive quality with transformer-based LLMs — reviving recurrent approaches at the billion-parameter scale.

The Transformer Bottleneck: Standard self-attention has O(N²) time and memory complexity in sequence length N. Even with Flash Attention (O(N) memory), the O(N²) compute remains. For sequence lengths of 100K-1M+ tokens, this quadratic cost becomes prohibitive. Recurrent architectures process sequences in O(N) time with O(1) memory per step.

RWKV (Receptance Weighted Key Value):

ComponentMechanismPurpose
Time-mixingWKV attention with linear complexitySequence mixing (replaces attention)
Channel-mixingGated FFN with shifted tokensFeature interaction
Token shiftLinear interpolation with previous tokenLocal context injection

RWKV replaces softmax attention with a weighted sum that can be computed recurrently: wkv_t = (Σ e^(w_s + k_s) · v_s) / (Σ e^(w_s + k_s)) where w provides exponential decay weights. This is computable as a running sum (RNN mode) or as a parallelizable scan (training mode). RWKV scales to 14B+ parameters with quality approaching transformer LLMs of similar size.

Mamba (Selective State Space Model):

Mamba builds on structured state space models (S4) but adds input-dependent (selective) parameters: the state transition matrices A, B, C vary based on the input at each step, enabling the model to selectively remember or forget information — unlike time-invariant SSMs where the same dynamics apply regardless of input content.

Mamba Architecture: Each Mamba block contains: a selective SSM layer (replaces attention), a gated MLP path, and residual connections. The selective SSM: h_t = A_t · h_{t-1} + B_t · x_t, y_t = C_t · h_t, where A_t, B_t, C_t are functions of the input x_t. This selectivity is crucial — it allows the model to decide what to store in its fixed-size state based on input content.

Training Efficiency: Despite being recurrent at inference, both RWKV and Mamba use parallel scan algorithms during training: the recurrence h_t = A_t · h_{t-1} + B_t · x_t is a linear recurrence that can be parallelized using the associative scan primitive, computing all hidden states in O(N log N) time on GPUs. This provides transformer-like training parallelism with RNN-like inference efficiency.

Inference Advantage:

AspectTransformerMamba/RWKV
Generation per tokenO(N) (KV cache lookup)O(1) (fixed state update)
Memory per tokenO(N) (growing KV cache)O(d²) (fixed state size)
Prefill costO(N²)O(N)
Long context costGrows linearly with NConstant

Quality Comparison: Mamba-2 (2024) matches transformer quality on language modeling up to ~3B parameters. At larger scales, pure recurrent models show a small but persistent gap on tasks requiring precise long-range retrieval (finding a specific fact buried deep in context). Hybrid architectures (interleaving attention and Mamba layers) close this gap while retaining most efficiency benefits.

Recurrent LLM architectures represent a fundamental challenge to the transformer's dominance — demonstrating that linear-time sequence models can achieve competitive quality while offering dramatically better inference efficiency for long sequences, potentially enabling a new generation of models that process books, codebases, and video streams as native context.

recurrent llmlinear rnn llmrwkv architectureretnet architecturelinear attention recurrence

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.