Encoder-decoder architecture uses both components for sequence-to-sequence tasks requiring input understanding and output generation. Architecture: Encoder processes input with bidirectional attention, decoder generates output with causal attention plus cross-attention to encoder. Cross-attention: Each decoder layer attends to encoder outputs, connecting input understanding to generation. Representative models: T5, BART, mT5, FLAN-T5, original Transformer (for translation). Training: Often uses denoising objectives (reconstruct corrupted text), span corruption (T5), or seq2seq tasks directly. Use cases: Translation, summarization, question answering, text-to-text tasks generally. T5 approach: Frame all tasks as text-to-text (same model for translation, summarization, QA, classification). Advantages: Natural fit for seq2seq, encoder provides rich input representation, decoder generates freely. Comparison: More complex than decoder-only, but potentially more efficient for conditional generation tasks. Current status: Less popular than decoder-only for general LLMs, but still used for specific applications like translation.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.