end-of-sequence token
**End-of-sequence token** is the **special vocabulary token that marks logical completion of a sequence during training and inference** - it is the canonical boundary signal in autoregressive language modeling.
**What Is End-of-sequence token?**
- **Definition**: Dedicated tokenizer symbol indicating sequence termination.
- **Training Role**: Teaches model when output should end in supervised objectives.
- **Inference Role**: Decoder typically stops when EOS token is generated.
- **Notation**: Often referenced as EOS in model and tokenizer configuration.
**Why End-of-sequence token Matters**
- **Completion Accuracy**: Reliable EOS behavior prevents needless continuation text.
- **Cost Efficiency**: Early natural stopping lowers token usage.
- **Format Correctness**: Supports clean boundaries in multi-turn and structured interactions.
- **Model Interoperability**: Consistent EOS handling is required across runtimes and checkpoints.
- **Safety**: Acts as one layer of bounded-generation control.
**How It Is Used in Practice**
- **Config Verification**: Ensure EOS IDs match tokenizer files and serving runtime settings.
- **Prompt Design**: Avoid accidental EOS-like patterns in special-control token spaces.
- **Behavior Monitoring**: Track EOS stop rates and long-tail generation anomalies.
End-of-sequence token is **a core termination token in all sequence-generation systems** - stable EOS handling is essential for predictable and efficient inference.