label flipping

**Label Flipping** is a **data poisoning attack that corrupts training data by changing the labels of selected examples** — the attacker flips a fraction of training labels (e.g., positive → negative) to degrade model performance or introduce targeted biases. **Label Flipping Strategies** - **Random Flipping**: Flip labels of a random subset of training data — degrades overall accuracy. - **Targeted Flipping**: Flip labels near a specific decision region — cause misclassification in targeted areas. - **Strategic Selection**: Use influence functions to select the most impactful examples to flip. - **Fraction**: Even flipping 5-10% of labels can significantly degrade model performance. **Why It Matters** - **Crowdsourced Labels**: Datasets with crowdsourced annotations are vulnerable to label corruption. - **Hard to Detect**: A few flipped labels in a large dataset are difficult to identify without clean reference data. - **Defense**: Data sanitization, robust loss functions (symmetric cross-entropy), and label noise detection methods mitigate flipping. **Label Flipping** is **poisoning through mislabeling** — corrupting training labels to trick the model into learning incorrect decision boundaries.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account