Spot/Preemptible Instances for ML
What are Spot Instances? Spare cloud capacity available at 60-90% discount, but can be terminated with short notice (2 minutes on AWS).
Use Cases for ML
| Use Case | Suitability |
|---|---|
| Training with checkpoints | Excellent |
| Batch inference | Good |
| Development/testing | Excellent |
| Real-time inference | Risky without fallback |
| Hyperparameter tuning | Excellent |
AWS Spot Configuration
# Boto3 spot instance request
ec2 = boto3.client("ec2")
response = ec2.request_spot_instances(
InstanceCount=1,
Type="persistent",
LaunchSpecification={
"ImageId": "ami-xxx",
"InstanceType": "p3.2xlarge",
"KeyName": "my-key",
},
SpotPrice="3.00" # Max price you will pay
)
EKS Spot Node Groups
# Terraform
resource "aws_eks_node_group" "spot_gpu" {
cluster_name = aws_eks_cluster.main.name
node_group_name = "spot-gpu"
capacity_type = "SPOT"
instance_types = ["g4dn.xlarge", "g4dn.2xlarge", "g5.xlarge"]
scaling_config {
desired_size = 3
max_size = 10
min_size = 0
}
labels = {
"capacity-type" = "spot"
}
taint {
key = "spot"
value = "true"
effect = "NO_SCHEDULE"
}
}
Kubernetes Spot Tolerations
apiVersion: apps/v1
kind: Deployment
spec:
template:
spec:
tolerations:
- key: "spot"
operator: "Equal"
value: "true"
effect: "NoSchedule"
nodeSelector:
capacity-type: spot
Handling Interruptions
Checkpointing
# Save checkpoints frequently during training
for epoch in range(epochs):
train_one_epoch(model)
# Save checkpoint every epoch
torch.save({
"epoch": epoch,
"model_state": model.state_dict(),
"optimizer_state": optimizer.state_dict(),
}, f"checkpoints/epoch_{epoch}.pt")
Interruption Handler
# AWS spot interruption handler
import requests
def check_interruption():
try:
response = requests.get(
"http://169.254.169.254/latest/meta-data/spot/instance-action",
timeout=1
)
if response.status_code == 200:
# 2-minute warning, save and shutdown
save_checkpoint()
return True
except:
pass
return False
Cost Comparison
| Instance Type | On-Demand | Spot | Savings |
|---|---|---|---|
| p3.2xlarge | $3.06/hr | $0.92/hr | 70% |
| g4dn.xlarge | $0.526/hr | $0.16/hr | 70% |
Best Practices
- Use multiple instance types for availability
- Checkpoint frequently during training
- Use on-demand for critical inference
- Set up interruption handlers
- Use diversified allocation strategies
spot instancepreemptiblecost
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.