Home Knowledge Base Spot/Preemptible Instances for ML

Spot/Preemptible Instances for ML

What are Spot Instances? Spare cloud capacity available at 60-90% discount, but can be terminated with short notice (2 minutes on AWS).

Use Cases for ML

Use CaseSuitability
Training with checkpointsExcellent
Batch inferenceGood
Development/testingExcellent
Real-time inferenceRisky without fallback
Hyperparameter tuningExcellent

AWS Spot Configuration

# Boto3 spot instance request
ec2 = boto3.client("ec2")

response = ec2.request_spot_instances(
    InstanceCount=1,
    Type="persistent",
    LaunchSpecification={
        "ImageId": "ami-xxx",
        "InstanceType": "p3.2xlarge",
        "KeyName": "my-key",
    },
    SpotPrice="3.00"  # Max price you will pay
)

EKS Spot Node Groups

# Terraform
resource "aws_eks_node_group" "spot_gpu" {
  cluster_name    = aws_eks_cluster.main.name
  node_group_name = "spot-gpu"

  capacity_type = "SPOT"
  instance_types = ["g4dn.xlarge", "g4dn.2xlarge", "g5.xlarge"]

  scaling_config {
    desired_size = 3
    max_size     = 10
    min_size     = 0
  }

  labels = {
    "capacity-type" = "spot"
  }

  taint {
    key    = "spot"
    value  = "true"
    effect = "NO_SCHEDULE"
  }
}

Kubernetes Spot Tolerations

apiVersion: apps/v1
kind: Deployment
spec:
  template:
    spec:
      tolerations:
      - key: "spot"
        operator: "Equal"
        value: "true"
        effect: "NoSchedule"
      nodeSelector:
        capacity-type: spot

Handling Interruptions

Checkpointing

# Save checkpoints frequently during training
for epoch in range(epochs):
    train_one_epoch(model)

    # Save checkpoint every epoch
    torch.save({
        "epoch": epoch,
        "model_state": model.state_dict(),
        "optimizer_state": optimizer.state_dict(),
    }, f"checkpoints/epoch_{epoch}.pt")

Interruption Handler

# AWS spot interruption handler
import requests

def check_interruption():
    try:
        response = requests.get(
            "http://169.254.169.254/latest/meta-data/spot/instance-action",
            timeout=1
        )
        if response.status_code == 200:
            # 2-minute warning, save and shutdown
            save_checkpoint()
            return True
    except:
        pass
    return False

Cost Comparison

Instance TypeOn-DemandSpotSavings
p3.2xlarge$3.06/hr$0.92/hr70%
g4dn.xlarge$0.526/hr$0.16/hr70%

Best Practices

spot instancepreemptiblecost

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.