Skip to main content
Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.

AWS Reserved Instances vs Savings Plans vs Spot — When Each Fits

KU
Kiril Urbonas
3 months ago 7 min readUpdated yesterday22 views

Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.

Key takeaways

  • Three discounting mechanisms, three different commitments.
  • The rules of thumb we use to pick, and the mistakes we made before settling on them.

AWS has three main ways to pay less than on-demand: Reserved Instances (RIs), Savings Plans, and Spot. They sound similar; they're not. Picking the wrong one costs money and locks you in. After running our cost discipline for ~three years, here's the split that works and the rules of thumb behind it.

The three mechanisms#

Reserved Instances. You commit to a specific instance type (e.g., m5.xlarge) in a specific region (or AZ) for 1 or 3 years. Discount: 30-60% off on-demand. Strictly bound to that instance type — if your workload moves to m6i, the RI is wasted.

Savings Plans (Compute). You commit to a dollar-per-hour spend rate for 1 or 3 years. The discount applies to any compute (EC2, Fargate, Lambda) regardless of instance type, family, or region. Discount: 25-50% off on-demand.

Spot. You bid for unused EC2 capacity. AWS can reclaim the instance with 2 minutes notice. Discount: up to 90% off on-demand, no commitment.

The shape of each commitment matters more than the discount percentage:

  • RI = commit to what.
  • Savings Plan = commit to how much.
  • Spot = no commit, accept interruption.

When RIs still make sense#

Honest answer: rarely, in the modern AWS pricing landscape. Savings Plans cover almost every RI use case more flexibly. We've shrunk our RI portfolio every quarter.

The remaining cases:

  • Predictable, immovable workload that won't change instance type. A specific old service that's frozen and won't be re-architected. RIs match it perfectly; Savings Plans charge the same. Either works.
  • RDS, ElastiCache, Redshift. Savings Plans don't cover these (they're EC2/Fargate/Lambda only). RIs for these services still earn their keep.

For new EC2 commitments, default to Savings Plans.

When Savings Plans fit (most of the time)#

The bulk of our compute commitment is Compute Savings Plans (the most flexible variant).

We size them at ~70% of our trailing 90-day average compute spend. That captures the steady-state workload while leaving room for:

  • Growth. Don't commit at 100% if you're growing; you'll over-commit.
  • Spot. Spot doesn't count against Savings Plan utilization, so we want some compute available for Spot.
  • Temporary spikes. Maintenance windows, migrations, load tests.

1-year vs 3-year: 1-year is ~25% off; 3-year is ~50% off. We default to 1-year. 3-year is a long time to predict your compute shape; the extra discount doesn't compensate for the lock-in risk in our case. Larger orgs with very stable workloads can do 3-year math; ours isn't stable enough.

No upfront vs partial upfront vs all upfront: the difference is small. We pay no upfront for cash-flow reasons; if your finance team prefers up-front, the discount is marginal.

When Spot is the answer#

Spot is the biggest discount but with the constraint that AWS can take the instance back at any time. Workloads that fit:

  • Stateless workers consuming a queue. If the worker dies mid-task, the message goes back to the queue and another worker picks it up. We run our entire async processing fleet on Spot.
  • Batch jobs with checkpointing. Long-running training jobs that periodically save state. Interruption = restart from the last checkpoint.
  • Stateless web servers behind a load balancer. If a Spot instance disappears, the LB removes it. Need over-provisioning (run more replicas than the minimum so an interruption doesn't drop you below capacity).
  • CI runners. Builds restart if interrupted.

Workloads that don't fit Spot:

  • Stateful workloads (databases, in-memory caches with no replicas).
  • Workloads with tight cold-start budgets.
  • Single-replica services.

Our split: ~40% of compute on Spot, ~50% covered by Savings Plans, ~10% on-demand for in-flight work or workloads not yet migrated.

The Spot architecture#

Spot is operationally heavier than RIs or Savings Plans. Things to get right:

Diversification. Don't run all your Spot on one instance type. AWS Spot pools are per-instance-type-per-AZ; if all your Spot is m5.xlarge in us-east-1a, a capacity event takes you all down at once. We use instance type "families" — accept any of m5.large, m5.xlarge, m5a.large, m6i.large, etc. — diversifying across pools.

Karpenter (or Cluster Autoscaler with mixed instance policies) handles this. We tell Karpenter "any instance with 4-16 vCPU and burst-capable network," and it picks whatever's cheapest in any AZ.

Graceful shutdown. Spot gives 2 minutes notice via the instance metadata endpoint. Listen for it; drain workloads cleanly. Karpenter does this for Kubernetes pods; for raw EC2 you handle it yourself.

Capacity Rebalance Notification. Newer than the 2-minute interruption notice; gives advanced warning when AWS expects to reclaim soon. Use it to pre-emptively drain.

Persistent Spot requests are deprecated. Don't try to use them. The current model is Spot via Auto Scaling Groups, Karpenter, or Fargate Spot.

Things we got wrong#

Over-committing on RIs early. First year we bought 3-year all-upfront RIs for our then-current instance types. By year 2 we'd migrated to a newer family and were paying for RIs we couldn't use. AWS lets you sell unused RIs on a marketplace, but recovery is partial.

Treating Spot as too risky. Early on we ran Spot only on dev clusters. The 70-90% savings vs on-demand are worth a lot of operational effort. Once we built the diversification + graceful-shutdown discipline, Spot became boring.

Buying Savings Plans during a growth spike. Committed to 80% of our then-current spend during a quarter when we were 50% over-provisioned. Spend dropped after rightsizing; we were stuck under-utilizing the Savings Plan. Now we model the commitment against our post-optimization baseline, not pre.

Ignoring rate differences across regions. Savings Plans are per-region (kind of — Compute SPs apply across regions but at the regional rate). Moving workload between regions can leave Savings Plans poorly utilized.

The operational discipline#

What we do continuously:

  • Monthly RI/SP review. Look at utilization (target > 95% of commitment used). Coverage (target ~70% of compute covered by SP or RI). Adjust.
  • Quarterly compute audit. Are we running on the right instance families? AWS releases new generations; the new ones are usually cheaper per unit performance. Migrate.
  • Spot interruption rate by pool. If a specific instance type/AZ is interrupting too often, drop it from the pool.
  • Tag every workload with cost-owner. Without this, you can't have informed conversations about who's spending what.

What we monitor#

  • Savings Plan utilization. Below 95% = wasted commitment.
  • On-demand spend. Should be small remainder; if it grows, we're under-committed.
  • Spot interruption rate. Healthy is 1-5% per day; > 10% means pool diversification needs work.
  • Cost per request / per service. The denominator changes; absolute cost can hide efficiency gains.

Things that surprised us#

Savings Plans are very flexible. When we migrated services from EC2 to Fargate, the same Savings Plan applied. From Fargate to Lambda — same. The "Compute" Savings Plan is genuinely cross-compute.

Spot capacity is plentiful — usually. Major capacity events (AWS reclaiming Spot broadly) are rare. We've had two in three years. Diversification covered them.

The biggest savings come from rightsizing, not discounts. A 30% discount on a wrong-sized instance is less than a 50% rightsizing. Audit usage first; then discount what's left.

AWS pricing is intentionally complex. The three discount mechanisms cover overlapping cases; picking is judgment. The framework above (Savings Plans for predictable steady-state, Spot for interruptible workloads, RIs only for non-EC2 services) has worked for us across orders of magnitude in spend. Your shape may differ; the meta-pattern of measure-commit-then-optimize is what transfers.

Explore topics:AWSCloud
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

About Kiril Urbonas

DevOps Engineer

537 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.