AWS Reserved Instances vs Savings Plans vs Spot — When Each Fits
Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.
Key takeaways
- Three discounting mechanisms, three different commitments.
- The rules of thumb we use to pick, and the mistakes we made before settling on them.
AWS has three main ways to pay less than on-demand: Reserved Instances (RIs), Savings Plans, and Spot. They sound similar; they're not. Picking the wrong one costs money and locks you in. After running our cost discipline for ~three years, here's the split that works and the rules of thumb behind it.
The three mechanisms#
Reserved Instances. You commit to a specific instance type (e.g., m5.xlarge) in a specific region (or AZ) for 1 or 3 years. Discount: 30-60% off on-demand. Strictly bound to that instance type — if your workload moves to m6i, the RI is wasted.
Savings Plans (Compute). You commit to a dollar-per-hour spend rate for 1 or 3 years. The discount applies to any compute (EC2, Fargate, Lambda) regardless of instance type, family, or region. Discount: 25-50% off on-demand.
Spot. You bid for unused EC2 capacity. AWS can reclaim the instance with 2 minutes notice. Discount: up to 90% off on-demand, no commitment.
The shape of each commitment matters more than the discount percentage:
- RI = commit to what.
- Savings Plan = commit to how much.
- Spot = no commit, accept interruption.
When RIs still make sense#
Honest answer: rarely, in the modern AWS pricing landscape. Savings Plans cover almost every RI use case more flexibly. We've shrunk our RI portfolio every quarter.
The remaining cases:
- Predictable, immovable workload that won't change instance type. A specific old service that's frozen and won't be re-architected. RIs match it perfectly; Savings Plans charge the same. Either works.
- RDS, ElastiCache, Redshift. Savings Plans don't cover these (they're EC2/Fargate/Lambda only). RIs for these services still earn their keep.
For new EC2 commitments, default to Savings Plans.
When Savings Plans fit (most of the time)#
The bulk of our compute commitment is Compute Savings Plans (the most flexible variant).
We size them at ~70% of our trailing 90-day average compute spend. That captures the steady-state workload while leaving room for:
- Growth. Don't commit at 100% if you're growing; you'll over-commit.
- Spot. Spot doesn't count against Savings Plan utilization, so we want some compute available for Spot.
- Temporary spikes. Maintenance windows, migrations, load tests.
1-year vs 3-year: 1-year is ~25% off; 3-year is ~50% off. We default to 1-year. 3-year is a long time to predict your compute shape; the extra discount doesn't compensate for the lock-in risk in our case. Larger orgs with very stable workloads can do 3-year math; ours isn't stable enough.
No upfront vs partial upfront vs all upfront: the difference is small. We pay no upfront for cash-flow reasons; if your finance team prefers up-front, the discount is marginal.
When Spot is the answer#
Spot is the biggest discount but with the constraint that AWS can take the instance back at any time. Workloads that fit:
- Stateless workers consuming a queue. If the worker dies mid-task, the message goes back to the queue and another worker picks it up. We run our entire async processing fleet on Spot.
- Batch jobs with checkpointing. Long-running training jobs that periodically save state. Interruption = restart from the last checkpoint.
- Stateless web servers behind a load balancer. If a Spot instance disappears, the LB removes it. Need over-provisioning (run more replicas than the minimum so an interruption doesn't drop you below capacity).
- CI runners. Builds restart if interrupted.
Workloads that don't fit Spot:
- Stateful workloads (databases, in-memory caches with no replicas).
- Workloads with tight cold-start budgets.
- Single-replica services.
Our split: ~40% of compute on Spot, ~50% covered by Savings Plans, ~10% on-demand for in-flight work or workloads not yet migrated.
The Spot architecture#
Spot is operationally heavier than RIs or Savings Plans. Things to get right:
Diversification. Don't run all your Spot on one instance type. AWS Spot pools are per-instance-type-per-AZ; if all your Spot is m5.xlarge in us-east-1a, a capacity event takes you all down at once. We use instance type "families" — accept any of m5.large, m5.xlarge, m5a.large, m6i.large, etc. — diversifying across pools.
Karpenter (or Cluster Autoscaler with mixed instance policies) handles this. We tell Karpenter "any instance with 4-16 vCPU and burst-capable network," and it picks whatever's cheapest in any AZ.
Graceful shutdown. Spot gives 2 minutes notice via the instance metadata endpoint. Listen for it; drain workloads cleanly. Karpenter does this for Kubernetes pods; for raw EC2 you handle it yourself.
Capacity Rebalance Notification. Newer than the 2-minute interruption notice; gives advanced warning when AWS expects to reclaim soon. Use it to pre-emptively drain.
Persistent Spot requests are deprecated. Don't try to use them. The current model is Spot via Auto Scaling Groups, Karpenter, or Fargate Spot.
Things we got wrong#
Over-committing on RIs early. First year we bought 3-year all-upfront RIs for our then-current instance types. By year 2 we'd migrated to a newer family and were paying for RIs we couldn't use. AWS lets you sell unused RIs on a marketplace, but recovery is partial.
Treating Spot as too risky. Early on we ran Spot only on dev clusters. The 70-90% savings vs on-demand are worth a lot of operational effort. Once we built the diversification + graceful-shutdown discipline, Spot became boring.
Buying Savings Plans during a growth spike. Committed to 80% of our then-current spend during a quarter when we were 50% over-provisioned. Spend dropped after rightsizing; we were stuck under-utilizing the Savings Plan. Now we model the commitment against our post-optimization baseline, not pre.
Ignoring rate differences across regions. Savings Plans are per-region (kind of — Compute SPs apply across regions but at the regional rate). Moving workload between regions can leave Savings Plans poorly utilized.
The operational discipline#
What we do continuously:
- Monthly RI/SP review. Look at utilization (target > 95% of commitment used). Coverage (target ~70% of compute covered by SP or RI). Adjust.
- Quarterly compute audit. Are we running on the right instance families? AWS releases new generations; the new ones are usually cheaper per unit performance. Migrate.
- Spot interruption rate by pool. If a specific instance type/AZ is interrupting too often, drop it from the pool.
- Tag every workload with cost-owner. Without this, you can't have informed conversations about who's spending what.
What we monitor#
- Savings Plan utilization. Below 95% = wasted commitment.
- On-demand spend. Should be small remainder; if it grows, we're under-committed.
- Spot interruption rate. Healthy is 1-5% per day; > 10% means pool diversification needs work.
- Cost per request / per service. The denominator changes; absolute cost can hide efficiency gains.
Things that surprised us#
Savings Plans are very flexible. When we migrated services from EC2 to Fargate, the same Savings Plan applied. From Fargate to Lambda — same. The "Compute" Savings Plan is genuinely cross-compute.
Spot capacity is plentiful — usually. Major capacity events (AWS reclaiming Spot broadly) are rare. We've had two in three years. Diversification covered them.
The biggest savings come from rightsizing, not discounts. A 30% discount on a wrong-sized instance is less than a 50% rightsizing. Audit usage first; then discount what's left.
What to read next#
- Karpenter node provisioning patterns at scale — the Spot diversification mechanism we use
- Kubernetes resource requests — right-sizing across the cluster — the rightsizing prerequisite
- Cross-cloud identity federation patterns — adjacent cloud architectural pattern
- Internal developer platforms — Backstage in practice — surfacing cost back to teams
AWS pricing is intentionally complex. The three discount mechanisms cover overlapping cases; picking is judgment. The framework above (Savings Plans for predictable steady-state, Spot for interruptible workloads, RIs only for non-EC2 services) has worked for us across orders of magnitude in spend. Your shape may differ; the meta-pattern of measure-commit-then-optimize is what transfers.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Linux Network Debugging — tcpdump, ss, and eBPF in Anger
When the service is slow and the network is suspect, these are the tools we reach for, in this order, with the exact flags that find the answer.
Incident Post-Mortems That Drive Change (Not Theater)
Most post-mortems produce a document and no follow-through. The format, the discipline, and the cultural moves that actually convert incidents into engineering improvements.
More from Cloud
Explore more articles in this category
The Cheapest Way to Centralize Logs at Scale
Cutting a log bill is not a procurement exercise. It is four decisions about what you drop at the agent, what you index, how long you keep it, and what you never send at all.
AWS Raised GPU Prices Twice in 2026: What to Do About It
EC2 Capacity Blocks went up around 15% in January and again in July. The increases track the memory shortage, and they change which GPU cloud is actually cheapest for your workload.
The RAM Shortage Is Now a Line Item on Your Cloud Bill
Memory makers moved their wafers to HBM for AI accelerators, and DDR5 spot prices tripled. Here is how that reaches your instance bill and what actually reduces the exposure.
You might have missed
Evergreen posts worth revisiting.