Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.
AWS has three main ways to pay less than on-demand: Reserved Instances (RIs), Savings Plans, and Spot. They sound similar; they're not. Picking the wrong one costs money and locks you in. After running our cost discipline for ~three years, here's the split that works and the rules of thumb behind it.
Reserved Instances. You commit to a specific instance type (e.g., m5.xlarge) in a specific region (or AZ) for 1 or 3 years. Discount: 30-60% off on-demand. Strictly bound to that instance type — if your workload moves to m6i, the RI is wasted.
Savings Plans (Compute). You commit to a dollar-per-hour spend rate for 1 or 3 years. The discount applies to any compute (EC2, Fargate, Lambda) regardless of instance type, family, or region. Discount: 25-50% off on-demand.
Spot. You bid for unused EC2 capacity. AWS can reclaim the instance with 2 minutes notice. Discount: up to 90% off on-demand, no commitment.
The shape of each commitment matters more than the discount percentage:
Honest answer: rarely, in the modern AWS pricing landscape. Savings Plans cover almost every RI use case more flexibly. We've shrunk our RI portfolio every quarter.
The remaining cases:
For new EC2 commitments, default to Savings Plans.
The bulk of our compute commitment is Compute Savings Plans (the most flexible variant).
We size them at ~70% of our trailing 90-day average compute spend. That captures the steady-state workload while leaving room for:
1-year vs 3-year: 1-year is ~25% off; 3-year is ~50% off. We default to 1-year. 3-year is a long time to predict your compute shape; the extra discount doesn't compensate for the lock-in risk in our case. Larger orgs with very stable workloads can do 3-year math; ours isn't stable enough.
No upfront vs partial upfront vs all upfront: the difference is small. We pay no upfront for cash-flow reasons; if your finance team prefers up-front, the discount is marginal.
Spot is the biggest discount but with the constraint that AWS can take the instance back at any time. Workloads that fit:
Workloads that don't fit Spot:
Our split: ~40% of compute on Spot, ~50% covered by Savings Plans, ~10% on-demand for in-flight work or workloads not yet migrated.
Spot is operationally heavier than RIs or Savings Plans. Things to get right:
Diversification. Don't run all your Spot on one instance type. AWS Spot pools are per-instance-type-per-AZ; if all your Spot is m5.xlarge in us-east-1a, a capacity event takes you all down at once. We use instance type "families" — accept any of m5.large, m5.xlarge, m5a.large, m6i.large, etc. — diversifying across pools.
Karpenter (or Cluster Autoscaler with mixed instance policies) handles this. We tell Karpenter "any instance with 4-16 vCPU and burst-capable network," and it picks whatever's cheapest in any AZ.
Graceful shutdown. Spot gives 2 minutes notice via the instance metadata endpoint. Listen for it; drain workloads cleanly. Karpenter does this for Kubernetes pods; for raw EC2 you handle it yourself.
Capacity Rebalance Notification. Newer than the 2-minute interruption notice; gives advanced warning when AWS expects to reclaim soon. Use it to pre-emptively drain.
Persistent Spot requests are deprecated. Don't try to use them. The current model is Spot via Auto Scaling Groups, Karpenter, or Fargate Spot.
Over-committing on RIs early. First year we bought 3-year all-upfront RIs for our then-current instance types. By year 2 we'd migrated to a newer family and were paying for RIs we couldn't use. AWS lets you sell unused RIs on a marketplace, but recovery is partial.
Treating Spot as too risky. Early on we ran Spot only on dev clusters. The 70-90% savings vs on-demand are worth a lot of operational effort. Once we built the diversification + graceful-shutdown discipline, Spot became boring.
Buying Savings Plans during a growth spike. Committed to 80% of our then-current spend during a quarter when we were 50% over-provisioned. Spend dropped after rightsizing; we were stuck under-utilizing the Savings Plan. Now we model the commitment against our post-optimization baseline, not pre.
Ignoring rate differences across regions. Savings Plans are per-region (kind of — Compute SPs apply across regions but at the regional rate). Moving workload between regions can leave Savings Plans poorly utilized.
What we do continuously:
Savings Plans are very flexible. When we migrated services from EC2 to Fargate, the same Savings Plan applied. From Fargate to Lambda — same. The "Compute" Savings Plan is genuinely cross-compute.
Spot capacity is plentiful — usually. Major capacity events (AWS reclaiming Spot broadly) are rare. We've had two in three years. Diversification covered them.
The biggest savings come from rightsizing, not discounts. A 30% discount on a wrong-sized instance is less than a 50% rightsizing. Audit usage first; then discount what's left.
AWS pricing is intentionally complex. The three discount mechanisms cover overlapping cases; picking is judgment. The framework above (Savings Plans for predictable steady-state, Spot for interruptible workloads, RIs only for non-EC2 services) has worked for us across orders of magnitude in spend. Your shape may differ; the meta-pattern of measure-commit-then-optimize is what transfers.
Get the latest tutorials, guides, and insights on AI, DevOps, Cloud, and Infrastructure delivered directly to your inbox.
When the service is slow and the network is suspect, these are the tools we reach for, in this order, with the exact flags that find the answer.
Most post-mortems produce a document and no follow-through. The format, the discipline, and the cultural moves that actually convert incidents into engineering improvements.
Explore more articles in this category
The observability market is huge and the pricing is a minefield. This is the map to the tools that matter, what each is best at, and how to avoid a runaway bill.
Cloud bills grow quietly until someone asks why. This is the map for cutting spend without cutting reliability: where the money actually goes, the levers that work, and the tools worth paying for.
Static keys leak. The question isn't if but how fast you notice and how clean your response runbook is when the pager goes off.
Evergreen posts worth revisiting.