AWS Cost Optimization: The Discipline That Actually Sticks
The initial cleanup is the easy part. The savings that last come from tagging, per-team budgets, and a monthly review, not a one-time audit.
Key takeaways
- The initial cleanup is the easy part.
- The savings that last come from tagging, per-team budgets, and a monthly review, not a one-time audit.
Most cost-optimization advice focuses on the initial cleanup: find the waste, fix it, save money. The reality is that costs creep back up when there is no ongoing discipline behind the cleanup. We have cut our AWS bill more than once over the years, and the savings that actually held were the ones backed by process, not the ones from a single sweep. This is the operational side: the practices that keep savings sticky after the first pass.
Why the bill grows even when usage does not#
Without discipline, AWS costs grow in predictable ways. New services get over-provisioned "to be safe." Reserved capacity expires and gets renewed at a higher level because consumption crept up. Old resources accumulate: snapshots, orphaned EBS volumes, idle load balancers nobody remembers creating. New AWS services get adopted without a cost pass. Engineers default to a bigger instance type because tuning is work nobody has time for.
Every one of these is invisible day to day. Over a year, the bill can grow 30 to 50 percent even with flat usage. A cleanup brings it back down. Without ongoing discipline, the cycle repeats on schedule.
Tagging is the foundation, not a nice-to-have#
Every resource carries tags: team, env, service, cost-center. Without them, attribution is impossible and accountability evaporates into "the AWS bill went up" with nobody able to say why.
We enforce this three ways: organization-level tag policies that block creation of untagged resources where the provider supports it, cost allocation tags activated in the billing dashboard, and a monthly script that flags anything untagged for its owner. After the initial enforcement push, untagged-resource volume drops close to zero. New resources get tagged at creation; the existing fleet got tagged in a batch cleanup that took about a month.
The payoff: the cost dashboard, filtered by team, becomes that team's bill. Actionable and specific, not an opaque line that went up.
Budgets per team, reviewed quarterly#
Set a budget per team, refreshed each quarter, with an alert at 80 percent. The team owner gets notified and a conversation happens before the number matters.
Budgets come from last quarter's actuals plus planned changes, not an arbitrary "cut 10 percent" target. When a team blows its budget, the conversation is not punitive: what changed, is it justified, what do we do about it. Sometimes the budget was wrong because nobody accounted for a new service. Sometimes there is a real anomaly, like a misconfigured Lambda or a forgotten test environment left running.
The discipline is not "stay under budget no matter what." It is "you know what you are spending, and surprises require an explanation."
The monthly review and the quarterly deep-dive#
A 30-minute meeting every month, same agenda every time: total cost against last month and the same month last year, the top five movers, any project spending more than expected, and status on last month's action items. This catches silent drift. A service whose cost grew 50 percent over three months gets caught at the first meeting, not discovered a year later.
Once a quarter, one team does a full cost deep-dive on their own services: right-sizing, spot opportunities, storage tiering, reservation coverage, zombie resources. This catches what the monthly review does not, the subtle waste that never produces a month-over-month spike but sits there costing money regardless.
Catch it before it ships#
New infrastructure gets a cost estimate in the pull request: expected monthly cost at baseline, cost at three times baseline, the reserved-versus-on-demand call and why, and a comparison against similar existing services. We generate these from Terraform plans with Infracost, so the estimate shows up as a comment on the diff itself. Reviewers can push back, "this is twice what similar services cost, why," and often there is a good answer. Sometimes there is not, and the resource gets right-sized before it ever runs.
AWS's own free tools do real work here too. Compute Optimizer looks at 14 days of utilization and suggests right-sizing; it is conservative, and when it says downsize, it is almost always right. Cost Anomaly Detection is ML-based spike detection wired to each team's Slack, and it has caught a service consuming ten times more compute after a deploy introduced a retry loop, an S3 PUT cost spike from unbatched small writes, and a Lambda invoking itself recursively. Each of those would have surfaced in the monthly review eventually; anomaly detection caught them within hours. Once native tooling stops being enough, and multi-account or multi-cloud spend usually gets there fast, a dedicated FinOps tools comparison is worth working through before you build your own.
Where the actual savings live#
For baseline workloads that do not churn, a one-year Compute Savings Plan covering roughly 70 percent of baseline EC2 spend, plus one-year RDS Reserved Instances for production databases. We do not go three years: workloads change too much for that to pay off. Reservation utilization gets reviewed monthly, and the rule is simple, above 90 percent buy more, 60 to 90 percent hold steady, below 60 percent let it lapse.
Spot covers batch jobs, CI runners, stateless web servers mixed with an on-demand baseline for reliability, and Kubernetes worker pools where pods can move freely. It does not cover stateful services, jobs that cannot checkpoint, or anything where an interruption cascades. On Kubernetes specifically, Karpenter handles spot well enough that we default to it and let specific pods opt into on-demand through tolerations.
S3 lifecycle policies run on every bucket by default through a shared Terraform module, tuned per bucket where the access pattern actually differs. Quarterly right-sizing pulls 30 days of CloudWatch metrics per service, computes p95 utilization, and flags anything under 30 percent as a candidate; most services right-size cleanly, and the exceptions get documented rather than forced.
The decision, concretely#
- Starting from zero? Tag everything before anything else. None of the rest of this works without attribution.
- Choosing between org-wide and per-team budgets? Per-team. Accountability lives at the team level or it lives nowhere.
- Only have time for one cadence? The monthly review catches drift faster than a quarterly one ever will; add the deep-dive once the monthly habit is established.
- Deciding what to automate? Automate detection and alerting, not deletion. An orphaned volume sometimes holds data someone needs; alert the owner and escalate after two weeks of silence.
The call we'd make#
Treat cost optimization as ongoing discipline, not a project with an end date. The teams that succeed tag everything from day one, run a monthly review and a quarterly deep-dive, and catch over-provisioning at the pull request instead of three months into production. The teams that struggle do a periodic cleanup, watch costs drift back up over the following year, and do the same cleanup again. The cleanup itself is the easy part. The discipline that prevents needing to repeat it is where the real savings live.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Building Production-Ready AI Applications with LangChain and Docker
We deploy LangChain apps in Docker on Kubernetes. The patterns that work, the LangChain-specific gotchas, and what we'd build differently next time.
Linux System Monitoring with Prometheus and Grafana
Set up comprehensive Linux system monitoring using Prometheus and Grafana. Monitor CPU, memory, disk, network, and application metrics with beautiful dashboards.
More from Cloud
Explore more articles in this category
Azure OpenAI's Sweden Central Outage: A Health Check Postmortem
A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster. The lesson has nothing to do with AI.
GKE Pod Snapshots: Cold Starts Drop 89%, If Your Nodes Match
Google's benchmarks show a 70B model restoring in 37 seconds instead of minutes. The catch is a hash and a hardware match that silently refuses to restore when either is off.
AWS Lost a Region for Good: Multi-AZ Is Not Disaster Recovery
AWS says it cannot restore data held only in Bahrain (me-south-1) or in one UAE zone. Multi-AZ gave availability, not recovery, and only cross-region copies survived.
You might have missed
Evergreen posts worth revisiting.