AWS Cost Audit: 7 Things We Found Wasting Money Every Month
A real cost audit uncovered idle load balancers, oversized RDS instances, and forgotten snapshots. Here's what we found and how we fixed each one.
Key takeaways
- A real cost audit uncovered idle load balancers, oversized RDS instances, and forgotten snapshots.
- Here's what we found and how we fixed each one.
On this page
AWS Cost Audit: 7 Things We Found Wasting Money Every Month#
After our AWS bill crossed $18,000/month for a 15-person startup, we did a proper audit. We found $6,200 in monthly waste. Here's every item.
1. Idle Load Balancers ($420/month)#
Three ALBs were still running from decommissioned staging environments. Each costs ~$16/month base plus LCU charges.
Fix: We added a Terraform lifecycle check that tags ALBs with the owning team and a TTL. A weekly Lambda deletes anything past its TTL with zero healthy targets.
2. Oversized RDS Instance ($1,800/month)#
Our production database was on db.r6g.2xlarge. CloudWatch showed average CPU at 12% and memory at 35%.
Fix: Downgraded to db.r6g.large during a maintenance window. Set up a CloudWatch alarm for CPU > 70% so we'll know when to scale back up.
3. Unattached EBS Volumes ($280/month)#
14 EBS volumes were sitting with status "available"—leftovers from terminated EC2 instances.
Fix: Scripted a check:
aws ec2 describe-volumes \
--filters Name=status,Values=available \
--query 'Volumes[*].{ID:VolumeId,Size:Size,Created:CreateTime}' \
--output table
Snapshot anything older than 30 days, then delete.
4. Forgotten Snapshots ($640/month)#
We had 2,400 EBS snapshots going back 3 years. Most were from AMIs we no longer use.
Fix: Implemented AWS Data Lifecycle Manager with a 90-day retention policy.
5. NAT Gateway Data Transfer ($1,400/month)#
Our NAT Gateway was processing 800GB/month. Much of it was S3 traffic from private subnets.
Fix: Added a VPC Gateway Endpoint for S3. Free, and it cut NAT traffic by 60%.
resource "aws_vpc_endpoint" "s3" {
vpc_id = aws_vpc.main.id
service_name = "com.amazonaws.us-east-1.s3"
route_table_ids = [aws_route_table.private.id]
}
6. Over-Provisioned Lambda Memory ($380/month)#
Every Lambda was set to 1024MB by default. AWS Power Tuning showed most needed 256MB.
Fix: Ran Power Tuning on our top 10 functions and right-sized them.
7. Missing Reserved Instances ($1,280/month)#
We were paying on-demand for 4 EC2 instances that had been running for 2 years.
Fix: Purchased 1-year no-upfront reserved instances for predictable workloads.
Best Practices for Ongoing Cost Control#
- Monthly cost review with a tagged cost allocation report
- Budget alerts at 80% and 100% of expected spend
- Terraform-managed resources so nothing is created outside of code
- Quarterly audit of unused resources using AWS Trusted Advisor
The $6,200/month we saved required about 8 hours of work. That's an annualized return of $74,400 for one day of effort.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
How We Cut Our Docker Image Size by 80% and Why It Matters
A real walkthrough of shrinking bloated Docker images from 1.2GB to 240MB using multi-stage builds, Alpine, and dependency auditing.
Prompt Engineering Patterns That Actually Work in Production
Battle-tested prompt patterns from running LLM features in production: structured output, chain-of-thought, and graceful failure handling.
More from Cloud
Explore more articles in this category
AWS Lost a Region for Good: Multi-AZ Is Not Disaster Recovery
AWS says it cannot restore data held only in Bahrain (me-south-1) or in one UAE zone. Multi-AZ gave availability, not recovery, and only cross-region copies survived.
The Cheapest Way to Centralize Logs at Scale
Cutting a log bill is not a procurement exercise. It is four decisions about what you drop at the agent, what you index, how long you keep it, and what you never send at all.
AWS Raised GPU Prices Twice in 2026: What to Do About It
EC2 Capacity Blocks went up around 15% in January and again in July. The increases track the memory shortage, and they change which GPU cloud is actually cheapest for your workload.
You might have missed
Evergreen posts worth revisiting.