Terraform Modules Done Right: Lessons from Managing 50+ Services
Practical patterns for Terraform modules at scale: versioning, composition, testing, and avoiding the monolith trap.
Key takeaways
Practical patterns for Terraform modules at scale: versioning, composition, testing, and avoiding the monolith trap.
Terraform Modules Done Right: Lessons from Managing 50+ Services#
After managing infrastructure for 50+ microservices with Terraform, we've learned which module patterns scale and which become nightmares. Here's what works.
The Monolith Trap#
Our first approach was one massive Terraform repo with everything in it. Plan took 12 minutes. A typo in a dev variable once triggered a production change. We split it up.
Pattern 1: Layered Modules#
We organize modules in three layers:
modules/
base/ # VPC, subnets, DNS zones
platform/ # EKS cluster, RDS, ElastiCache
service/ # Per-service: ALB, task def, IAM role
Each layer depends only on the layer below via remote state data sources:
data "terraform_remote_state" "platform" {
backend = "s3"
config = {
bucket = "terraform-state-prod"
key = "platform/terraform.tfstate"
region = "us-east-1"
}
}
resource "aws_lb_target_group" "service" {
vpc_id = data.terraform_remote_state.platform.outputs.vpc_id
# ...
}
Pattern 2: Versioned Module Registry#
We publish reusable modules to a private registry with semantic versioning:
module "service" {
source = "app.terraform.io/ourorg/service/aws"
version = "~> 2.0"
name = "payment-api"
environment = "production"
cpu = 512
memory = 1024
}
Rules we follow:
- Breaking changes = major version bump
- New optional variables = minor version bump
- Bug fixes = patch version bump
- Teams pin to major version (
~> 2.0), not exact
Pattern 3: Composition Over Configuration#
Instead of one module with 40 variables and 15 conditional blocks, we compose small modules:
module "alb" {
source = "./modules/alb"
# ...
}
module "ecs_service" {
source = "./modules/ecs-service"
target_group_arn = module.alb.target_group_arn
# ...
}
module "monitoring" {
source = "./modules/cloudwatch-alarms"
service_name = module.ecs_service.name
# ...
}
Each module does one thing. Connecting them is explicit, not hidden behind flags.
Pattern 4: Automated Testing#
We test modules with terraform validate, tflint, and integration tests:
# In CI pipeline
cd modules/service
terraform init -backend=false
terraform validate
tflint --init
tflint
# Integration test (creates real resources, then destroys)
cd tests/
go test -v -timeout 30m ./...
Best Practices Summary#
- Keep blast radius small: one service per state file
- Version your modules: semantic versioning with a changelog
- Compose, don't configure: small modules > one mega-module
- Test in CI: validate + lint + integration tests
- Use remote state carefully: read-only data sources, not cross-stack references
- Document inputs/outputs: a module without docs is a module no one trusts
Terraform at scale is a software engineering problem, not just an infrastructure problem. Treat your modules like libraries.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Linux Performance Troubleshooting: A Real Incident Walkthrough
Step-by-step debugging of a production Linux server hitting 100% CPU. From top to perf to the actual fix.
Incident Postmortems That Actually Prevent Repeat Failures
We wrote pretty postmortems for two years and kept hitting the same incidents. Here's what changed when we started writing ugly ones.
More from Infrastructure
Explore more articles in this category
Perplexity Left DynamoDB for CobbleDB: When Should You?
Perplexity built its own key-value store because DynamoDB's read path and bill stopped fitting 50 KB search items. Here is the checklist for when leaving is justified.
The Terraform Lock File Is Code: Review It Before You Init
A DPRK-linked group is mailing DevOps candidates Terraform take-home repos whose lock file points at a fake registry. terraform init then runs the attacker's provider.
Redis vs Memcached: Choosing a Cache in 2026
Both are fast in-memory stores, and both get picked by habit more than by requirements. Here is what actually differs and when each one is the right call.
You might have missed
Evergreen posts worth revisiting.