Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,102 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,059 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 939 views
Topics
Latest Articles
View All →Linux Performance Tuning for Containers and Kubernetes Nodes
Container performance problems usually live in the node kernel, not your app. Here is what we tune, why, and how we measure before touching anything.
Best Practices: Blue-Green Deployment Guardrails
Blue/green sounds simple until your green cluster has a memory leak and you've already sent 50% of traffic there. The guardrails are what make it safe.
Terraform Cloud Cost Controls: Budgets, Policies, and Tagging
Manage cloud spend with Terraform: cost estimation, tagging, and policy-as-code.
Best Practices: Kernel and Package Patch Management
We had four different patch cadences across our fleet and routinely missed CVEs by weeks. The unified workflow that finally caught up.
Docker Security Best Practices: Images, Runtime, and Supply Chain
Harden container images and runtime. Image scanning, minimal base, and supply chain security.
Best Practices: AWS Cost Control with Tagging and Budgets
AWS bill grew 40% YoY for two years before we got serious. Tagging, scoped budgets, and a weekly review meeting did 80% of the work.
Best Practices: Ansible Role Design for Large Teams
A team of 30 engineers all editing the same monolithic Ansible repo doesn't work. Here's the role taxonomy and review process that did.
Observability with OpenTelemetry: Traces, Metrics, and Logs
Unify traces, metrics, and logs with OpenTelemetry. Instrumentation, sampling, and backend-agnostic pipelines.
Best Practices: Terraform State Isolation by Environment
One Terraform state file per environment sounds obvious until you watch a dev plan touch a prod resource. Here's how we actually isolate state and the mistakes we made getting there.
Best Practices: GitHub Actions Pipeline Reliability
Our CI was 73% green at the worst point. People trusted it less than coin flips. Six things we did to get to 96%, in rough order of impact.
Best Practices: Docker Image Hardening for Production
Our base image went from 1.2 GB and 200+ CVEs to 80 MB and 4 CVEs. Most of the work wasn't clever — it was deletion.
Best Practices: Kubernetes Cluster Upgrade Strategy
We upgraded a 60-node EKS cluster from 1.27 to 1.31 over six months. Four minor versions, one bad surprise, zero customer impact. Here's the playbook.