Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,102 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,059 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 939 views
Topics
Latest Articles
View All →Cloud Cost Monitoring: Tracking and Optimizing AWS Spending
Building visibility into cloud costs that actually drives action. The dashboards we look at, the alerts that fire, and the queries we run.
Multi-Region Deployment: Building Resilient Cloud Applications
We run our app in two AWS regions for failover. The hard parts aren't the deployment — they're data consistency, traffic shifting, and the assumptions that break when "primary" is suddenly the wrong region.
AWS Lambda Optimization: Reducing Costs and Improving Performance
We run ~200 Lambda functions. Cold starts, memory tuning, and the cost-vs-latency trade-offs that actually move the bill.
DevOps Metrics and KPIs: Measuring Success
We track the four DORA metrics plus a handful of others. The trade-off between what's measurable and what's meaningful, and how we use the numbers.
Multi-Region Resilience: Failover, Data, and DNS
Design for region failure. Active/passive and active/active, data replication, and failover testing.
Canary Releases: Gradual Rollout Strategy
We've run canary deploys on most services for two years. The mechanics are easy; the metrics that decide "promote or roll back" are where the design is.
Blue-Green Deployments: Zero-Downtime Releases
We use blue-green for stateful services where canary doesn't fit. The actual mechanics, the data-layer subtleties, and when blue-green isn't the right answer.
Log Aggregation Strategies: Centralizing Your Logs
We collect ~800GB of logs per day across our fleet. The shape of our logging stack, what we keep, what we drop, and what we'd build differently.
Infrastructure Monitoring with Prometheus: Complete Setup Guide
A working Prometheus stack for a 40-node cluster: what we deploy, what we tune, and what we wish we'd known about cardinality two years ago.
Docker Multi-Stage Builds: Optimizing Image Size
A focused look at the techniques that shrink container images: which actually pay off, which are folklore, and the discipline that keeps images small over time.
Kubernetes Backup Strategies: Protecting Your Cluster Data
We've had to restore a Kubernetes cluster from backup twice. Once it worked. Once it took 14 hours. Here's the strategy we run now.
MLOps Pipelines: From Experiment to Production Models
Build MLOps pipelines for training, evaluation, and deployment. Reproducibility and monitoring.