Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,100 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,055 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 933 views
Topics
Latest Articles
View All →Distributed Tracing with OpenTelemetry — What We Ship, What We Skip
How we run OpenTelemetry across ~40 services. The instrumentation that earns its place, the patterns we abandoned, and what tracing actually catches that metrics don't.
Postgres Autovacuum — Tuning From Production Stalls
A 2 AM incident, the autovacuum settings that caused it, and the parameter changes that prevented the next one. The discipline that took our biggest Postgres host from periodic stalls to steady.
Database Connection Pooling at Scale: PgBouncer, RDS Proxy, Application Pool
Three layers of pooling, three different jobs. We learned the hard way which to use when. Real numbers from a 8k-connection workload.
Backstage Adoption: From Demo to 80% Service Coverage in 6 Months
We launched Backstage in October. Six months in, 80% of services are catalogued, on-boarding takes a third of the time, and we mostly know what owns what.
Cloudflare Workers vs Vercel Edge: A Latency-Cost Comparison
We deployed the same edge function on both platforms and measured for a quarter. Where each wins, where each loses, and the surprises along the way.
eBPF for SREs: Three Real Diagnoses That Saved Hours
We started using eBPF tooling for ad-hoc production debugging six months ago. Three real incidents where it cut investigation time from hours to minutes.
LLM Output Validation: Schema-First Prompt Engineering Patterns
We invalidate ~6% of LLM outputs before they reach a downstream system. Here's how we structure prompts and validators to catch malformed responses early.
Argo Rollouts: Canary Deployments That Caught a $40k Bug
A two-line config change to an Argo Rollouts analysis template caught a regression that would have cost ~$40k in API spend before we noticed. Here's the pattern.
Pulumi vs Terraform: What 18 Months of Production Taught Us
We ran Pulumi in TypeScript and Terraform in HCL side by side across 60+ services. Each won different categories of work. Here's the breakdown.
GCP Workload Identity Federation: Replacing Service Account Keys
We deleted every static GCP service account key in our org over six weeks. Here's the migration plan, the gotchas, and the policies we now enforce.
Linux Memory Management: When OOM Killer Strikes Your K8s Pods
Three production OOM incidents that taught us how kubelet, containerd, and the kernel actually decide which process dies. With debugging commands you'll wish you had earlier.
GitHub Actions Self-Hosted Runners: Why We Switched and What Broke
Bills hit $3,400/mo for runner minutes. We moved to self-hosted on EKS spot. The savings were real; the surprises were too.