Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,100 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,055 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 933 views
Topics
Latest Articles
View All →Blameless Postmortems: The Template and Facilitation That Works
Our early postmortems quietly assigned blame and taught people to hide mistakes. Here's the template and the facilitation rules that finally made them honest and useful.
Token Budgeting for Long-Context Prompts: What to Cut First
A 180k-token context window is not a license to stuff everything in. Here's how we cut prompt size 60% without hurting answer quality, and what to trim first.
Multi-Provider LLM Gateways: Routing, Fallback, and Cost Control
When our single LLM provider had a 40-minute outage, every AI feature went dark. A gateway with routing and fallback fixed that, and cut spend 30% as a bonus.
Error Budgets to Roadmap: Turning Reliability Into Prioritization
Reliability arguments used to be shouting matches between SRE and product. An error budget turned them into arithmetic. Here's how we made the number drive the roadmap.
Four Signals That Matter: Choosing SLIs Users Actually Feel
Most SLI dashboards track things nobody notices. Here's how we picked the handful of signals that map to real user pain, and dropped the vanity metrics.
Serverless Cold Starts: Measuring and Fixing Them on Lambda
A p99 that jumped to 3.4 seconds during traffic ramps turned out to be cold starts. Here's how we measured them properly and cut the tail, with real init timings.
Multi-Region Failover with Route 53: Health Checks and Gotchas
Our failover config looked perfect in the console and did nothing during a real outage. Here's the health-check design that actually flipped regions when it mattered.
Streaming LLM Responses: SSE, Backpressure, and Cancellation
Users hit stop, but our server kept paying for tokens for another 40 seconds. Here's how we wired real cancellation and backpressure into an SSE streaming endpoint.
Cloudflare R2 vs S3: Egress-Free Object Storage in Practice
We moved 40 TB of user media off S3 and cut the bill by 70 percent, mostly by killing egress fees. Here's where R2 won and where we kept S3 anyway.
GitHub Actions Reusable Workflows: DRY Pipelines at Org Scale
We had the same 180-line build workflow copy-pasted into 60 repos. Fixing one bug meant 60 PRs. Here's the reusable-workflow setup that made it one.
Kustomize Overlays That Scale Across Environments
Our overlay tree grew to seven environments and started copy-pasting the same patch into each. Here's the component-based layout that stopped the drift.
Kubernetes Ingress vs Gateway API: Migrating Without Downtime
We moved 40 services off the nginx Ingress controller onto Gateway API without a single dropped connection. Here's the routing overlap trick that made it boring.