Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,100 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,055 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 933 views
Topics
Latest Articles
View All →Terraform Drift Detection in CI — Catching Out-of-Band Changes Before They Bite
State drift is silent until a deploy fails or an outage reveals it. The scheduled plan-and-diff pipeline that surfaces console hotfixes and manual edits while they're still cheap to reconcile.
RAG Retrieval Evaluation — Building an Offline Eval Harness Before You Ship
You can't improve retrieval you don't measure. The offline eval harness that lets us change embeddings, chunking, and rerankers with confidence instead of vibes — with the metrics that actually predict production quality.
Alert on Symptoms, Not Causes — SLO Burn-Rate Alerting in Practice
Cause-based alerts page you for things that don't matter and miss things that do. How we rebuilt alerting around SLO burn rates — multi-window, multi-burn-rate — and cut pages while catching more real pain.
Edge Caching with Stale-While-Revalidate — Fast and Fresh at the CDN
The cache-control header most teams under-use. How stale-while-revalidate and stale-if-error turned our CDN from a freshness liability into a latency and resilience win — with the gotchas.
LLM Output Validation — Schema-Constrained Generation in Production
Parsing model output with a regex and a prayer doesn't survive contact with traffic. The validation layers that keep structured LLM output reliable — constrained decoding, schema validation, and the repair loop.
CI Pipeline Caching That Actually Pays Off
Most CI caches either miss constantly or restore stale junk. The cache-key discipline, scope boundaries, and measurements that turned our pipeline cache from theatre into real minutes saved.
Observability — Correlating Logs, Metrics, and Traces in Anger
The "three pillars" framing misses the point — what matters is correlating across them. The patterns that earn their place and the tooling decisions that pay back.
Multi-Region — Active-Active vs Active-Passive, And What We Actually Run
The architectural choice is presented as binary; the practical answer is "depends on the workload." The patterns that earn their place and the failure modes we've hit.
RAG vs Fine-Tuning — Picking the Right Tool, Honestly
They solve different problems. RAG injects knowledge; fine-tuning changes behavior. The decision criteria, the hybrid pattern, and what we'd do over.
Kubernetes NetworkPolicies in Practice
Default-deny, namespace isolation, egress control — the patterns we use, the gotchas around DNS, and where Cilium changed our calculus.
Database Sharding — The Choices We Wish We'd Made Earlier
Sharding isn't just "split the table" — the shard key choice cascades through queries, joins, rebalancing, and operations. The decisions that pay off and the ones we redid.
Incident Post-Mortems That Drive Change (Not Theater)
Most post-mortems produce a document and no follow-through. The format, the discipline, and the cultural moves that actually convert incidents into engineering improvements.