Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,100 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,055 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 933 views
Topics
Latest Articles
View All →Distroless Docker Images: Smaller, Safer Production Containers
Our node image shipped 240 CVEs, most from OS packages we never called. Moving to distroless dropped the count to single digits and cut image size by 70%.
Zero-Downtime Postgres Migrations: Expand-Contract in Practice
A single ALTER TABLE took a lock and stalled every write for 40 seconds during peak traffic. Expand-contract is how we stopped shipping outages.
Reranking in RAG: When a Cross-Encoder Earns Its Latency
Our RAG answers kept citing the wrong paragraph even when the right one was retrieved. A cross-encoder reranker fixed relevance but added 180ms. Here's when that trade pays off.
LLM Evals in CI: Catching Prompt Regressions Before They Ship
A prompt tweak that helped one case quietly broke twenty others. Here's the CI eval harness we built so that never ships silently again.
Debugging Latency with eBPF: bpftrace One-Liners That Find It
When a service is slow and every dashboard looks green, bpftrace lets you watch the kernel directly. These one-liners found our tail latency.
Multi-Arch Docker Builds with Buildx: One Image, Every Platform
Our M-series laptops built arm64, our CI built amd64, and prod pulled whichever tag won the race. Buildx and a manifest list ended the chaos.
AWS IAM Roles Anywhere: Workloads Outside AWS Without Static Keys
We had long-lived AWS keys sitting in a datacenter we don't own. IAM Roles Anywhere let us delete every one of them. Here's the real setup.
Hunting Slow Queries with pg_stat_statements
The dashboard said the database was fine. It wasn't. Here's how pg_stat_statements found the query eating 40% of our Postgres CPU.
Semantic Caching for LLM Apps: Cutting Cost on Repeated Queries
Users kept asking the same questions in slightly different words, and we paid full price every time. Semantic caching cut our LLM bill by a third.
Vercel Edge Middleware Patterns for Auth and A/B Testing
How we cut auth redirect latency to single-digit milliseconds and ran A/B tests without a flash of wrong content, using Vercel Edge Middleware.
Time-Series Postgres: Declarative Partitioning in Practice
A 900GB events table where every query scanned four years to read one day. Range partitioning by month cut our dashboard queries from 8s to under 200ms.
Argo Rollouts: Canary Analysis That Auto-Aborts on Bad Metrics
A bad deploy used to mean a pager at 2am and a manual rollback. Now Argo Rollouts watches the error rate and aborts the canary itself before anyone wakes up.