Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,100 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,055 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 933 views
Topics
Latest Articles
View All →Time-Series Postgres: Declarative Partitioning in Practice
A 900GB events table where every query scanned four years to read one day. Range partitioning by month cut our dashboard queries from 8s to under 200ms.
Flux vs Argo CD: Picking a GitOps Engine in 2026
After running both in production across a dozen clusters, here's where Flux and Argo CD actually differ and which one we'd reach for now.
Argo CD ApplicationSets: Managing Many Clusters Without Copy-Paste
Twenty-three clusters, one app, and a folder of near-identical Application YAMLs that drifted constantly. ApplicationSets killed the copy-paste and the drift.
RAG Chunking Strategies: Fixed, Semantic, and Recursive Compared
Our support bot kept citing half a sentence and missing the answer that sat two lines below. The culprit wasn't the model, it was how we split the docs.
cgroup v2 Limits: What Actually Constrains Your Containers
A container with a 2-core limit was pegged at 100% CPU yet running slow. The throttling counter told the real story, and it wasn't the number we set.
Postgres Index Bloat: Detecting and Fixing It Before It Hurts
A 40GB index on a 6GB table was the first sign. Queries were fine until they weren't. Here's how we found the bloat and cleared it with zero downtime.
Turso vs Cloudflare D1: Choosing an Edge SQLite Database
Both put SQLite near your users, but they solve replication and write latency very differently. We ran the same schema on both for a month and picked one.
Cloudflare Workers vs AWS Lambda@Edge: Where Each Wins
We moved a rewrite-heavy request path off Lambda@Edge to Workers and cut p95 from 340ms to 41ms. Here's when that swap pays off and when it doesn't.
Cloud IAM Least-Privilege Without Breaking Everything
Least privilege fails when it's a one-time audit that locks things down until something breaks, then gets reverted. The iterative, log-driven approach that tightens permissions safely — and the policies we stopped writing by hand.
Prompt Caching for Production LLM Apps — Cutting Cost and Latency at the Token Layer
A long, stable system prompt re-billed on every request is money on fire. How prompt caching works, where the cache boundary belongs, and the structuring discipline that got us a big cost and latency cut without changing behavior.
Linux Memory Pressure — Reading PSI Before the OOM Killer Reads You
Free memory is a lie and load average doesn't see memory stalls. How Pressure Stall Information gives you a direct, early signal of memory contention — and how we wired it into alerts and autoscaling.
Kubernetes Pod Disruption Budgets — Surviving Node Drains Without an Outage
Node upgrades, autoscaler scale-downs, and spot reclaims all drain nodes. Without PDBs they can take all your replicas at once. The budgets, probes, and graceful-shutdown handling that keep voluntary disruptions invisible to users.