109 articles tagged with Monitoring.
We cut our average CI build time from 28 minutes to 6 minutes. The changes that mattered, ranked by impact.
We migrated 40+ services to GitOps with Argo CD. Two years in, here's what works and what required workarounds.
How a packet actually gets from the internet to a pod, walked layer by layer. Plus the things that surprise people the first time they hit them.
We've shipped three end-to-end ML systems. The pieces that look obvious in slides and turn out to be the actual work.
Standard APM doesn't tell you when your LLM-powered features are silently degrading. The signals we track and the dashboards that catch the regressions standard tools miss.
How we deploy LLM-powered features. The deployment patterns are mostly normal; the validation is where the differences are.
We cut LLM inference cost 47% over a quarter while improving p95 latency. Six changes, ranked by what each one actually delivered.
Blue/green sounds simple until your green cluster has a memory leak and you've already sent 50% of traffic there. The guardrails are what make it safe.
Manage cloud spend with Terraform: cost estimation, tagging, and policy-as-code.
We had four different patch cadences across our fleet and routinely missed CVEs by weeks. The unified workflow that finally caught up.
AWS bill grew 40% YoY for two years before we got serious. Tagging, scoped budgets, and a weekly review meeting did 80% of the work.
Unify traces, metrics, and logs with OpenTelemetry. Instrumentation, sampling, and backend-agnostic pipelines.