Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Production monitoring catches user-facing issues. CI failures stay invisible until someone notices the merge queue is stuck. The metrics and alerts that make pipelines observable.
••4 months ago

Pipeline Observability — Why CI Failures Don't Trigger Alerts (And Should)

Production monitoring catches user-facing issues. CI failures stay invisible until someone notices the merge queue is stuck. The metrics and alerts that make pipelines observable.

KU
Kiril Urbonas·3 min read·27
Read article
Version-pinned modules across many repos. The release process, semver discipline, and the breaking-change communication that keeps a shared registry sane.
••4 months ago

Terraform Module Versioning and Shared Registries

Version-pinned modules across many repos. The release process, semver discipline, and the breaking-change communication that keeps a shared registry sane.

KU
Kiril Urbonas·3 min read·11
Read article
Most LLM eval suites correlate poorly with what real users experience. The eval patterns we run that move with prod metrics — and the ones that lied to us.
••4 months ago

LLM Evals That Actually Predict Production Quality

Most LLM eval suites correlate poorly with what real users experience. The eval patterns we run that move with prod metrics — and the ones that lied to us.

KU
Kiril Urbonas·3 min read·19
Read article
Static thresholds on error rate produce noisy alerts. Burn-rate alerting flips the question to "are we burning the error budget faster than we can sustain?" — and pages only on real problems.
••4 months ago

Burn-Rate Alerting — The SLO Discipline That Prevents Alert Fatigue

Static thresholds on error rate produce noisy alerts. Burn-rate alerting flips the question to "are we burning the error budget faster than we can sustain?" — and pages only on real problems.

KU
Kiril Urbonas·3 min read·35
Read article
cpu.shares vs cpu.cfs_quota_us vs memory.max — the cgroup mechanics behind Kubernetes resource limits, and the surprises that explain the weird symptoms you've seen.
••4 months ago

Container Resource Limits — What They Actually Do at the Kernel Level

cpu.shares vs cpu.cfs_quota_us vs memory.max — the cgroup mechanics behind Kubernetes resource limits, and the surprises that explain the weird symptoms you've seen.

KU
Kiril Urbonas·3 min read·31
Read article
Bad resource requests waste money or trigger OOMs. The methodology we use to right-size requests based on actual usage, and the gotchas the autoscalers don't fix.
••4 months ago

Kubernetes Resource Requests — Right-Sizing Without Guessing

Bad resource requests waste money or trigger OOMs. The methodology we use to right-size requests based on actual usage, and the gotchas the autoscalers don't fix.

KU
Kiril Urbonas·3 min read·26
Read article
SBOMs and signed attestations sound like checkboxes until you need to answer "did this artifact come from our pipeline?" The minimum viable supply-chain story we run.
••4 months ago

Supply Chain Security — SBOMs, Attestation, and What to Actually Verify

SBOMs and signed attestations sound like checkboxes until you need to answer "did this artifact come from our pipeline?" The minimum viable supply-chain story we run.

KU
Kiril Urbonas·3 min read·20
Read article
Edge compute is useless without an edge data layer. Three serverless databases that put data within ms of your edge functions, with the tradeoffs that aren't on the marketing pages.
••4 months ago

Edge Databases for Low-Latency Apps: D1, Turso, Neon Serverless

Edge compute is useless without an edge data layer. Three serverless databases that put data within ms of your edge functions, with the tradeoffs that aren't on the marketing pages.

KU
Kiril Urbonas·3 min read·109
Read article
Single-provider LLM apps fail when the provider does. Multi-provider routing isn't just resilience — it's also a cost lever. The patterns we run.
••4 months ago

Multi-Provider LLM Routing — Failover, Cost Routing, and Load Balancing

Single-provider LLM apps fail when the provider does. Multi-provider routing isn't just resilience — it's also a cost lever. The patterns we run.

KU
Kiril Urbonas·3 min read·23
Read article
EXPLAIN ANALYZE output is dense and intimidating. Once you can read it, most slow-query investigations finish in minutes. The patterns we keep seeing.
••4 months ago

Postgres Query Plans — Reading Them and the Indexes We Wish We'd Added Sooner

EXPLAIN ANALYZE output is dense and intimidating. Once you can read it, most slow-query investigations finish in minutes. The patterns we keep seeing.

KU
Kiril Urbonas·3 min read·23
Read article
Argo CD ships your manifests; Argo Rollouts ships them gradually with automated quality gates. The setup, the analysis templates that earn their place, and what we measure.
••4 months ago

Argo Rollouts — Progressive Delivery Beyond Argo CD

Argo CD ships your manifests; Argo Rollouts ships them gradually with automated quality gates. The setup, the analysis templates that earn their place, and what we measure.

KU
Kiril Urbonas·3 min read·25
Read article
bpftrace one-liners replace strace, perf top, and a half-dozen ad-hoc debugging scripts. The patterns that actually earn their place when you're troubleshooting at 2 AM.
••4 months ago

eBPF Tools for Everyday Ops — bpftrace Patterns We Use

bpftrace one-liners replace strace, perf top, and a half-dozen ad-hoc debugging scripts. The patterns that actually earn their place when you're troubleshooting at 2 AM.

KU
Kiril Urbonas·3 min read·31
Read article
Page 30 of 47 · 559 posts