Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
How we run OpenTelemetry across ~40 services. The instrumentation that earns its place, the patterns we abandoned, and what tracing actually catches that metrics don't.
••5 months ago

Distributed Tracing with OpenTelemetry — What We Ship, What We Skip

How we run OpenTelemetry across ~40 services. The instrumentation that earns its place, the patterns we abandoned, and what tracing actually catches that metrics don't.

KU
Kiril Urbonas·3 min read·25
Read article
A 2 AM incident, the autovacuum settings that caused it, and the parameter changes that prevented the next one. The discipline that took our biggest Postgres host from periodic stalls to steady.
••5 months ago

Postgres Autovacuum — Tuning From Production Stalls

A 2 AM incident, the autovacuum settings that caused it, and the parameter changes that prevented the next one. The discipline that took our biggest Postgres host from periodic stalls to steady.

KU
Kiril Urbonas·2 min read·17
Read article
Three layers of pooling, three different jobs. We learned the hard way which to use when. Real numbers from a 8k-connection workload.
••5 months ago

Database Connection Pooling at Scale: PgBouncer, RDS Proxy, Application Pool

Three layers of pooling, three different jobs. We learned the hard way which to use when. Real numbers from a 8k-connection workload.

KU
Kiril Urbonas·3 min read·26
Read article
We launched Backstage in October. Six months in, 80% of services are catalogued, on-boarding takes a third of the time, and we mostly know what owns what.
••5 months ago

Backstage Adoption: From Demo to 80% Service Coverage in 6 Months

We launched Backstage in October. Six months in, 80% of services are catalogued, on-boarding takes a third of the time, and we mostly know what owns what.

KU
Kiril Urbonas·3 min read·31
Read article
We deployed the same edge function on both platforms and measured for a quarter. Where each wins, where each loses, and the surprises along the way.
••5 months ago

Cloudflare Workers vs Vercel Edge: A Latency-Cost Comparison

We deployed the same edge function on both platforms and measured for a quarter. Where each wins, where each loses, and the surprises along the way.

KU
Kiril Urbonas·3 min read·207
Read article
We started using eBPF tooling for ad-hoc production debugging six months ago. Three real incidents where it cut investigation time from hours to minutes.
••5 months ago

eBPF for SREs: Three Real Diagnoses That Saved Hours

We started using eBPF tooling for ad-hoc production debugging six months ago. Three real incidents where it cut investigation time from hours to minutes.

KU
Kiril Urbonas·2 min read·27
Read article
We invalidate ~6% of LLM outputs before they reach a downstream system. Here's how we structure prompts and validators to catch malformed responses early.
••5 months ago

LLM Output Validation: Schema-First Prompt Engineering Patterns

We invalidate ~6% of LLM outputs before they reach a downstream system. Here's how we structure prompts and validators to catch malformed responses early.

KU
Kiril Urbonas·3 min read·51
Read article
A two-line config change to an Argo Rollouts analysis template caught a regression that would have cost ~$40k in API spend before we noticed. Here's the pattern.
••5 months ago

Argo Rollouts: Canary Deployments That Caught a $40k Bug

A two-line config change to an Argo Rollouts analysis template caught a regression that would have cost ~$40k in API spend before we noticed. Here's the pattern.

KU
Kiril Urbonas·3 min read·30
Read article
We ran Pulumi in TypeScript and Terraform in HCL side by side across 60+ services. Each won different categories of work. Here's the breakdown.
••5 months ago

Pulumi vs Terraform: What 18 Months of Production Taught Us

We ran Pulumi in TypeScript and Terraform in HCL side by side across 60+ services. Each won different categories of work. Here's the breakdown.

KU
Kiril Urbonas·3 min read·18
Read article
We deleted every static GCP service account key in our org over six weeks. Here's the migration plan, the gotchas, and the policies we now enforce.
••5 months ago

GCP Workload Identity Federation: Replacing Service Account Keys

We deleted every static GCP service account key in our org over six weeks. Here's the migration plan, the gotchas, and the policies we now enforce.

KU
Kiril Urbonas·2 min read·110
Read article
Three production OOM incidents that taught us how kubelet, containerd, and the kernel actually decide which process dies. With debugging commands you'll wish you had earlier.
••5 months ago

Linux Memory Management: When OOM Killer Strikes Your K8s Pods

Three production OOM incidents that taught us how kubelet, containerd, and the kernel actually decide which process dies. With debugging commands you'll wish you had earlier.

KU
Kiril Urbonas·2 min read·31
Read article
Bills hit $3,400/mo for runner minutes. We moved to self-hosted on EKS spot. The savings were real; the surprises were too.
••5 months ago

GitHub Actions Self-Hosted Runners: Why We Switched and What Broke

Bills hit $3,400/mo for runner minutes. We moved to self-hosted on EKS spot. The savings were real; the surprises were too.

KU
Kiril Urbonas·2 min read·17
Read article
Page 34 of 47 · 559 posts