Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Our node image shipped 240 CVEs, most from OS packages we never called. Moving to distroless dropped the count to single digits and cut image size by 70%.
••3 months ago

Distroless Docker Images: Smaller, Safer Production Containers

Our node image shipped 240 CVEs, most from OS packages we never called. Moving to distroless dropped the count to single digits and cut image size by 70%.

KU
Kiril Urbonas·2 min read·10
Read article
A single ALTER TABLE took a lock and stalled every write for 40 seconds during peak traffic. Expand-contract is how we stopped shipping outages.
••3 months ago

Zero-Downtime Postgres Migrations: Expand-Contract in Practice

A single ALTER TABLE took a lock and stalled every write for 40 seconds during peak traffic. Expand-contract is how we stopped shipping outages.

KU
Kiril Urbonas·2 min read·23
Read article
Our RAG answers kept citing the wrong paragraph even when the right one was retrieved. A cross-encoder reranker fixed relevance but added 180ms. Here's when that trade pays off.
••3 months ago

Reranking in RAG: When a Cross-Encoder Earns Its Latency

Our RAG answers kept citing the wrong paragraph even when the right one was retrieved. A cross-encoder reranker fixed relevance but added 180ms. Here's when that trade pays off.

KU
Kiril Urbonas·2 min read·23
Read article
A prompt tweak that helped one case quietly broke twenty others. Here's the CI eval harness we built so that never ships silently again.
••3 months ago

LLM Evals in CI: Catching Prompt Regressions Before They Ship

A prompt tweak that helped one case quietly broke twenty others. Here's the CI eval harness we built so that never ships silently again.

KU
Kiril Urbonas·2 min read·23
Read article
When a service is slow and every dashboard looks green, bpftrace lets you watch the kernel directly. These one-liners found our tail latency.
••3 months ago

Debugging Latency with eBPF: bpftrace One-Liners That Find It

When a service is slow and every dashboard looks green, bpftrace lets you watch the kernel directly. These one-liners found our tail latency.

KU
Kiril Urbonas·2 min read·40
Read article
Our M-series laptops built arm64, our CI built amd64, and prod pulled whichever tag won the race. Buildx and a manifest list ended the chaos.
••3 months ago

Multi-Arch Docker Builds with Buildx: One Image, Every Platform

Our M-series laptops built arm64, our CI built amd64, and prod pulled whichever tag won the race. Buildx and a manifest list ended the chaos.

KU
Kiril Urbonas·2 min read·41
Read article
We had long-lived AWS keys sitting in a datacenter we don't own. IAM Roles Anywhere let us delete every one of them. Here's the real setup.
••3 months ago

AWS IAM Roles Anywhere: Workloads Outside AWS Without Static Keys

We had long-lived AWS keys sitting in a datacenter we don't own. IAM Roles Anywhere let us delete every one of them. Here's the real setup.

KU
Kiril Urbonas·2 min read·12
Read article
The dashboard said the database was fine. It wasn't. Here's how pg_stat_statements found the query eating 40% of our Postgres CPU.
••3 months ago

Hunting Slow Queries with pg_stat_statements

The dashboard said the database was fine. It wasn't. Here's how pg_stat_statements found the query eating 40% of our Postgres CPU.

KU
Kiril Urbonas·2 min read·22
Read article
Users kept asking the same questions in slightly different words, and we paid full price every time. Semantic caching cut our LLM bill by a third.
••3 months ago

Semantic Caching for LLM Apps: Cutting Cost on Repeated Queries

Users kept asking the same questions in slightly different words, and we paid full price every time. Semantic caching cut our LLM bill by a third.

KU
Kiril Urbonas·2 min read·24
Read article
How we cut auth redirect latency to single-digit milliseconds and ran A/B tests without a flash of wrong content, using Vercel Edge Middleware.
••3 months ago

Vercel Edge Middleware Patterns for Auth and A/B Testing

How we cut auth redirect latency to single-digit milliseconds and ran A/B tests without a flash of wrong content, using Vercel Edge Middleware.

KU
Kiril Urbonas·2 min read·37
Read article
A 900GB events table where every query scanned four years to read one day. Range partitioning by month cut our dashboard queries from 8s to under 200ms.
••3 months ago

Time-Series Postgres: Declarative Partitioning in Practice

A 900GB events table where every query scanned four years to read one day. Range partitioning by month cut our dashboard queries from 8s to under 200ms.

KU
Kiril Urbonas·1 min read·31
Read article
A bad deploy used to mean a pager at 2am and a manual rollback. Now Argo Rollouts watches the error rate and aborts the canary itself before anyone wakes up.
••3 months ago

Argo Rollouts: Canary Analysis That Auto-Aborts on Bad Metrics

A bad deploy used to mean a pager at 2am and a manual rollback. Now Argo Rollouts watches the error rate and aborts the canary itself before anyone wakes up.

KU
Kiril Urbonas·2 min read·26
Read article
Page 26 of 47 · 559 posts