Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
We ran the same RAG workload across three vector stores for a quarter each. Here's what we learned about latency, cost, and operational overhead.
••5 months ago

Vector Database Selection: Pinecone, pgvector, Qdrant After 6 Months in Production

We ran the same RAG workload across three vector stores for a quarter each. Here's what we learned about latency, cost, and operational overhead.

KU
Kiril Urbonas·2 min read·26
Read article
Every hook on this list caught a bug or a security issue in the last twelve months. The configs are short. The savings have been considerable.
••5 months ago

Pre-Commit Hooks That Saved Our Repo: 7 Real Examples

Every hook on this list caught a bug or a security issue in the last twelve months. The configs are short. The savings have been considerable.

KU
Kiril Urbonas·3 min read·41
Read article
We moved a 60-node production EKS cluster to Auto Mode. Some pain points evaporated, others got harder. The cost picture is more nuanced than the marketing suggests.
••6 months ago

EKS Auto Mode: What Worked, What Broke in Our Migration

We moved a 60-node production EKS cluster to Auto Mode. Some pain points evaporated, others got harder. The cost picture is more nuanced than the marketing suggests.

KU
Kiril Urbonas·3 min read·29
Read article
We ran the same workload on both for half a year. The break-even point isn't where most blog posts say it is — and the latency story has more nuance than throughput-per-dollar charts admit.
••6 months ago

Self-Hosted LLMs vs OpenAI API: A Cost-vs-Latency Analysis After 6 Months

We ran the same workload on both for half a year. The break-even point isn't where most blog posts say it is — and the latency story has more nuance than throughput-per-dollar charts admit.

KU
Kiril Urbonas·2 min read·68
Read article
We've been running the OTel Collector at the edge of every cluster for 18 months. The config patterns that lasted, the ones we ripped out, and a few processors that quietly saved us money.
••6 months ago

OpenTelemetry Collector Pipelines: Real Configs That Survived Production

We've been running the OTel Collector at the edge of every cluster for 18 months. The config patterns that lasted, the ones we ripped out, and a few processors that quietly saved us money.

KU
Kiril Urbonas·2 min read·2K
Read article
Blue/green is easy for stateless services. We did it for our primary Postgres cluster with 3.2TB of data and ~8k connections. Here's exactly how — and what almost went wrong.
••6 months ago

Blue/Green Deploys for Stateful Services: A Postgres Cutover Story

Blue/green is easy for stateless services. We did it for our primary Postgres cluster with 3.2TB of data and ~8k connections. Here's exactly how — and what almost went wrong.

KU
Kiril Urbonas·2 min read·20
Read article
We replaced 14 long-lived IAM users with SSO + temporary credentials. The migration plan, the gotchas, and the policies we now enforce.
••6 months ago

Zero Trust on AWS: Lessons From Implementing IAM Identity Center

We replaced 14 long-lived IAM users with SSO + temporary credentials. The migration plan, the gotchas, and the policies we now enforce.

KU
Kiril Urbonas·2 min read·22
Read article
Six months running RAG in production taught us that the retrieval step matters far more than the model. Concrete techniques that moved the needle, with before/after numbers.
••6 months ago

Embedding Quality in RAG: How We Cut Hallucinations by 60%

Six months running RAG in production taught us that the retrieval step matters far more than the model. Concrete techniques that moved the needle, with before/after numbers.

KU
Kiril Urbonas·2 min read·41
Read article
How we shipped three schema migrations with zero customer impact. Expand-then-contract, dual-writes, and the rollback plan we never had to use — but tested anyway.
••6 months ago

Database Migrations Without Downtime: Patterns From Three Real Cutovers

How we shipped three schema migrations with zero customer impact. Expand-then-contract, dual-writes, and the rollback plan we never had to use — but tested anyway.

KU
Kiril Urbonas·2 min read·27
Read article
We were drowning in 200 alerts a week. Most got ignored. After a quarter of triage and rework, we're at about 15 — and on-call actually responds to them.
••6 months ago

Monitoring That Actually Helps On-Call: Alerts, Dashboards, and Runbooks

We were drowning in 200 alerts a week. Most got ignored. After a quarter of triage and rework, we're at about 15 — and on-call actually responds to them.

KU
Kiril Urbonas·2 min read·30
Read article
We had .env files in three repos, AWS keys in Slack DMs, and a postgres password etched into a Confluence page. Cleaning it up took a sprint and changed how we think about secrets.
••6 months ago

Secrets Management in Practice: From .env Files to Vault

We had .env files in three repos, AWS keys in Slack DMs, and a postgres password etched into a Confluence page. Cleaning it up took a sprint and changed how we think about secrets.

KU
Kiril Urbonas·2 min read·25
Read article
We wrote pretty postmortems for two years and kept hitting the same incidents. Here's what changed when we started writing ugly ones.
••6 months ago

Incident Postmortems That Actually Prevent Repeat Failures

We wrote pretty postmortems for two years and kept hitting the same incidents. Here's what changed when we started writing ugly ones.

KU
Kiril Urbonas·3 min read·20
Read article
Page 35 of 47 · 559 posts