Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.
••3 months ago

AWS Reserved Instances vs Savings Plans vs Spot — When Each Fits

Three discounting mechanisms, three different commitments. The rules of thumb we use to pick, and the mistakes we made before settling on them.

KU
Kiril Urbonas·3 min read·45
Read article
When the service is slow and the network is suspect, these are the tools we reach for, in this order, with the exact flags that find the answer.
••4 months ago

Linux Network Debugging — tcpdump, ss, and eBPF in Anger

When the service is slow and the network is suspect, these are the tools we reach for, in this order, with the exact flags that find the answer.

KU
Kiril Urbonas·3 min read·51
Read article
Token caching, model routing, prompt compression, and the boring discipline of measuring. The levers that cut our LLM bill 60% without touching feature scope.
••4 months ago

LLM Cost Optimization in Production — What Actually Moves the Bill

Token caching, model routing, prompt compression, and the boring discipline of measuring. The levers that cut our LLM bill 60% without touching feature scope.

KU
Kiril Urbonas·3 min read·13
Read article
pg_upgrade is fast but takes downtime; logical replication lets you cut over while the old DB still serves traffic. The runbook, the gotchas, and the post-cutover checklist.
••4 months ago

Postgres Logical Replication for Zero-Downtime Major Upgrades

pg_upgrade is fast but takes downtime; logical replication lets you cut over while the old DB still serves traffic. The runbook, the gotchas, and the post-cutover checklist.

KU
Kiril Urbonas·3 min read·28
Read article
Horizontal and vertical autoscalers solve different problems and break in different ways. The thresholds, cooldowns, and conflicts we learned the hard way.
••4 months ago

Kubernetes HPA and VPA — Tuning From Production Pain

Horizontal and vertical autoscalers solve different problems and break in different ways. The thresholds, cooldowns, and conflicts we learned the hard way.

KU
Kiril Urbonas·3 min read·18
Read article
Tracking experiments and shipping models are different problems. The MLOps tooling assumes one solution; production splits them. The patterns we use.
••4 months ago

MLOps — Model Registry vs MLflow Tracking, And When You Need Both

Tracking experiments and shipping models are different problems. The MLOps tooling assumes one solution; production splits them. The patterns we use.

KU
Kiril Urbonas·3 min read·31
Read article
Vault + Kubernetes auth + Vault Agent Injector. The setup, the failure modes during pod startup, and the patterns that beat raw Kubernetes Secrets.
••4 months ago

HashiCorp Vault as a Secrets Backend for Kubernetes

Vault + Kubernetes auth + Vault Agent Injector. The setup, the failure modes during pod startup, and the patterns that beat raw Kubernetes Secrets.

KU
Kiril Urbonas·3 min read·29
Read article
The single most useful Postgres extension you might not be using. The queries it surfaces, the indexes it implies, and the operational discipline of reading it weekly.
••4 months ago

pg_stat_statements — Postgres Query Analysis Without Guessing

The single most useful Postgres extension you might not be using. The queries it surfaces, the indexes it implies, and the operational discipline of reading it weekly.

KU
Kiril Urbonas·3 min read·32
Read article
io_uring replaces epoll for new high-throughput services. The patterns that earn their place, the gotchas in older kernels, and where we'd still pick epoll.
••4 months ago

Linux io_uring — Async I/O Patterns We Use

io_uring replaces epoll for new high-throughput services. The patterns that earn their place, the gotchas in older kernels, and where we'd still pick epoll.

KU
Kiril Urbonas·3 min read·68
Read article
Three caching patterns, three failure modes. The one we use most, the one that bit us, and the rule that decides which pattern fits which workload.
••4 months ago

Caching Patterns — Read-Through, Write-Through, Cache-Aside in Practice

Three caching patterns, three failure modes. The one we use most, the one that bit us, and the rule that decides which pattern fits which workload.

KU
Kiril Urbonas·3 min read·23
Read article
Picking partition counts and keys decides whether your Kafka consumers scale linearly or hit a wall. The patterns that survived rebalances, partition-count changes, and consumer-group ops.
••4 months ago

Kafka Partition Strategies — Scaling Consumers Without Reshuffling Everything

Picking partition counts and keys decides whether your Kafka consumers scale linearly or hit a wall. The patterns that survived rebalances, partition-count changes, and consumer-group ops.

KU
Kiril Urbonas·3 min read·35
Read article
AI agents for incident triage sound great in demos. We've tried it in production. The patterns that earn their keep, the ones that backfire, and where humans still beat agents.
••4 months ago

Agentic Ops — When (and When Not) to Use AI Agents for Incident Response

AI agents for incident triage sound great in demos. We've tried it in production. The patterns that earn their keep, the ones that backfire, and where humans still beat agents.

KU
Kiril Urbonas·3 min read·18
Read article
Page 29 of 47 · 559 posts