Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Building visibility into cloud costs that actually drives action. The dashboards we look at, the alerts that fire, and the queries we run.
••10 months ago

Cloud Cost Monitoring: Tracking and Optimizing AWS Spending

Building visibility into cloud costs that actually drives action. The dashboards we look at, the alerts that fire, and the queries we run.

KU
Kiril Urbonas·3 min read·17
Read article
We run our app in two AWS regions for failover. The hard parts aren't the deployment — they're data consistency, traffic shifting, and the assumptions that break when "primary" is suddenly the wrong region.
••11 months ago

Multi-Region Deployment: Building Resilient Cloud Applications

We run our app in two AWS regions for failover. The hard parts aren't the deployment — they're data consistency, traffic shifting, and the assumptions that break when "primary" is suddenly the wrong region.

KU
Kiril Urbonas·3 min read·27
Read article
We run ~200 Lambda functions. Cold starts, memory tuning, and the cost-vs-latency trade-offs that actually move the bill.
••11 months ago

AWS Lambda Optimization: Reducing Costs and Improving Performance

We run ~200 Lambda functions. Cold starts, memory tuning, and the cost-vs-latency trade-offs that actually move the bill.

KU
Kiril Urbonas·3 min read·25
Read article
We track the four DORA metrics plus a handful of others. The trade-off between what's measurable and what's meaningful, and how we use the numbers.
••11 months ago

DevOps Metrics and KPIs: Measuring Success

We track the four DORA metrics plus a handful of others. The trade-off between what's measurable and what's meaningful, and how we use the numbers.

KU
Kiril Urbonas·3 min read·18
Read article
Design for region failure. Active/passive and active/active, data replication, and failover testing.
••11 months ago

Multi-Region Resilience: Failover, Data, and DNS

Design for region failure. Active/passive and active/active, data replication, and failover testing.

KU
Kiril Urbonas·1 min read·25
Read article
We've run canary deploys on most services for two years. The mechanics are easy; the metrics that decide "promote or roll back" are where the design is.
••11 months ago

Canary Releases: Gradual Rollout Strategy

We've run canary deploys on most services for two years. The mechanics are easy; the metrics that decide "promote or roll back" are where the design is.

KU
Kiril Urbonas·3 min read·27
Read article
We use blue-green for stateful services where canary doesn't fit. The actual mechanics, the data-layer subtleties, and when blue-green isn't the right answer.
••11 months ago

Blue-Green Deployments: Zero-Downtime Releases

We use blue-green for stateful services where canary doesn't fit. The actual mechanics, the data-layer subtleties, and when blue-green isn't the right answer.

KU
Kiril Urbonas·3 min read·171
Read article
We collect ~800GB of logs per day across our fleet. The shape of our logging stack, what we keep, what we drop, and what we'd build differently.
••11 months ago

Log Aggregation Strategies: Centralizing Your Logs

We collect ~800GB of logs per day across our fleet. The shape of our logging stack, what we keep, what we drop, and what we'd build differently.

KU
Kiril Urbonas·3 min read·34
Read article
A working Prometheus stack for a 40-node cluster: what we deploy, what we tune, and what we wish we'd known about cardinality two years ago.
••11 months ago

Infrastructure Monitoring with Prometheus: Complete Setup Guide

A working Prometheus stack for a 40-node cluster: what we deploy, what we tune, and what we wish we'd known about cardinality two years ago.

KU
Kiril Urbonas·3 min read·36
Read article
A focused look at the techniques that shrink container images: which actually pay off, which are folklore, and the discipline that keeps images small over time.
••11 months ago

Docker Multi-Stage Builds: Optimizing Image Size

A focused look at the techniques that shrink container images: which actually pay off, which are folklore, and the discipline that keeps images small over time.

KU
Kiril Urbonas·3 min read·40
Read article
We've had to restore a Kubernetes cluster from backup twice. Once it worked. Once it took 14 hours. Here's the strategy we run now.
••October 13, 2025

Kubernetes Backup Strategies: Protecting Your Cluster Data

We've had to restore a Kubernetes cluster from backup twice. Once it worked. Once it took 14 hours. Here's the strategy we run now.

KU
Kiril Urbonas·4 min read·21
Read article
Build MLOps pipelines for training, evaluation, and deployment. Reproducibility and monitoring.
••October 12, 2025

MLOps Pipelines: From Experiment to Production Models

Build MLOps pipelines for training, evaluation, and deployment. Reproducibility and monitoring.

KU
Kiril Urbonas·1 min read·10
Read article
Page 41 of 47 · 559 posts