Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Container performance problems usually live in the node kernel, not your app. Here is what we tune, why, and how we measure before touching anything.
••July 22, 2025

Linux Performance Tuning for Containers and Kubernetes Nodes

Container performance problems usually live in the node kernel, not your app. Here is what we tune, why, and how we measure before touching anything.

KU
Kiril Urbonas·3 min read·78
Read article
Blue/green sounds simple until your green cluster has a memory leak and you've already sent 50% of traffic there. The guardrails are what make it safe.
••July 17, 2025

Best Practices: Blue-Green Deployment Guardrails

Blue/green sounds simple until your green cluster has a memory leak and you've already sent 50% of traffic there. The guardrails are what make it safe.

KU
Kiril Urbonas·2 min read·26
Read article
Manage cloud spend with Terraform: cost estimation, tagging, and policy-as-code.
••July 1, 2025

Terraform Cloud Cost Controls: Budgets, Policies, and Tagging

Manage cloud spend with Terraform: cost estimation, tagging, and policy-as-code.

KU
Kiril Urbonas·1 min read·36
Read article
We had four different patch cadences across our fleet and routinely missed CVEs by weeks. The unified workflow that finally caught up.
••June 12, 2025

Best Practices: Kernel and Package Patch Management

We had four different patch cadences across our fleet and routinely missed CVEs by weeks. The unified workflow that finally caught up.

KU
Kiril Urbonas·3 min read·23
Read article
Harden container images and runtime. Image scanning, minimal base, and supply chain security.
••June 11, 2025

Docker Security Best Practices: Images, Runtime, and Supply Chain

Harden container images and runtime. Image scanning, minimal base, and supply chain security.

KU
Kiril Urbonas·1 min read·34
Read article
AWS bill grew 40% YoY for two years before we got serious. Tagging, scoped budgets, and a weekly review meeting did 80% of the work.
••May 26, 2025

Best Practices: AWS Cost Control with Tagging and Budgets

AWS bill grew 40% YoY for two years before we got serious. Tagging, scoped budgets, and a weekly review meeting did 80% of the work.

KU
Kiril Urbonas·3 min read·35
Read article
A team of 30 engineers all editing the same monolithic Ansible repo doesn't work. Here's the role taxonomy and review process that did.
••May 23, 2025

Best Practices: Ansible Role Design for Large Teams

A team of 30 engineers all editing the same monolithic Ansible repo doesn't work. Here's the role taxonomy and review process that did.

KU
Kiril Urbonas·4 min read·29
Read article
Unify traces, metrics, and logs with OpenTelemetry. Instrumentation, sampling, and backend-agnostic pipelines.
••May 21, 2025

Observability with OpenTelemetry: Traces, Metrics, and Logs

Unify traces, metrics, and logs with OpenTelemetry. Instrumentation, sampling, and backend-agnostic pipelines.

KU
Kiril Urbonas·1 min read·28
Read article
One Terraform state file per environment sounds obvious until you watch a dev plan touch a prod resource. Here's how we actually isolate state and the mistakes we made getting there.
••May 19, 2025

Best Practices: Terraform State Isolation by Environment

One Terraform state file per environment sounds obvious until you watch a dev plan touch a prod resource. Here's how we actually isolate state and the mistakes we made getting there.

KU
Kiril Urbonas·2 min read·17
Read article
Our CI was 73% green at the worst point. People trusted it less than coin flips. Six things we did to get to 96%, in rough order of impact.
••May 15, 2025

Best Practices: GitHub Actions Pipeline Reliability

Our CI was 73% green at the worst point. People trusted it less than coin flips. Six things we did to get to 96%, in rough order of impact.

KU
Kiril Urbonas·3 min read·34
Read article
Our base image went from 1.2 GB and 200+ CVEs to 80 MB and 4 CVEs. Most of the work wasn't clever — it was deletion.
••May 11, 2025

Best Practices: Docker Image Hardening for Production

Our base image went from 1.2 GB and 200+ CVEs to 80 MB and 4 CVEs. Most of the work wasn't clever — it was deletion.

KU
Kiril Urbonas·2 min read·31
Read article
We upgraded a 60-node EKS cluster from 1.27 to 1.31 over six months. Four minor versions, one bad surprise, zero customer impact. Here's the playbook.
••May 7, 2025

Best Practices: Kubernetes Cluster Upgrade Strategy

We upgraded a 60-node EKS cluster from 1.27 to 1.31 over six months. Four minor versions, one bad surprise, zero customer impact. Here's the playbook.

KU
Kiril Urbonas·2 min read·38
Read article
Page 44 of 47 · 559 posts