Skip to main content

Practical DevOps, Cloud, AI & Linux engineering guides

Featured Article

GitLab's New Rate Limits: What to Fix Before Oct 19

GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.

KU
Kiril UrbonasAI Engineer
|Oct 4, 2026
GitLab's New Rate Limits: What to Fix Before Oct 19

Most read

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05

Topics

Latest Articles

View All →
Standard APM doesn't tell you when your LLM-powered features are silently degrading. The signals we track and the dashboards that catch the regressions standard tools miss.
••September 3, 2025

AI Observability and Monitoring: Tracking Model Performance in Production

Standard APM doesn't tell you when your LLM-powered features are silently degrading. The signals we track and the dashboards that catch the regressions standard tools miss.

KU
Kiril Urbonas·3 min read·26
Read article
Evolve CI/CD toward autonomous pipelines that detect issues and roll back safely.
••September 1, 2025

Autonomous CI/CD Pipelines: Self-Healing and AI-Assisted Deployments

Evolve CI/CD toward autonomous pipelines that detect issues and roll back safely.

KU
Kiril Urbonas·1 min read·39
Read article
Multi-agent systems are mostly hype. The patterns we've seen actually deliver value, plus the ones we'd avoid until the tooling is more mature.
••August 31, 2025

Multi-Agent AI Systems: Building Collaborative AI Applications

Multi-agent systems are mostly hype. The patterns we've seen actually deliver value, plus the ones we'd avoid until the tooling is more mature.

KU
Kiril Urbonas·3 min read·52
Read article
We have ~40 prompts in production. The patterns that improved quality, the ones that turned out to be folklore, and how we test prompts now.
••August 27, 2025

Prompt Engineering Best Practices: Maximizing LLM Performance

We have ~40 prompts in production. The patterns that improved quality, the ones that turned out to be folklore, and how we test prompts now.

KU
Kiril Urbonas·3 min read·67
Read article
How we deploy LLM-powered features. The deployment patterns are mostly normal; the validation is where the differences are.
••August 23, 2025

AI Model Deployment Strategies: From Development to Production

How we deploy LLM-powered features. The deployment patterns are mostly normal; the validation is where the differences are.

KU
Kiril Urbonas·3 min read·29
Read article
We tried four quantization techniques on Llama-3 and Mistral models. The quality vs cost trade-offs we found, plus what works for production inference.
••August 20, 2025

Model Quantization Techniques: Reducing LLM Size and Cost

We tried four quantization techniques on Llama-3 and Mistral models. The quality vs cost trade-offs we found, plus what works for production inference.

KU
Kiril Urbonas·3 min read·41
Read article
We benchmarked four vector databases on the same workload. Each has a place. Here's how we'd pick today.
••August 16, 2025

Vector Databases for AI: Comparing Pinecone, Weaviate, and ChromaDB

We benchmarked four vector databases on the same workload. Each has a place. Here's how we'd pick today.

KU
Kiril Urbonas·3 min read·36
Read article
We've shipped four production RAG applications. Each one taught us something. The end-to-end pattern that works.
••August 13, 2025

Building RAG Applications: A Complete Guide to Retrieval Augmented Generation

We've shipped four production RAG applications. Each one taught us something. The end-to-end pattern that works.

KU
Kiril Urbonas·2 min read·48
Read article
Run retrieval-augmented generation at scale. Chunking, caching, and observability.
••August 12, 2025

RAG in Production: Reliability, Latency, and Cost for LLM Apps

Run retrieval-augmented generation at scale. Chunking, caching, and observability.

KU
Kiril Urbonas·1 min read·24
Read article
We cut LLM inference cost 47% over a quarter while improving p95 latency. Six changes, ranked by what each one actually delivered.
••August 9, 2025

Best Practices: AI Inference Cost Optimization

We cut LLM inference cost 47% over a quarter while improving p95 latency. Six changes, ranked by what each one actually delivered.

KU
Kiril Urbonas·2 min read·47
Read article
Wikis rot. We moved every operational doc into the repo it describes. Six months in, the docs are mostly correct because the only people who can update them are the ones who change the system.
••July 28, 2025

Best Practices: Infrastructure Documentation as Code

Wikis rot. We moved every operational doc into the repo it describes. Six months in, the docs are mostly correct because the only people who can update them are the ones who change the system.

KU
Kiril Urbonas·3 min read·36
Read article
A flat VPC is fine until you need to prove who can reach what. Five segmentation patterns that work in AWS without requiring a service mesh.
••July 24, 2025

Best Practices: Cloud Networking Segmentation Patterns

A flat VPC is fine until you need to prove who can reach what. Five segmentation patterns that work in AWS without requiring a service mesh.

KU
Kiril Urbonas·3 min read·35
Read article
Page 43 of 47 · 559 posts