Practical DevOps, Cloud, AI & Linux engineering guides
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Most read
- 01How to Reduce Datadog Costs Without Losing CoverageCloud · 2,102 views
- 02OpenTelemetry Collector Pipelines: Real Configs That Survived ProductionDevOps · 2,045 views
- 03Azure DevOps Best Practices in 2026: Build Pipelines You Can TrustDevOps · 1,059 views
- 04
- 05A Pragmatic Multi-Region Strategy for Small TeamsCloud · 939 views
Topics
Latest Articles
View All →Ansible Playbook Optimization: Writing Efficient Playbooks
We cut our largest playbook's runtime from 14 minutes to 4 minutes. The specific changes that mattered, plus the ones that didn't.
Pulumi vs Terraform Deep Dive: Choosing the Right IaC Tool
We tried Pulumi for a quarter and went back to Terraform. Both are real options. Why we picked one and what would change our mind.
Operational Checklist: Kubernetes Secrets and External Vault Integration
K8s Secrets are barely encrypted. We moved every secret to Vault with the Vault Agent injector and never went back. The setup checklist.
Infrastructure Testing Strategies: Validating Your IaC
We test infrastructure code with three layers: validation, plan review, and integration tests. The setup that catches real bugs without slowing down PRs.
Terraform Modules Best Practices: Building Reusable Infrastructure
We have a private module registry with ~25 modules used across 12 accounts. Versioning, interface design, and the over-modularization mistake we keep making.
Linux Container Internals: Understanding How Containers Work
A container is a process with extra kernel features applied. Walking through namespaces, cgroups, and the actual mechanics — the level of detail that makes "container weirdness" debuggable.
Shell Scripting Best Practices: Writing Maintainable Scripts
We have a few hundred shell scripts in production. The patterns that make them survive contact with reality, and the ones we've stopped writing.
Prompt Engineering for DevOps: Consistency and Safety
Use prompts to get reliable, safe outputs from LLMs for runbooks, code, and ops tasks.
File System Optimization: Improving Disk Performance
Filesystem choice, mount options, IO schedulers — the per-host tweaks that actually moved disk performance for our database and storage workloads.
Process Management and Monitoring in Linux
How processes actually live and die on Linux, the tools that show what's happening, and the patterns we use for monitoring service health.
Linux Security Hardening: Protecting Your System
A practical Linux hardening checklist for production hosts. The settings that earn their place via real production reasons, not the cargo-cult version.
Operational Checklist: Systemd Service Reliability Patterns
A condensed checklist of the systemd unit-file patterns we now use everywhere, with the production reasons each one matters.