Skip to main content
A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster. The lesson has nothing to do with AI.

Azure OpenAI's Sweden Central Outage: A Health Check Postmortem

KU
Kiril Urbonas
6 days ago • 6 min read•1 view

A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster. The lesson has nothing to do with AI.

Key takeaways

  • A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster.
  • The lesson has nothing to do with AI.

Between 10:03 and 15:58 UTC on September 29, 2026, Azure OpenAI Service, Foundry Agent Service, Foundry Models, and Cognitive Services went intermittent in Sweden Central, and Microsoft's own incident report describes a failure that has nothing to do with AI. A backend service timed out reading from its database and cache, got marked unhealthy, restarted, got marked unhealthy again, and an automated health check piled on more restarts until there weren't enough healthy instances left to serve requests. That is a restart storm, the same shape of outage that takes down ordinary web services every week.

What actually broke#

The affected component was a backend service that looks up service information and resource metadata on the request path for Sweden Central, not a model, not a GPU, not inference capacity. Per Microsoft's status history, that service started timing out reading from its dependent database and caching layers. Instances hit utilization thresholds, were marked unhealthy, and restarted repeatedly. An automated health check then triggered additional restarts on top of that, shrinking the healthy pool further. Customers saw intermittent request failures, increased latency, and HTTP 5XX errors on data-plane APIs in the region for close to six hours. Mitigation meant scaling out the backend service, raising resource limits, restarting infrastructure nodes, and disabling the health check.

That last detail is the whole story. The fix was not "make the database faster." It was "stop restarting instances."

The restart storm, mechanically#

A slow dependency and a dead dependency look identical to a health check that only measures response time. The database and cache were not down, they were degraded, and every instance depending on them took longer to answer. A fixed-timeout check read that as "not responding" and marked the instance unhealthy.

Restarting an unhealthy instance does not fix a slow downstream dependency, it just removes one more instance from the pool. The remaining instances carry more load each, which makes them slower too, which fails the same health check, which restarts more of them. Each restart cycle also adds reconnection and cache-warming load back onto the struggling database, so the dependency worsens at the exact moment the fleet needs it to recover. It self-reinforces until something outside the loop intervenes, which here meant an engineer manually disabling the health check.

The question every health check should answer, and usually doesn't#

Most health checks answer one question: did this instance respond within N milliseconds. That conflates two different states. A dead instance should be restarted. A slow-but-alive one needs load shed from it, not a restart, since restarting does nothing for the slowness and briefly makes capacity worse.

A check that distinguishes the two looks less like a ping and more like a circuit breaker:

python.python
# health_check.py: distinguish "slow" from "dead" before acting on it
import time

class DependencyHealth:
    def __init__(self, failure_threshold=5, reset_after_s=30):
        self.consecutive_timeouts = 0
        self.failure_threshold = failure_threshold
        self.reset_after_s = reset_after_s
        self.opened_at = None

    def record(self, latency_ms, timed_out):
        if timed_out:
            self.consecutive_timeouts += 1
        else:
            self.consecutive_timeouts = 0

    def should_shed_load(self):
        # Trips on a pattern, not one slow call, and recovers on a timer
        # rather than needing a human to flip it.
        if self.consecutive_timeouts >= self.failure_threshold:
            self.opened_at = self.opened_at or time.time()
            return True
        if self.opened_at and time.time() - self.opened_at < self.reset_after_s:
            return True
        self.opened_at = None
        return False

The instance stays in the pool, stops hitting the slow dependency for a cooldown window, and serves what it can from cache or a degraded response. Nobody restarts it. Restarts stay reserved for a process that's actually crashed or refusing connections, a narrower definition of "unhealthy" than "took longer than usual to answer."

Rate-limit your own remediation#

The second, independent mistake: nothing capped how many restarts could happen in a window. A health check with no rate limit on the action it triggers is a feedback loop waiting for a slow dependency to set it off:

yaml.yaml
# restart-policy.yaml: cap how fast automated remediation can act
restart_policy:
  max_restarts_per_window: 3
  window_seconds: 300
  on_limit_exceeded: page_oncall   # not: restart anyway
  min_healthy_fraction: 0.6        # never restart below this fleet-wide

min_healthy_fraction is the part teams skip. A restart policy that only looks at one instance at a time has no concept of "the fleet is already thin." Tie the restart decision to current fleet health, not just the instance in front of you, and the storm can't self-reinforce past a floor you chose on purpose.

What this means if you run production workloads against Azure OpenAI#

The lesson is not "leave Azure." A single-region data-plane dependency inherits whatever restart-storm risk that region's internal services carry, invisibly, until it's failing. The practical response is regional redundancy for the data plane: a secondary deployment behind health-aware routing, so a Sweden Central-specific incident degrades throughput rather than availability. We cover picking that pattern in multi-region active-active vs active-passive, and the routing layer that makes a second region useful for LLM traffic in multi-provider LLM gateways and fallback routing.

This is a different incident from the September 3 Azure outage we covered earlier, where a regional infrastructure failure took multiple LLM providers down together because they shared a cloud substrate. That one was about provider correlation; this one is about a health check mismanaging a degraded dependency inside a single provider's region. Same category of fix either way: don't assume one region.

The decision, concretely#

  • Does your health check distinguish slow-but-alive from actually-dead? If it only checks response time against a fixed timeout, no. Add a failure-count threshold before acting.
  • Is your restart policy rate-limited? Unlimited restarts in a short window turn a slow dependency into an outage. Cap restarts per window and require sign-off past that cap.
  • Do you have a circuit breaker on your database or cache dependency? If a slow read just blocks the request thread, shed load during a cooldown instead of queuing every instance behind the same slow call.
  • Are you running production inference against a single Azure region? A six-hour regional incident taking down your product means you need a second region in the routing path, not a second vendor.

The call we'd make#

Audit your own autohealing before you audit anyone else's. The Sweden Central incident is a clean, publicly documented case of a safety mechanism amplifying the failure it was built to catch, and that pattern shows up in internal systems more often than postmortems admit. If your health checks can restart an instance without checking fleet health first, fix that this week. Regional redundancy matters too, but it's the slower fix, and the restart-rate-limit is the one you can ship today.

Explore topics:Cloud
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

560 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.