Skip to main content
A single Availability Zone in AWS's Spain region degraded for roughly 2.5 hours, and most "multi-AZ" architectures would have failed anyway.

AWS EU-SOUTH-2 Packet Loss: The Single-AZ Lesson Nobody Learned

KU
Kiril Urbonas
6 days ago • 5 min read•0 views

A single Availability Zone in AWS's Spain region degraded for roughly 2.5 hours, and most "multi-AZ" architectures would have failed anyway.

Key takeaways

A single Availability Zone in AWS's Spain region degraded for roughly 2.5 hours, and most "multi-AZ" architectures would have failed anyway.

On October 4-5, 2026, AWS's EU-SOUTH-2 region (Spain) reportedly saw elevated packet loss confined to a single Availability Zone, eus2-az1, for roughly two and a half hours, with knock-on latency and error-rate increases for EC2, EKS, and DynamoDB according to third-party status trackers like IsDown. AWS has not published its own incident report for this one, which is itself the point: most single-AZ brownouts never get a postmortem, they just get quietly absorbed by whichever customers happened to have real multi-AZ failover and quietly eaten by everyone else.

"Multi-AZ" is a checkbox, not a guarantee#

Every team we've audited says they're multi-AZ. Most of them mean their RDS instance has a standby, or their ASG spans three subnets. That is multi-AZ in the billing sense, not multi-AZ in the failover sense. A brownout like this one doesn't take an AZ fully offline, it just makes it slow and flaky, which is the scenario almost nobody tests. Full outages fail loudly. Partial degradation fails quietly, and your systems keep sending traffic to the sick AZ because nothing told them to stop.

This is not the Bahrain story again#

We've written before about the Bahrain region outage and why multi-AZ isn't DR, and it's tempting to file this under the same headline. It isn't the same failure. Bahrain was effectively a full region event, where multi-AZ never had a chance because the whole region was compromised. EU-SOUTH-2 is narrower and more common: one AZ degrading while its siblings run fine. That's a harder failure mode to defend against, not an easier one, because the system looks healthy in aggregate dashboards while a third of your capacity is quietly timing out.

Your EKS node groups are probably not spread the way you think#

Pull the actual AZ distribution instead of trusting the Terraform variable name.

bash.bash
# Check which AZs your EKS node groups actually run in
$ aws eks describe-nodegroup \
    --cluster-name prod-cluster \
    --nodegroup-name prod-workers \
    --query 'nodegroup.subnets' --output text | \
  xargs -I{} aws ec2 describe-subnets --subnet-ids {} \
    --query 'Subnets[0].AvailabilityZone' --output text

If every result comes back eus2-az1, your "multi-AZ cluster" is a single-AZ cluster with a multi-AZ label. We see this constantly: the node group subnet list was copied once at cluster creation and never revisited as capacity grew in one AZ because it happened to have cheaper spot availability that week. Cluster Autoscaler and Karpenter both respect whatever subnet list you hand them, and neither one will rebalance existing pods across AZs after the fact unless you explicitly configure topology spread constraints. A cluster that started balanced in January can drift to lopsided by October without anyone changing a config file on purpose.

DynamoDB and SageMaker retries make the brownout worse, not better#

Default SDK retry behavior assumes failures are rare and independent. During an AZ-level brownout they are neither. A client that retries aggressively against a degraded AZ amplifies the packet loss into a self-inflicted thundering herd, and DynamoDB's adaptive capacity won't save you if your retry policy is fixed-interval instead of exponential with jitter.

python.python
# boto3 config: exponential backoff with jitter, capped retries
import boto3
from botocore.config import Config

config = Config(
    retries={
        "max_attempts": 5,
        "mode": "adaptive",  # throttles client-side during elevated error rates
    },
    connect_timeout=2,
    read_timeout=5,
)

dynamodb = boto3.client("dynamodb", config=config)

adaptive mode matters here specifically because it rate-limits the client when it detects a rising error rate, instead of hammering a sick AZ harder. Most teams are still on standard or legacy because nobody revisited the default after the SDK upgraded past it.

Your health checks probably can't see "degraded"#

Binary health checks answer "is it up," which is the wrong question during a brownout. Elevated packet loss and latency rarely trip a liveness probe. You need a check that measures tail latency or error rate against a threshold, not just a 200 response.

yaml.yaml
# kubernetes readiness probe tuned for degradation, not just downtime
readinessProbe:
  httpGet:
    path: /healthz/dependencies
    port: 8080
  periodSeconds: 5
  failureThreshold: 2
  timeoutSeconds: 2

Pair that with an endpoint that actually checks downstream latency against a budget, not just connectivity. A probe that only checks "can I open a socket" will happily keep routing traffic into an AZ that's technically reachable and functionally useless.

The decision, concretely#

  • Are your EKS node groups actually spread across AZs, or just configured to be? Verify with the subnet query above quarterly, not once at launch.
  • Does your retry policy assume independent, rare failures? If you're on boto3 legacy or standard retry mode, switch to adaptive before the next brownout, not during it.
  • Would your health checks catch "slow" or only "down"? If the answer is only "down," you have a monitoring gap, not a reliability guarantee.
  • Is your actual fix multi-AZ, or does the workload need multi-region? If the answer is the latter, read active-active versus active-passive before your next architecture review, because bolting on a second AZ won't solve a problem that needs a second region.

The call we'd make#

A 2.5-hour single-AZ brownout in a smaller AWS region is not newsworthy on its own, and that's exactly why it's worth your attention: these happen more often than the headline outages, they're harder to detect because nothing fully dies, and most "multi-AZ" setups we've reviewed would have taken a real latency hit from this one. Before your next incident review, go verify your actual AZ spread and your actual retry config, because the gap between "configured multi-AZ" and "resilient to AZ degradation" is where these incidents do their damage.

React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

567 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.