Hope Is Not a Resilience Strategy
Every architecture diagram has a failover arrow on it. Far fewer teams have watched that arrow work under real traffic. Multi-zone databases, retry logic, circuit breakers and autoscaling policies are all bets on how a system will behave when something breaks, and most of those bets are never tested until an actual outage settles them at the worst possible moment.
Chaos engineering is the practice of settling those bets on your own schedule. The Principles of Chaos Engineering describe it as running controlled experiments on a system in order to build confidence in its ability to withstand turbulent conditions. Netflix made the idea famous with Chaos Monkey, which randomly terminates production instances so that engineers have no choice but to design services that survive instance loss. The idea has since matured from a single blunt tool into a discipline, and SLOs are what make that discipline safe to practise.
Your SLI Is the Steady State
The principles start with a steady state: a measurable output that indicates the system is behaving normally. Teams new to chaos engineering often reach for infrastructure signals such as CPU or pod counts, but those describe the system rather than the user. The better answer is already sitting in your SLO definitions. If checkout availability and p95 latency are the indicators your users care about, they are the indicators an experiment should watch.
That turns a vague intention into a falsifiable hypothesis: terminating one replica of the checkout service will keep availability above 99.9% and p95 latency under 400 milliseconds for the duration of the test. Either the hypothesis holds and you have earned real confidence, or it fails and you have found a weakness before a customer did. If your indicators do not yet map to user journeys, work through critical user journeys, SLIs and SLOs first; experiments against a meaningless metric teach you nothing.
Size the Blast Radius With the Error Budget
An experiment that runs in production spends error budget, and it should. The budget exists precisely to fund calculated risks, and a chaos experiment is the most calculated risk there is. The remaining budget also gives you a principled answer to the question of how much chaos is acceptable this month. A service with 70% of its budget left can afford an availability-zone failure test. A service that burned through its budget last week should be fixing the causes, not adding new ones. The error budget calculator converts a target into the minutes you actually have to spend.
Start with the smallest blast radius that can still disprove the hypothesis: one pod, then one node, then one zone, then a dependency. Each step only proceeds if the previous one held.
Abort Conditions Come Free With Burn Rate Alerts
The scariest part of running experiments in production is the moment things go worse than expected. Good tooling handles this with automatic halts, and the trigger should be the same signal you already page on. AWS Fault Injection Service ties experiments to stop conditions backed by CloudWatch alarms, so an alarm on your SLO burn rate rolls the fault back without a human in the loop. LitmusChaos and Chaos Mesh, both CNCF incubating projects, support probes and status checks that fail an experiment when a metric crosses a threshold. If you have already built multi-window burn rate alerts, you have the abort condition ready to wire in.
Pick the Tool That Matches Your Platform
- Kubernetes: LitmusChaos offers a hub of reusable experiments and chains them into workflows; Chaos Mesh defines faults as Kubernetes custom resources with a dashboard and workflow engine, and injects network delay, packet loss, pod kills and CPU or memory stress.
- AWS: Fault Injection Service runs managed experiments against EC2, EKS, RDS, Lambda and network paths from a template that declares targets, actions and stop conditions.
- Azure: Azure Chaos Studio offers service-direct faults such as forced database failovers alongside agent-based faults inside virtual machines.
- Any environment: Gremlin is a commercial platform with a library of scenarios and a halt button that reverts every attack immediately.
From Single Experiments to Game Days
Individual experiments verify components. Game days verify people. Google's annual disaster recovery testing programme, described in the ACM Queue article Weathering the Unexpected, deliberately breaks systems and business processes across the company so that responders, runbooks and escalation paths are exercised as well as code. You do not need Google's scale to copy the format: pick one scenario a quarter, tell the on-call engineer nothing beyond the time window, and run it as a real incident with a real incident management process. Then write the post-incident review as if it had been unplanned. The failure you rehearsed on a Tuesday afternoon is the one that will not page you at 3 a.m.