Every chaos experiment answers a question. The blast radius decides how expensive the wrong answer is.

Blast radius is the share of a target's instances a fault is allowed to touch. Set it to 100% and a bad hypothesis becomes a real outage. Set it to 5% and the same fault becomes a data point — one you can act on before your users ever notice.

Small is not timid — it's scientific

The point of an experiment is to learn something you did not already know. A single failed pod in a fleet of forty tells you almost everything a full outage would: whether retries kick in, whether the load balancer sheds the dead instance, whether your dashboards even light up. It just tells you at a fraction of the risk.

So the first run of any new fault should be deliberately tiny. If steady state holds, you have earned the right to widen. If it breaks, you found the weakness with a scalpel instead of a sledgehammer.

A ladder, not a leap

Treat blast radius as a ladder you climb one rung at a time:

  • 1 instance — prove the fault injects and recovers cleanly.
  • ~10% — confirm redundancy and automated remediation actually engage.
  • ~25% — stress the system enough to see latency and error-budget impact.
  • Higher — only for services with a proven track record, ideally during a Game Day.

Each rung is a checkpoint. You climb only when the rung below held.

How FaultMesh keeps you honest

In FaultMesh every experiment declares its blast radius up front, and a dry-run counts the live pods it would touch before anything happens. If a run would exceed your safety envelope, the safety monitor aborts it. And if the error budget for a service is already spent, the budget gate refuses new experiments entirely — because the worst time to widen the blast radius is when you are already burning reliability.

Pick the smallest radius that can still answer your question. Then let the evidence, not your nerve, decide when to climb.