From incident to experiment in one click
Turn every post-mortem into a permanent regression test for your reliability.
Every chaos experiment answers a question. The blast radius decides how expensive the wrong answer is.
Blast radius is the share of a target's instances a fault is allowed to touch. Set it to 100% and a bad hypothesis becomes a real outage. Set it to 5% and the same fault becomes a data point — one you can act on before your users ever notice.
The point of an experiment is to learn something you did not already know. A single failed pod in a fleet of forty tells you almost everything a full outage would: whether retries kick in, whether the load balancer sheds the dead instance, whether your dashboards even light up. It just tells you at a fraction of the risk.
So the first run of any new fault should be deliberately tiny. If steady state holds, you have earned the right to widen. If it breaks, you found the weakness with a scalpel instead of a sledgehammer.
Treat blast radius as a ladder you climb one rung at a time:
Each rung is a checkpoint. You climb only when the rung below held.
In FaultMesh every experiment declares its blast radius up front, and a dry-run counts the live pods it would touch before anything happens. If a run would exceed your safety envelope, the safety monitor aborts it. And if the error budget for a service is already spent, the budget gate refuses new experiments entirely — because the worst time to widen the blast radius is when you are already burning reliability.
Pick the smallest radius that can still answer your question. Then let the evidence, not your nerve, decide when to climb.
Turn every post-mortem into a permanent regression test for your reliability.
Translating burn rate and remaining budget into decisions the whole business can make.
Platform, control, observation, remediation and scoring — and how they close the loop.