Why blast radius is the most important setting you'll pick
Start small, expand with evidence. How to scope an experiment so it teaches you something without waking anyone up.
"How reliable are we?" is a terrible question. "How much unreliability can we still afford this month?" is one a business can actually answer.
That is what an error budget is: a spending limit on failure. If your SLO says a service should be available 99.9% of the time, the remaining 0.1% is a budget you are allowed to spend — on deploys, on migrations, and yes, on chaos experiments.
A raw availability number tells you nothing about what to do next. A budget does. When most of the budget is left, you ship faster and run bolder experiments. When it is nearly gone, you slow down, freeze risky changes, and protect what you have. Same signal, opposite actions — and every team, from engineering to finance, understands "we are running low."
Remaining budget tells you where you are; burn rate tells you how fast you are getting there. Burning budget slowly across a month is healthy. Burning a week's worth in an hour is an incident in progress. FaultMesh tracks burn rate over 1-hour, 6-hour and 24-hour windows, so a slow leak and a sudden gush look different at a glance — long before either becomes a breached SLO.
A budget nobody enforces is just a chart. FaultMesh makes it a control: the budget gate automatically blocks new experiments when a service drops below 20% remaining budget. You cannot go stress-test a service that is already one bad afternoon from breaching its SLO — the platform simply will not let you.
Better than reacting is anticipating. FaultMesh forecasts when a service is on track to exhaust its budget, using its recent burn rate to project the exhaustion date. That turns a reliability conversation into a planning one: "at this rate we run out in eleven days" is a sentence a CFO, a product lead and an SRE can all act on together.
Reliability stops being a vibe and becomes a number with a direction, a speed, and a deadline. That is a budget everyone in the building can reason about.
Start small, expand with evidence. How to scope an experiment so it teaches you something without waking anyone up.
Turn every post-mortem into a permanent regression test for your reliability.
Platform, control, observation, remediation and scoring — and how they close the loop.