Why blast radius is the most important setting you'll pick
Start small, expand with evidence. How to scope an experiment so it teaches you something without waking anyone up.
Injecting a fault is the easy part. Turning it into a lasting reliability gain takes a whole loop — and FaultMesh organises that loop into five cooperating planes.
The foundation: the gateway that routes and rate-limits every request, authentication and access control, licensing, and notifications. It is the plane you rarely think about and would immediately miss — the connective tissue that lets everything else stay isolated, authorised and observable.
Where chaos actually happens. The control plane orchestrates experiments and Game Days and injects faults through Chaos Mesh — latency, packet loss, pod kills, CPU and memory stress and more. It owns the experiment lifecycle: create, approve, execute, abort, each transition recorded on a timeline you can replay.
A fault you cannot see is just an outage. The observation plane provides real-time anomaly detection and a live map of your service topology, so when something breaks you know exactly what broke and how far the damage could spread. It is how you see failure the way your users would — and catch it first.
Detecting a problem is not the same as fixing it. The remediation plane runs rule-driven auto-healing: match on an alert, execute an ordered chain of actions — restart, scale, failover — gated by policy and recorded in an audit log. And when automation is not enough, the kill switch aborts every active experiment across your organisation, instantly.
The plane that makes progress undeniable. Scoring tracks SLOs and error budgets and rolls six dimensions into a single resilience score per service, mapped to a maturity level from L1 to L5. It is what turns a pile of experiments into a trend line leadership can act on.
Each plane hands off to the next: Control injects, Observation watches, Remediation heals, Scoring measures, and Platform holds it all together. Run that loop often enough and reliability stops being an accident — it becomes something you engineer on purpose.
Start small, expand with evidence. How to scope an experiment so it teaches you something without waking anyone up.
Turn every post-mortem into a permanent regression test for your reliability.
Translating burn rate and remaining budget into decisions the whole business can make.