Most incidents are paid for twice: once in the outage, and again in the post-mortem that gathers dust in a wiki.

The document is not the point. The point is making sure the same failure never surprises you again. And the only honest proof that a fix works is to reproduce the failure on demand — safely, on your terms.

The post-mortem loop that never closes

A typical retro ends with action items: add a retry, raise a limit, fix an alert. They get filed, half get done, and nobody ever re-runs the exact scenario to confirm the system now survives it. Six months later a similar incident lands, and the retro reads eerily familiar.

The missing step is a repeatable test that encodes the failure itself — not a checkbox, but an experiment you can run again next quarter.

Incident in, experiment out

FaultMesh closes that loop. When you file an incident, you record what actually broke: a network partition, a crashing pod, a latency spike, a CPU storm. From that record, FaultMesh can generate a matching experiment template in one click — mapping the root cause to the right fault type:

  • network partition → a network fault
  • pod crash → a pod-kill fault
  • high latency → an HTTP latency fault
  • CPU spike → a CPU-stress fault

Now the incident is not just a story — it is a runnable, versioned experiment that lives next to your others.

Regression tests for reliability

Fold that experiment into a Game Day and it becomes a regression test: every release, every quarter, you re-run the exact failure that once hurt you and confirm the system now shrugs it off. Your resilience score reflects the improvement, and your error budget shows the headroom you have earned.

An incident you can reproduce on demand is an incident that has lost its power to surprise you.