Build systems that survive reality
Everything you need to run FaultMesh — from your first injected fault to a rising resilience score across every service.
Introduction
FaultMesh is an enterprise chaos engineering platform. It injects controlled failures into your systems, detects the resulting anomalies, remediates them automatically, and scores how resilient each service is over time.
It is built as a Kubernetes-native, modular platform with five clear control planes. Run it as a managed SaaS, or deploy the Kubernetes Operator inside your own cluster.
Quickstart
Authenticate, create an experiment, approve it, and execute — the full loop takes three calls against the API gateway.
# 1 · Authenticate and get a token curl -X POST https://api.faultmesh.com/auth/login \ -H "Content-Type: application/json" \ -d '{"email":"you@company.com","password":"••••••••"}' # 2 · Create your first experiment curl -X POST https://api.faultmesh.com/api/experiments \ -H "Authorization: Bearer $TOKEN" \ -d '{"name":"pod-kill checkout","type":"PodChaos", "targetService":"checkout","blastRadiusPercent":25}' # 3 · Approve, then execute curl -X POST https://api.faultmesh.com/api/experiments/$ID/approve -H "Authorization: Bearer $TOKEN" curl -X POST https://api.faultmesh.com/api/experiments/$ID/execute -H "Authorization: Bearer $TOKEN"
Prefer a UI? The FaultMesh console walks you through the same flow with a guided experiment wizard, live timeline and one-click abort.
Core concepts
A shared vocabulary keeps chaos engineering rigorous instead of reckless.
Experiment
A single controlled fault against a target service, with a defined type, blast radius and duration.
Hypothesis
What you expect to stay true while the fault is active — your steady-state assumption.
Blast radius
The percentage of a target's instances the fault is allowed to touch. Start small, expand with confidence.
Steady state
The measurable normal behaviour that must hold. If it breaks, the experiment surfaced a real weakness.
Resilience score
A composite, weighted score per service built from six dimensions, mapped to a maturity level.
Error budget
How much unreliability an SLO still permits. Burn rate tells you how fast you are spending it.
The five planes
FaultMesh is organised into five cooperating planes:
- Platform — gateway, authentication, licensing and notifications.
- Control — experiment orchestration and fault injection through Chaos Mesh.
- Observation — anomaly detection and live service topology.
- Remediation — rule-driven auto-healing and action execution.
- Scoring — SLO tracking, error budgets and resilience scoring.
Experiments
Experiments are the atomic unit of chaos. Each one injects a specific fault type into a target, powered by Chaos Mesh custom resources:
Every experiment supports a dry-run that counts live pods and validates the safety envelope before anything is touched.
Lifecycle
create → approve → execute → (abort). A timeline and an event stream record every state transition, so you can replay exactly what happened.
Game Days
A Game Day chains multiple experiments into a single orchestrated exercise, with optional inter-step validation that checks steady state between steps.
Progress and a timestamped, cursor-paginated log stream are persisted, so a Game Day can be followed live and audited afterwards.
Remediation
Remediation rules match on alerts and run an ordered action chain — restart, scale, failover or a custom executor — with a cooldown to prevent flapping.
Every rule can be gated by an OPA policy and leaves a full entry in the remediation audit log.
SLO & Scoring
Scoring turns raw signals into decisions leadership can act on:
- SLI — availability, p99 latency, error rate and throughput, evaluated over 1h to 30d windows.
- Error budget — total, consumed and remaining, with 1h / 6h / 24h burn rates.
- Maturity L1–L5 — a resilience maturity level derived from the composite score.
- Budget gate — automatically blocks new experiments when remaining budget drops below 20%.
Deployment
FaultMesh runs the same platform in two deployment modes:
Managed SaaS
A fully managed service with strict per-organisation data isolation. No infrastructure to run — sign in and go.
Private Cloud Operator
A native Kubernetes Operator reconciles every FaultMesh component as Deployments, Services and config inside your own cluster. Your data never leaves your perimeter.
Safety
Chaos without guardrails is just an outage. FaultMesh makes safety structural, not optional:
- Kill switch — one action aborts every active experiment across your organisation, instantly.
- Budget gate — experiments are refused when the error budget is already exhausted.
- Safety monitor — a background monitor enforces per-experiment timeouts and aborts on safety-envelope violations.
- RBAC & audit — Admin / Operator / Viewer roles, with every governance decision written to an immutable audit log.