MTTR — Mean Time to Recovery / Repair / Resolve
- MTTR measures the average time required to restore or resolve a service after a failure.
- Lower MTTR generally indicates faster incident recovery.
- Depending on the organization, MTTR may mean:
- Mean Time to Recovery
- Mean Time to Repair
- Mean Time to Resolve
MTBF — Mean Time Between Failures
- MTBF measures the average operating time between failures of a system.
- Higher MTBF generally indicates fewer failures.
- It is commonly used as a measure of system reliability.
MTTD — Mean Time to Detect
- MTTD measures the average time required to detect that an incident or failure has occurred.
- Lower MTTD means incidents are detected more quickly.
- Monitoring, alerting, and observability help reduce MTTD.
MTTA — Mean Time to Acknowledge
- MTTA measures the average time between an alert or incident being reported and someone acknowledging it.
- Lower MTTA means teams acknowledge incidents more quickly.
- It is commonly associated with incident response.
SLA — Service Level Agreement
- SLA is a formal agreement between a service provider and customer that defines expected service levels and commitments.
- An SLA may specify:
- Availability
- Response time
- Support commitments
- Resolution expectations
- Example: A provider may commit to 99.9% monthly availability.
SLO — Service Level Objective
- SLO defines a specific target for a service's reliability or performance.
- SLOs are typically used internally to guide engineering and operations.
- Example: 99.9% availability over a given period.
SLI — Service Level Indicator
- SLI is the actual measurement used to evaluate a service's performance or reliability.
- Common SLIs include:
- Availability
- Latency
- Error rate
- Throughput
SLI vs SLO vs SLA
- SLI → What do we measure?
- SLO → What target do we want?
- SLA → What do we promise the customer?
Example
- SLI: Actual availability = 99.95%
- SLO: Target availability = 99.9%
- SLA: Customer commitment = 99.9%
Incident Management Metrics
- MTTD → How quickly did we detect the incident?
- MTTA → How quickly did we acknowledge the incident?
- MTTR → How quickly did we recover or resolve the incident?
- MTBF → How long did the system operate before another failure?