learn

Key Terms

Explore this topic in Graph →

MTTR — Mean Time to Recovery / Repair / Resolve

  • MTTR measures the average time required to restore or resolve a service after a failure.
  • Lower MTTR generally indicates faster incident recovery.
  • Depending on the organization, MTTR may mean:
    • Mean Time to Recovery
    • Mean Time to Repair
    • Mean Time to Resolve

MTBF — Mean Time Between Failures

  • MTBF measures the average operating time between failures of a system.
  • Higher MTBF generally indicates fewer failures.
  • It is commonly used as a measure of system reliability.

MTTD — Mean Time to Detect

  • MTTD measures the average time required to detect that an incident or failure has occurred.
  • Lower MTTD means incidents are detected more quickly.
  • Monitoring, alerting, and observability help reduce MTTD.

MTTA — Mean Time to Acknowledge

  • MTTA measures the average time between an alert or incident being reported and someone acknowledging it.
  • Lower MTTA means teams acknowledge incidents more quickly.
  • It is commonly associated with incident response.

SLA — Service Level Agreement

  • SLA is a formal agreement between a service provider and customer that defines expected service levels and commitments.
  • An SLA may specify:
    • Availability
    • Response time
    • Support commitments
    • Resolution expectations
  • Example: A provider may commit to 99.9% monthly availability.

SLO — Service Level Objective

  • SLO defines a specific target for a service's reliability or performance.
  • SLOs are typically used internally to guide engineering and operations.
  • Example: 99.9% availability over a given period.

SLI — Service Level Indicator

  • SLI is the actual measurement used to evaluate a service's performance or reliability.
  • Common SLIs include:
    • Availability
    • Latency
    • Error rate
    • Throughput

SLI vs SLO vs SLA

  • SLI → What do we measure?
  • SLO → What target do we want?
  • SLA → What do we promise the customer?

Example

  • SLI: Actual availability = 99.95%
  • SLO: Target availability = 99.9%
  • SLA: Customer commitment = 99.9%

Incident Management Metrics

  • MTTD → How quickly did we detect the incident?
  • MTTA → How quickly did we acknowledge the incident?
  • MTTR → How quickly did we recover or resolve the incident?
  • MTBF → How long did the system operate before another failure?

Learning checkpoint

Mark this guide complete to include it in your local Engineering Journey.

Knowledge path

Connected concepts

Explore the knowledge graph →