SLA vs SLO vs SLI: Definitions, Error Budgets and Examples

SLA, SLO and SLI explained in plain English: how they differ, how to calculate an error budget with a worked example, and how to measure reliability.

Updated 8 min readBy the Uptime Tracker team

Short answer

An SLI (service level indicator) is a measurement, such as the percentage of successful requests. An SLO (service level objective) is an internal target for that measurement, such as 99.9% over 30 days. An SLA (service level agreement) is a contract with customers that promises a level of service and defines consequences, typically service credits, if it is missed. The error budget is the amount of failure the SLO allows: 0.1% of a 30-day window is 43 minutes 12 seconds.

SLI, SLO and SLA are three layers of the same idea: measure reliability, set a target, and, optionally, promise it to customers. The terms were popularized by Google's Site Reliability Engineering books and are now standard across the industry.

Definitions

TermWhat it isExampleAudience
SLI (indicator)A quantitative measurement of service behaviorShare of HTTP requests that return a non-5xx response in under 1 secondEngineering
SLO (objective)A target value for an SLI over a time window99.9% of requests succeed over a rolling 30 daysEngineering and product
SLA (agreement)A contractual promise with consequences99.5% monthly uptime, or a 10% service creditCustomers, legal, sales

The relationship is: you measure SLIs, aim for SLOs, and promise SLAs. The SLA should always be looser than the SLO, so you have a safety margin before contractual penalties apply.

Service level indicators (SLIs)

A good SLI measures something users actually feel. It is usually expressed as a ratio of good events to total events, which gives a number between 0% and 100%. Common SLIs:

  • Availability: successful requests divided by total requests, or successful checks divided by total checks.
  • Latency: share of requests served faster than a threshold, such as 300 ms.
  • Freshness: share of time data is no older than a threshold, useful for pipelines and dashboards.
  • Correctness: share of responses that pass a validation check.

Avoid SLIs based on server metrics like CPU usage. They describe the machine, not the user experience.

Service level objectives (SLOs)

An SLO combines an SLI, a target and a window: "99.9% of checkout requests succeed, measured over a rolling 30 days." Guidelines for setting one:

  1. Measure the SLI for several weeks before choosing a target.
  2. Choose a target users would notice if you missed it, not 100%.
  3. Prefer rolling windows (last 30 days) over calendar months, so a bad day at month-end does not reset artificially.
  4. Keep the number of SLOs small, a few per service.
  5. Review targets quarterly.

Service level agreements (SLAs)

An SLA is a legal and commercial document. It defines the metric, how it is measured, what is excluded (scheduled maintenance, customer-caused issues, force majeure), how customers claim compensation, and what the compensation is. Service credits, a percentage of the monthly fee, are the most common remedy. Because an SLA carries financial consequences, it is usually set below the internal SLO.

Error budgets

An error budget is the amount of unreliability an SLO permits: 100% minus the SLO. It turns reliability into a resource the team can spend. If the budget is healthy, the team can ship faster and take more risk. If it is nearly exhausted, the team slows down releases and focuses on stability.

Worked example: time-based

SLO: 99.9% availability over a rolling 30-day window.

Window:        30 days x 24 h x 60 min = 43,200 minutes
Error budget:  0.1% x 43,200 min       = 43.2 minutes (43 min 12 s)

Incidents this window:
  Deploy rollback    12 min
  Database failover   9 min
  Total              21 min

Budget consumed:   21 / 43.2  = 48.6%
Budget remaining:  43.2 - 21  = 22.2 minutes

Half the budget is gone with time still left in the window. A reasonable policy is to keep shipping but require extra review for risky changes until the window rolls forward.

Worked example: request-based

SLO: 99.9% of API requests succeed over 30 days. The service handles 20,000,000 requests in that period.

Error budget:   0.1% x 20,000,000 = 20,000 failed requests
Failed so far:  14,500
Remaining:       5,500 (27.5% of the budget)

Request-based budgets are fairer for services with uneven traffic: an outage at 3 a.m. with little traffic consumes less budget than one at peak.

Burn rate

Burn rate is how fast you are consuming the error budget relative to the rate that would exactly exhaust it at the end of the window. A burn rate of 1 means you will use the whole budget by the end of the window. A burn rate of 10 means a 30-day budget would be gone in 3 days. Alerting on high burn rates catches serious problems quickly while ignoring small, slow consumption.

How to measure in practice

  • External uptime checks measure availability as users see it from outside, including DNS, TLS and network problems that server logs miss. Check interval limits precision: a 1-minute interval cannot see a 20-second outage.
  • Server or load balancer logs give request-based SLIs with full traffic coverage, but miss failures that happen before requests reach you.
  • Real-user monitoring captures actual user experience in browsers and apps.

Many teams combine external checks for availability with logs for request success and latency. Uptime Tracker measures time-weighted availability from external checks, with outages counted from the first failed check and maintenance windows excluded. The Team plan adds SLOs with error budget tracking. See the nines table for the downtime each target allows, or use the uptime calculator.

Putting it together: one service, three layers

Here is how the three terms fit together for a hypothetical public API:

LayerDefinition for this API
SLIShare of requests to the public API that return a non-5xx response, measured at the load balancer, plus share of 1-minute external checks that succeed
SLO99.95% of requests succeed over a rolling 30 days (internal error budget: 21 minutes 36 seconds of full outage equivalent)
SLA99.9% monthly availability, excluding announced maintenance; 10% service credit below 99.9%, 25% below 99%

The gap between the 99.95% SLO and the 99.9% SLA gives the team about 21 minutes per month of warning before customers are owed credits. When the error budget is half spent, the team reviews risky releases; when it is gone, feature releases pause.

Common mistakes

  • Setting the SLA equal to the SLO, leaving no margin.
  • Targeting 100%, which leaves no room for change and is unachievable.
  • Measuring only the homepage while customers depend on the API.
  • Defining many SLOs that nobody reviews.
  • Not agreeing in advance what happens when the error budget runs out.
FAQ

Frequently asked questions

What is the difference between an SLA and an SLO?

An SLO is an internal reliability target, such as 99.9% availability over 30 days. An SLA is a contract with customers that promises a level of service and specifies compensation, usually service credits, if it is missed. SLAs are normally set looser than SLOs to leave a safety margin.

What is an SLI example?

A typical SLI is the percentage of HTTP requests that return a successful (non-5xx) response within 1 second. Other examples are the percentage of successful uptime checks, the share of requests faster than 300 ms, and the share of time a data pipeline is up to date.

How do you calculate an error budget?

Subtract the SLO from 100% and apply it to the window. For a 99.9% SLO over 30 days, the error budget is 0.1% of 43,200 minutes, or 43.2 minutes. For a request-based SLO with 20 million requests, it is 20,000 failed requests.

What happens when an error budget is exhausted?

That depends on the policy the team agreed in advance. Common responses are freezing non-essential releases, prioritizing reliability work, and requiring extra review for changes until the budget recovers as the rolling window moves forward.

Is 99.9% a good SLO?

99.9% is a common SLO for SaaS products and business websites. It allows 43 minutes 12 seconds of downtime in a 30-day window. Whether it is right depends on user expectations, dependencies and the cost of higher reliability.

Uptime Tracker

Start monitoring in under five minutes

Start on the free plan — commercial use allowed. No credit card, no password, just your email address.

  • Free forever plan
  • No credit card
  • Cancel anytime