Skip to main content

SaaS SLI vs SLO vs SLA: examples and error budgets

Understand the difference between an SLI, SLO and SLA, choose a user-focused reliability target, calculate an error budget and review results without promising perfect uptime.

In this guide

What do SLI, SLO, SLA and error budget mean?

An SLI is a defined measurement of a service property, such as the share of valid requests that succeed. An SLO is the target or range for that indicator over a stated window. An SLA is an agreement with a customer about service commitments and what happens if they are missed. An error budget is the amount of unreliability allowed by an SLO during its window. Google SRE uses these terms separately because each should drive a different measurement or decision.

SLI: decide what customer experience to measure

Choose a measurable signal close to a customer outcome: successful core requests, completed transactions, acceptable response time or durable data writes. Define which requests count, how failures are recognised and how missing telemetry is treated. A server uptime probe can be a useful signal, but it may not show whether a customer completed the task.

SLO: set a target and a measurement window

State the indicator, target and window together. For example: “At least 99.9% of eligible document-save requests succeed in each rolling 30-day window.” Define the eligible request set, status codes, exclusions and data source before reporting a result. A different denominator or window can produce a different percentage.

SLA: review an external commitment as a contract

A public or signed service-level agreement may include measurement rules, exclusions, notification steps and remedies such as service credits. It can have legal and financial effects. Do not copy a cloud provider's uptime number into your own agreement without checking product dependencies, monitoring, support capacity and reviewed contract language.

Reliability objective worksheet
Customer journeySLI definition / eligible eventsSLO target and windowData sourceDecision if budget is low
Core save or transaction
Sign-in or API workflow
Data durability or export

How do you choose an SLO for a SaaS product?

Start with the journeys customers rely on

Ask what must work for a user to receive the product's main value, then identify the service boundary and failure the team can measure. Separate a background report delay from a failed payment or lost write if their customer impact differs. Use a small number of targets that can change an engineering or product decision.

Choose a denominator and window your team can explain

Define eligible requests, successful outcomes, latency thresholds, regions and measurement windows in plain language. Check that the data source sees the same event consistently and that retries do not accidentally count one user action several times. Keep exclusions narrow and written down rather than removing inconvenient failures after the fact.

Set a target from customer needs and observed reliability

Review customer expectations, actual service behaviour, dependencies and the cost of more reliability. A 100% target is not a practical way to plan every service: it can hide tradeoffs or encourage teams to spend without improving a meaningful user outcome. Do not adopt a popular percentage without validating its meaning for your product.

How should a team use its error budget?

Calculate the budget from the SLO, not from a slogan

For a request-based objective, the allowed unsuccessful share is the complement of the target over the defined eligible requests and window. If the objective is 99.9% successful eligible requests, the corresponding budget is 0.1% of those requests in that window. A time-based availability SLI uses a different interpretation, so do not convert percentages into minutes unless the SLI is actually time-based.

Use budget burn to guide release and reliability decisions

When the budget is being consumed faster than expected, investigate the affected journey and recent changes. A team may slow risky releases, focus on reliability work or revise a mitigation plan. Agree on a proportionate response in advance; one brief error should not automatically freeze all work.

SaaS reliability objective questions

Is an SLO the same as an SLA?

No. An SLO is an internal or product reliability target measured by an SLI. An SLA is an agreement with customers that can define remedies and other obligations. An SLA may reference an SLO, but the terms should not be treated as interchangeable.

Does a 99.9% SLO mean a fixed amount of downtime?

Only if the SLI and its window measure availability as time. A request-success SLI measures the share of eligible requests that succeed, so its error budget is a share of requests instead of a fixed number of outage minutes. Always state the indicator and window.

How many SLOs should a startup have?

There is no universal number. Begin with a few customer journeys where reliability changes a real product or operations decision. Add a target only when the team can define it, measure it and act on the result.

Can we advertise our SLO as uptime?

Use wording that matches the actual measurement and whether it is a target or contractual commitment. Have legal and product owners review public or signed promises, measurement exclusions and remedies before publishing them.