Skip to main content

SaaS API rate limiting: tenant quotas, 429 responses and fair use

Design tenant-aware API rate limits and quotas, return useful 429 responses, protect shared resources and test limits against real request costs.

In this guide

What are API rate limits and quotas in SaaS?

A rate limit controls how much request activity a caller can send during a short period or burst. A quota caps usage over a longer period, such as a daily export allowance or monthly API allocation. SaaS teams use them to protect capacity, contain abusive traffic and keep one tenant from degrading a shared service. Limits should reflect the cost and purpose of an operation, not only a single request count.

Choose the identity that represents a caller

Apply limits to an authenticated tenant, API credential, user or a combination that matches the risk. An IP-only limit can affect many legitimate users behind a shared network, while a credential-only limit may miss a compromised account creating many keys. Keep authentication and tenant authorisation separate from the counting mechanism.

Account for bursts, request cost and concurrency

A token bucket can allow a bounded burst while limiting the average rate. For expensive queries, large payloads or long-running jobs, also consider weighted cost, maximum concurrency, queue limits or a cumulative quota. Two endpoints at the same requests-per-second rate may consume very different database or compute capacity.

Return a clear response when a caller reaches a limit

HTTP 429 means the caller has sent too many requests in a period. Explain the limit in a useful way and, when practical, include a Retry-After value so a client can wait before trying again. Avoid leaking internal capacity details or another tenant's usage. The HTTP standard does not prescribe how a server identifies or counts a caller.

Tenant rate-limit design worksheet
Operation / resource costCaller identityRate / burstQuota or concurrency cap429 and retry behaviour
Read API
Expensive report or export
Write or asynchronous job

How should a SaaS team set API limits?

Measure normal and peak demand first

Use traffic and load-test evidence to estimate supported request volume, burst size, payload size and downstream cost. Set initial thresholds with headroom for legitimate customer workflows and dependencies. Record the tested limit and review it when data volume, infrastructure or customer plans change.

Apply controls at more than one shared layer

An edge limit can stop a flood before it reaches the application, but a request already admitted may consume queue, worker, database or third-party capacity. Add suitable concurrency and resource controls deeper in the system. Keep tenant identity attached to asynchronous work so queued jobs cannot bypass the intended budget.

Make limits visible and consistent with the plan

Document the unit, time window, burst behaviour, reset method, fair-use terms and support route for each plan. Ensure product copy, API documentation, dashboards and enforcement use the same values. A plan label such as “unlimited” should not conceal a restrictive undisclosed cap; get contract review for customer-facing terms.

How should API clients handle a throttled response?

Pause before retrying and respect Retry-After

A client should not immediately resend every rejected request. If the response includes Retry-After, wait at least that long; otherwise use a bounded backoff strategy and add jitter when many clients could retry together. Set a retry cap and surface a useful error if the operation still cannot complete.

Make retried writes safe

A network timeout can leave a client uncertain whether a write succeeded. Use idempotency keys or another application-level duplicate protection for operations that may be retried, and distinguish a throttled response from an unknown result. Do not retry a non-idempotent payment or data mutation blindly.

Monitor rejected work as well as accepted work

Track 429 rates by endpoint, tenant plan and caller type, while protecting tenant privacy. A sudden spike can signal a client retry bug, a poorly chosen limit or abuse. Watch for rejected users with legitimate workflows and tune the policy based on evidence without removing capacity safeguards.

SaaS API rate-limit questions

What is the difference between a rate limit and a quota?

A rate limit controls request speed or burst over a short interval. A quota limits total usage over a longer interval. A robust API may need both, plus cost and concurrency limits for expensive operations.

Should we rate-limit only by IP address?

Usually not as the only control. IP addresses can represent many users or change for one user. Combine network protection with authenticated tenant, credential or user limits appropriate to your API and abuse model.

Should every 429 response include Retry-After?

RFC 6585 says a 429 response may include Retry-After. Providing it is useful when the server can give a safe wait time; document client behaviour when the header is absent and avoid suggesting a retry time that the system cannot support.