Handle LLM API outages: retries, fallback and graceful degradation
Keep SaaS AI features reliable during provider errors with error classification, bounded retries, circuit breakers, clear degraded behavior and tested provider failover.
In this guide
How should a SaaS application handle an LLM provider outage?
Separate temporary provider trouble from errors that will not recover by retrying, then apply a bounded recovery plan. Respect provider retry guidance, use a small retry budget with backoff, open a circuit when failures persist, and return a clear degraded response when the AI feature cannot safely run. A provider outage should not trigger an unbounded agent loop, duplicate customer actions or a silent switch that changes data handling or answer quality.
Classify errors before deciding whether to retry
Distinguish temporary overload, network timeout and rate limits from invalid requests, missing permissions, exhausted credits and spend limits. Retryable conditions may merit a bounded delay; authentication, schema and billing errors need correction, not repeated calls. Read documented status, error code and Retry-After information for your provider and SDK version instead of treating every non-success as a transient outage.
Bound attempts and respect retry-after guidance
Set a maximum attempts count, total time deadline and per-tenant retry budget. When the provider supplies a Retry-After delay, wait at least that long; otherwise use exponential backoff with jitter and avoid synchronized retry storms. Coordinate retries at one layer in the call stack, and account for automatic SDK retries so your application does not multiply them unknowingly.
Make uncertain outcomes safe to resume
A timeout may occur after a provider or downstream tool completed work. Track request state and use idempotency keys for side effects so a retry does not send duplicate messages, charge an account or modify the same record twice. Do not retry invalid, unauthorized or explicitly non-retryable work merely because the UI displayed an error.
| Error class | Retry? and limit | User-visible behavior | Fallback/stop trigger | Owner and test |
|---|---|---|---|---|
| Temporary overload | ||||
| Rate limit | ||||
| Auth/configuration error |
How do circuit breakers and degraded modes protect customers?
Open a circuit when provider failures persist
Measure timeouts, overload responses and dependency health. Stop sending new traffic for a defined window after a failure threshold, then probe recovery with limited requests. Keep the circuit scoped to the failing provider, model route or feature so one outage does not unnecessarily disable unrelated SaaS functions.
Define a useful and honest degraded experience
When safe, allow a user to save a draft, browse already-authorized source material, retry later or complete the task manually. Clearly label stale or partial results and do not present a fallback answer as current, verified or equivalent when it is not. If a safe partial result is unavailable, return a clear error and a next step rather than inventing an answer.
Control queues and shed load before a retry cascade
Limit queue depth, concurrency and waiting time; expire abandoned tasks; and prioritize work according to a documented customer need. Reject or defer excess work cheaply when the provider is overloaded instead of keeping workers occupied with requests that cannot succeed. Monitor retry ratio and queue age as early signs of a growing incident.
When is a second AI provider a safe fallback?
Validate behavior and compatibility before enabling failover
Test the alternate model against the same quality, structured output, safety, latency and tool-use checks as the primary. A common API shape does not guarantee identical behavior. Keep the same authorization, tenant filtering, moderation and output validation in place on both routes.
Review data terms and location before moving requests
Confirm the alternate provider, endpoint, region, retention and subprocessors are approved for the data and customer promise. Do not send a customer's prompt to a second service during an outage unless the service's processing and transfer terms allow it. If approval is absent, degrade or pause the AI feature instead.
Exercise failover and recovery without production side effects
Run scheduled tests using synthetic or approved traffic. Verify the circuit opens, the alternate route works, usage remains within budget, customers see accurate status and recovery does not switch traffic back and forth rapidly. Document how to disable the fallback and how to reconcile work that was in flight during the transition.
LLM API outage FAQs
Should every LLM API error be retried?
No. Retry only conditions documented as temporary, within a bounded policy. Correct authentication, request and billing errors instead of repeating them.
How many times should an application retry an LLM call?
There is no universal number. Set a small per-request and service-wide budget, a total deadline, and backoff with jitter. Include retries performed by your SDK and follow the provider's Retry-After instruction.
Should we automatically fail over to another model provider?
Only after the alternative has passed quality, safety, data-processing, regional and cost review. Otherwise, use a clear degraded mode or pause the feature.
What should users see during an AI outage?
Give a timely, honest status and a useful next step—such as saving work, trying later or using an approved manual path. Do not imply an incomplete or substituted answer is verified or equivalent.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; engineering, security and editorial review pending · Sources checked .
- Error codesOpenAI API documentation
- Google SRE: Handling OverloadGoogle SRE Book
- Google SRE: Addressing Cascading FailuresGoogle SRE Book
- LLM10:2025 Unbounded ConsumptionOWASP Gen AI Security Project
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)National Institute of Standards and Technology
- Evaluation best practicesOpenAI API documentation
- Your data and model usage policies by endpointOpenAI Platform Documentation
- Production best practicesOpenAI API documentation