KeyoAPI

← Blog ·

OpenAI API Rate Limit Exceeded: Retry, Backoff, and Queue Design

Learn how to handle OpenAI API rate limits with exponential backoff, bounded retries, durable queues, model discovery, authentication, and production-safe migration patterns.

Rate-limit errors are rarely solved by adding a longer timeout or retrying immediately. A reliable integration treats rate limits as a scheduling problem: requests must be retried carefully, work must be queued when demand exceeds capacity, and the system must remain useful when a model or upstream service is unavailable.

This article presents a provider-neutral design for handling rate limits and explains how the same approach applies when migrating an application to an OpenAI-compatible API gateway such as KeyoAPI.

What a Rate Limit Error Means

A rate-limit response usually indicates that the service is temporarily refusing additional work. Common causes include:

These causes require different responses. A short-lived request burst may recover after a delay. An exhausted account balance will not be fixed by retrying. Your application should therefore distinguish between transient failures and configuration or billing failures.

A useful error classification is:

Error category Retry? Application action
Temporary rate limit Yes, with bounded backoff Delay and retry
Temporary upstream failure Usually Retry with jitter
Request timeout Usually Retry with a reasonable timeout
Invalid authentication No Fix credentials or configuration
Account quota or balance exhausted No until resolved Alert and stop unnecessary retries
Model unavailable No with the same model Discover available models and select a supported one
Invalid request No Fix the request before retrying

Do not assume that every error containing the words “rate limit” has the same cause. Inspect the response status, error body, request metadata, and account usage information.

Retry With Exponential Backoff

Immediate retries amplify the original traffic spike. If hundreds of workers retry at the same time, the service can remain overloaded even after the initial limit window ends.

Exponential backoff increases the delay after each failed attempt:

delay = min(max_delay, base_delay * 2^attempt) + random_jitter

For example, a client might wait approximately 1 second, then 2, then 4, then 8 seconds. The random component prevents multiple workers from retrying simultaneously.

Use Full Jitter

A practical strategy is full jitter:

exponential_delay = min(max_delay, base_delay * 2^attempt)
delay = random_number_between(0, exponential_delay)

If the server provides a retry delay, prefer that value when it is valid and within your configured safety limit. Otherwise, use your own exponential-backoff policy.

Bound the Number of Attempts

Retries need a hard limit. A request that can retry indefinitely may consume worker capacity, increase cost, and create a backlog that never clears.

A language-neutral retry loop looks like this:

for attempt from 0 through max_retries: response = send_request() if response succeeded: return response if response is a permanent error: fail immediately if response is transient: if attempt == max_retries: fail and record the exhausted retry delay = server_retry_hint_or_exponential_backoff(attempt) sleep(delay) raise retry_exhausted_error

Set the retry limit according to the operation. Interactive requests generally need a short budget so users are not left waiting. Background jobs can tolerate longer delays, provided the queue and worker limits are controlled.

Do Not Retry Every Failure

Retry only when the failure is plausibly temporary. Retrying an invalid model name, malformed request, invalid API key, or exhausted account balance wastes time and may produce additional charges.

For timeout errors, retrying can be appropriate when the operation is safe to repeat. Use a reasonable client timeout and distinguish connection failures from a response that confirms the request was accepted but took too long to complete.

Add a Queue for Bursty Work

Backoff helps one request recover. A queue manages many requests competing for limited capacity.

A production queue should separate request acceptance from provider execution:

  1. The application validates and records a job.
  2. A worker takes a job from the queue.
  3. The worker sends the API request under a concurrency limit.
  4. Transient failures return the job to the queue with a scheduled retry time.
  5. Permanent failures move to a dead-letter queue or terminal-error state.
  6. Completed results are stored and made available to the caller.

A queue record should normally include:

Control Concurrency

A queue without a worker limit simply moves the burst from the web tier to the worker tier. Start with a conservative concurrency value and increase it only after observing successful throughput, latency, error rates, and account usage.

Useful controls include:

A token-aware limiter is often more accurate than a request-count limiter because one large prompt can consume substantially more capacity than one small prompt.

Apply Backpressure

When the queue grows beyond a safe size, the system must stop accepting unlimited work. Options include:

Backpressure is a product decision as well as an infrastructure decision. Make the behavior explicit instead of allowing memory usage and retry latency to grow without limit.

Design for Duplicate Requests

Retries can create duplicate provider requests when the client times out after the provider has accepted the request. This is especially important for operations that incur usage charges or trigger side effects.

Use an idempotency strategy appropriate to the provider and operation:

If the provider supports an idempotency mechanism, verify its behavior in the live documentation before depending on it. Otherwise, enforce deduplication within your own job system.

Migrate Through an OpenAI-Compatible Interface

An OpenAI-compatible API can reduce application changes because the request shape and authentication pattern may resemble an existing integration. Compatibility should still be treated as an interface claim to verify, not as a guarantee that every model, parameter, response field, streaming mode, or error behavior is identical.

A safe migration workflow is:

  1. Isolate the provider base URL in configuration.
  2. Keep the API key in a secret manager.
  3. Replace hard-coded model assumptions with configuration or discovery.
  4. Run contract tests against the new service.
  5. Compare output shape, latency, error behavior, and usage.
  6. Roll out gradually with clear rollback controls.

For KeyoAPI, the documented base URL is:

https://www.keyoapi.xyz/v1

Requests use Bearer-token authentication. The API key should remain server-side and must not be placed in browser code, public articles, screenshots, or source repositories.

A documented chat request uses this endpoint:

curl https://www.keyoapi.xyz/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "MODEL_ID_FROM_LIVE_CATALOG", "messages": [ { "role": "user", "content": "Hello" } ] }'

The model value in this example is intentionally a placeholder. Do not copy a model name from an old integration and assume that it is available on the new service.

Discover Models at Runtime

Model availability can change. A model may be renamed, removed, restricted, or unavailable for a particular account. Applications should not silently assume that a model is permanently supported.

KeyoAPI documents a model discovery endpoint:

curl https://www.keyoapi.xyz/v1/models \
  -H "Authorization: Bearer YOUR_API_KEY"

Use an exact model ID returned by the live response. The current model catalog and pricing information are available at:

A robust deployment can validate its configured model during startup, deployment checks, or a scheduled health check. For request-time fallback, use a predefined policy that has been tested for quality and cost. Do not automatically select an arbitrary model merely because it appears in the catalog.

When model discovery fails, treat that as a dependency or authentication problem rather than repeatedly retrying every application request.

Authentication and Security

Rate-limit handling must not weaken security controls.

Follow these practices:

Be careful with retry logs. Logging the full request body on every failed attempt can expose sensitive content and multiply storage costs. Prefer request IDs, model IDs, status categories, latency, attempt count, and a redacted error summary.

Error Handling and Observability

A useful error-handling layer should return stable internal error categories even when providers use different response formats.

Recommended categories include:

Record metrics such as:

Alert on sustained queue growth, repeated quota failures, increasing retry rates, and model discovery failures. A single rate-limit response is usually not an incident; a backlog that continues to grow is.

Cost and Capacity Controls

Retries consume resources. A timeout followed by a successful retry may represent two provider requests, depending on when the original request was processed. Track usage across attempts and include retry costs in feature-level budgets.

Practical controls include:

Do not describe a provider’s price, quota, or service guarantee unless it is confirmed in the current pricing and documentation pages. Pricing and availability can change, so operational configuration should be reviewed against the live catalog before deployment.

A Production Evaluation Workflow

Before migrating a live workload or changing retry behavior, test the full path:

1. Validate the contract

Confirm the base URL, authentication method, model discovery endpoint, request format, response format, and supported parameters in the live documentation.

2. Test failure classification

Simulate or capture representative cases for rate limits, timeouts, invalid credentials, unavailable models, malformed requests, and quota exhaustion. Verify that permanent errors do not enter an endless retry loop.

3. Test queue behavior

Measure throughput and queue age under normal load, short bursts, sustained overload, and provider recovery. Verify that backpressure activates before worker memory or database capacity is exhausted.

4. Test duplicate protection

Force a client timeout after request submission and confirm that the resulting job does not create an unintended duplicate result.

5. Compare quality and cost

Use a representative evaluation set. Compare output quality, latency, retry frequency, model availability, and total cost per completed task rather than comparing only the successful response time.

6. Roll out gradually

Start with a small traffic percentage or an internal workload. Keep the provider configuration switchable and define rollback conditions before increasing traffic.

Practical Checklist

A reliable API integration does more than retry failed requests. It controls admission, schedules work, identifies permanent failures, protects credentials, measures cost, and adapts to changing model availability. That combination keeps rate limits from becoming an outage while giving developers a clear path for evaluating or migrating to an OpenAI-compatible service.

← Blog · Home · Docs