Rate-limit errors are rarely solved by adding a longer timeout or retrying immediately. A reliable integration treats rate limits as a scheduling problem: requests must be retried carefully, work must be queued when demand exceeds capacity, and the system must remain useful when a model or upstream service is unavailable.
This article presents a provider-neutral design for handling rate limits and explains how the same approach applies when migrating an application to an OpenAI-compatible API gateway such as KeyoAPI.
What a Rate Limit Error Means
A rate-limit response usually indicates that the service is temporarily refusing additional work. Common causes include:
- Too many requests in a short period
- Too many tokens or large payloads over a time window
- Account quota exhaustion
- Insufficient prepaid balance
- Temporary upstream capacity limits
These causes require different responses. A short-lived request burst may recover after a delay. An exhausted account balance will not be fixed by retrying. Your application should therefore distinguish between transient failures and configuration or billing failures.
A useful error classification is:
| Error category | Retry? | Application action |
|---|---|---|
| Temporary rate limit | Yes, with bounded backoff | Delay and retry |
| Temporary upstream failure | Usually | Retry with jitter |
| Request timeout | Usually | Retry with a reasonable timeout |
| Invalid authentication | No | Fix credentials or configuration |
| Account quota or balance exhausted | No until resolved | Alert and stop unnecessary retries |
| Model unavailable | No with the same model | Discover available models and select a supported one |
| Invalid request | No | Fix the request before retrying |
Do not assume that every error containing the words “rate limit” has the same cause. Inspect the response status, error body, request metadata, and account usage information.
Retry With Exponential Backoff
Immediate retries amplify the original traffic spike. If hundreds of workers retry at the same time, the service can remain overloaded even after the initial limit window ends.
Exponential backoff increases the delay after each failed attempt:
delay = min(max_delay, base_delay * 2^attempt) + random_jitter
For example, a client might wait approximately 1 second, then 2, then 4, then 8 seconds. The random component prevents multiple workers from retrying simultaneously.
Use Full Jitter
A practical strategy is full jitter:
exponential_delay = min(max_delay, base_delay * 2^attempt)
delay = random_number_between(0, exponential_delay)
If the server provides a retry delay, prefer that value when it is valid and within your configured safety limit. Otherwise, use your own exponential-backoff policy.
Bound the Number of Attempts
Retries need a hard limit. A request that can retry indefinitely may consume worker capacity, increase cost, and create a backlog that never clears.
A language-neutral retry loop looks like this:
for attempt from 0 through max_retries: response = send_request() if response succeeded: return response if response is a permanent error: fail immediately if response is transient: if attempt == max_retries: fail and record the exhausted retry delay = server_retry_hint_or_exponential_backoff(attempt) sleep(delay) raise retry_exhausted_error
Set the retry limit according to the operation. Interactive requests generally need a short budget so users are not left waiting. Background jobs can tolerate longer delays, provided the queue and worker limits are controlled.
Do Not Retry Every Failure
Retry only when the failure is plausibly temporary. Retrying an invalid model name, malformed request, invalid API key, or exhausted account balance wastes time and may produce additional charges.
For timeout errors, retrying can be appropriate when the operation is safe to repeat. Use a reasonable client timeout and distinguish connection failures from a response that confirms the request was accepted but took too long to complete.
Add a Queue for Bursty Work
Backoff helps one request recover. A queue manages many requests competing for limited capacity.
A production queue should separate request acceptance from provider execution:
- The application validates and records a job.
- A worker takes a job from the queue.
- The worker sends the API request under a concurrency limit.
- Transient failures return the job to the queue with a scheduled retry time.
- Permanent failures move to a dead-letter queue or terminal-error state.
- Completed results are stored and made available to the caller.
A queue record should normally include:
- A unique job ID
- An idempotency or deduplication key
- Request type and model ID
- Payload reference, rather than unnecessarily duplicating large content
- Attempt count
- Next eligible retry time
- Creation and expiration timestamps
- Current status
- Error classification and last error message
Control Concurrency
A queue without a worker limit simply moves the burst from the web tier to the worker tier. Start with a conservative concurrency value and increase it only after observing successful throughput, latency, error rates, and account usage.
Useful controls include:
- Maximum active requests per model
- Maximum total provider concurrency
- Maximum jobs per tenant
- Maximum token or payload budget per time window
- Queue age limits
- Separate priority queues for interactive and batch work
A token-aware limiter is often more accurate than a request-count limiter because one large prompt can consume substantially more capacity than one small prompt.
Apply Backpressure
When the queue grows beyond a safe size, the system must stop accepting unlimited work. Options include:
- Return a temporary “try again later” response
- Accept the job but provide a status endpoint
- Reduce optional work such as summarization or enrichment
- Defer low-priority jobs
- Reject requests that exceed a tenant’s quota
- Serve cached or previously generated results where appropriate
Backpressure is a product decision as well as an infrastructure decision. Make the behavior explicit instead of allowing memory usage and retry latency to grow without limit.
Design for Duplicate Requests
Retries can create duplicate provider requests when the client times out after the provider has accepted the request. This is especially important for operations that incur usage charges or trigger side effects.
Use an idempotency strategy appropriate to the provider and operation:
- Generate a stable job ID before the first attempt.
- Store the request and result in your database.
- Check whether a completed result already exists before sending work.
- Ensure that a worker lease can expire and be safely reclaimed.
- Keep result writes atomic with the job state transition.
- Do not assume that a network timeout means the provider did not process the request.
If the provider supports an idempotency mechanism, verify its behavior in the live documentation before depending on it. Otherwise, enforce deduplication within your own job system.
Migrate Through an OpenAI-Compatible Interface
An OpenAI-compatible API can reduce application changes because the request shape and authentication pattern may resemble an existing integration. Compatibility should still be treated as an interface claim to verify, not as a guarantee that every model, parameter, response field, streaming mode, or error behavior is identical.
A safe migration workflow is:
- Isolate the provider base URL in configuration.
- Keep the API key in a secret manager.
- Replace hard-coded model assumptions with configuration or discovery.
- Run contract tests against the new service.
- Compare output shape, latency, error behavior, and usage.
- Roll out gradually with clear rollback controls.
For KeyoAPI, the documented base URL is:
https://www.keyoapi.xyz/v1
Requests use Bearer-token authentication. The API key should remain server-side and must not be placed in browser code, public articles, screenshots, or source repositories.
A documented chat request uses this endpoint:
curl https://www.keyoapi.xyz/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_LIVE_CATALOG", "messages": [ { "role": "user", "content": "Hello" } ] }'
The model value in this example is intentionally a placeholder. Do not copy a model name from an old integration and assume that it is available on the new service.
Discover Models at Runtime
Model availability can change. A model may be renamed, removed, restricted, or unavailable for a particular account. Applications should not silently assume that a model is permanently supported.
KeyoAPI documents a model discovery endpoint:
curl https://www.keyoapi.xyz/v1/models \
-H "Authorization: Bearer YOUR_API_KEY"
Use an exact model ID returned by the live response. The current model catalog and pricing information are available at:
A robust deployment can validate its configured model during startup, deployment checks, or a scheduled health check. For request-time fallback, use a predefined policy that has been tested for quality and cost. Do not automatically select an arbitrary model merely because it appears in the catalog.
When model discovery fails, treat that as a dependency or authentication problem rather than repeatedly retrying every application request.
Authentication and Security
Rate-limit handling must not weaken security controls.
Follow these practices:
- Store API keys in a secret manager or protected environment configuration.
- Send keys only over HTTPS.
- Never log the
Authorizationheader. - Redact prompts and generated content when they may contain personal or confidential data.
- Restrict access to queue payloads and stored results.
- Rotate keys according to your operational policy.
- Use separate credentials for development, staging, and production where supported.
- Apply tenant-level authorization before enqueueing work.
- Avoid placing provider keys in client-side JavaScript.
Be careful with retry logs. Logging the full request body on every failed attempt can expose sensitive content and multiply storage costs. Prefer request IDs, model IDs, status categories, latency, attempt count, and a redacted error summary.
Error Handling and Observability
A useful error-handling layer should return stable internal error categories even when providers use different response formats.
Recommended categories include:
authentication_errorinvalid_requestmodel_unavailablerate_limitedquota_exhaustedtimeoutupstream_unavailableretry_exhaustedunknown_provider_error
Record metrics such as:
- Requests attempted and completed
- Retry count by error category
- Queue depth and oldest job age
- Success rate after retry
- Provider latency
- Timeout rate
- Token or usage estimates
- Cost by tenant, model, and feature
- Number of jobs sent to the dead-letter queue
Alert on sustained queue growth, repeated quota failures, increasing retry rates, and model discovery failures. A single rate-limit response is usually not an incident; a backlog that continues to grow is.
Cost and Capacity Controls
Retries consume resources. A timeout followed by a successful retry may represent two provider requests, depending on when the original request was processed. Track usage across attempts and include retry costs in feature-level budgets.
Practical controls include:
- Maximum attempts per job
- Maximum wall-clock age for a job
- Per-tenant daily or monthly budgets
- Maximum prompt and output sizes
- Lower-cost models for non-critical work, after quality testing
- Caching for deterministic or repeatable requests
- Sampling or batching for offline workloads
- Separate budgets for interactive and background traffic
Do not describe a provider’s price, quota, or service guarantee unless it is confirmed in the current pricing and documentation pages. Pricing and availability can change, so operational configuration should be reviewed against the live catalog before deployment.
A Production Evaluation Workflow
Before migrating a live workload or changing retry behavior, test the full path:
1. Validate the contract
Confirm the base URL, authentication method, model discovery endpoint, request format, response format, and supported parameters in the live documentation.
2. Test failure classification
Simulate or capture representative cases for rate limits, timeouts, invalid credentials, unavailable models, malformed requests, and quota exhaustion. Verify that permanent errors do not enter an endless retry loop.
3. Test queue behavior
Measure throughput and queue age under normal load, short bursts, sustained overload, and provider recovery. Verify that backpressure activates before worker memory or database capacity is exhausted.
4. Test duplicate protection
Force a client timeout after request submission and confirm that the resulting job does not create an unintended duplicate result.
5. Compare quality and cost
Use a representative evaluation set. Compare output quality, latency, retry frequency, model availability, and total cost per completed task rather than comparing only the successful response time.
6. Roll out gradually
Start with a small traffic percentage or an internal workload. Keep the provider configuration switchable and define rollback conditions before increasing traffic.
Practical Checklist
- Classify rate limits, timeouts, quota failures, and permanent request errors separately.
- Use exponential backoff with jitter.
- Honor a valid server retry hint when appropriate.
- Set a maximum retry count and wall-clock retry budget.
- Limit worker concurrency and request rates.
- Use a durable queue for bursty or background work.
- Add backpressure when queue depth or age exceeds safe limits.
- Track job IDs and protect against duplicate processing.
- Keep API keys server-side and redact them from logs.
- Discover live model IDs instead of relying on hard-coded assumptions.
- Verify compatibility for parameters, responses, streaming, and errors.
- Monitor queue age, retry rates, latency, usage, and cost.
- Stop retrying when the account quota or prepaid balance is exhausted.
- Review the current documentation, model catalog, and pricing before deployment.
- Roll out changes gradually with a tested rollback path.
A reliable API integration does more than retry failed requests. It controls admission, schedules work, identifies permanent failures, protects credentials, measures cost, and adapts to changing model availability. That combination keeps rate limits from becoming an outage while giving developers a clear path for evaluating or migrating to an OpenAI-compatible service.