Rate-limit errors are rarely solved by adding a longer timeout or retrying immediately. A production client needs to control concurrency, delay retries, preserve request ordering where necessary, and distinguish temporary throttling from permanent failures such as invalid credentials, exhausted account quota, or an unavailable model.
This article presents a provider-agnostic design for handling Gemini API rate limits and explains how the same architecture can support an alternative API gateway such as KeyoAPI. The examples avoid assuming compatibility between providers. Before migrating, verify the target provider's current endpoint, model catalog, request schema, limits, and pricing in its live documentation.
What a Rate Limit Error Means
A rate-limit response generally means that the service is temporarily refusing additional work. The cause may be:
- Too many requests in a short period
- Too many concurrent requests
- Token or input-volume limits
- Account quota exhaustion
- Insufficient prepaid balance
- A provider-side capacity constraint
These cases do not all have the same remedy.
A temporary rate limit may succeed after a delay. Account quota exhaustion may require an account or billing change. Insufficient balance cannot be fixed by retrying. An invalid model ID or authentication failure should fail immediately rather than enter a retry loop.
Your client should therefore classify failures before deciding whether to retry.
The Core Design: Queue, Limit, Retry
A reliable request path has four stages:
- Accept work into a bounded queue.
- Control the number of active requests.
- Retry only transient failures.
- Apply exponential backoff with jitter.
The queue protects the upstream API from traffic bursts. The concurrency limit controls in-flight requests. Backoff spreads retries over time. Jitter prevents many workers from retrying simultaneously after receiving the same error.
A simplified flow looks like this:
application request | v
bounded work queue | v
worker pool with concurrency limit | v
provider request | +--> success | +--> transient error --> delayed retry | +--> permanent error --> fail and report
Why a queue is better than direct retries
Suppose 1,000 application requests arrive at once and every request immediately calls the provider. If the provider rejects most of them, the application may generate another burst when all clients retry together.
A queue changes the traffic shape:
- The application absorbs a controlled amount of pending work.
- Workers process requests at a known rate.
- Retry delays reduce pressure during an incident.
- The queue can reject excess work before memory usage grows without bound.
A queue does not increase your quota. It gives your system a way to respect the quota consistently.
Choose the Queue Boundary
The correct queue boundary depends on your workload.
In-process queue
An in-process queue is appropriate for:
- A single service instance
- Short-lived requests
- Best-effort background work
- Low operational complexity
Its main limitation is durability. Queued work disappears when the process restarts unless the application can safely recreate it.
Durable queue
Use a durable queue when:
- Requests must survive deployments or crashes
- Processing may take minutes
- You need multiple consumers
- You need visibility into pending, delayed, and dead-lettered work
- The producer and worker need to scale independently
The queue should store enough information to reconstruct the request, including a request ID, model identifier, input reference, attempt count, and deadline. Avoid storing sensitive prompts or outputs unless the data-handling policy permits it.
Bounded admission
Every queue needs a capacity policy. When the queue is full, choose explicitly among:
- Rejecting the request with a retryable application error
- Returning an accepted status and processing asynchronously
- Dropping low-priority work
- Applying backpressure to the caller
An unbounded queue only moves the failure from the provider to your own memory and latency budget.
Use Exponential Backoff With Jitter
A typical retry delay is:
delay = min(max_delay, base_delay * 2^attempt) + random(0, jitter_window)
For example, a service might use:
- A small base delay
- A maximum delay of tens of seconds
- A small random jitter window
- A maximum retry count or overall deadline
The exact values should be measured against the provider's documented limits and your application's latency requirements. Do not assume that a particular delay is appropriate for every provider.
Honor server guidance
When the response includes a Retry-After value, treat it as the primary delay signal unless your operational policy imposes a stricter maximum. If the value is absent, use your exponential backoff policy.
Do not retry immediately because the previous request failed quickly. A fast failure is still a signal that the upstream service is refusing work.
Avoid synchronized retries
Without jitter, workers often follow this pattern:
request fails at 10:00:00
all workers retry at 10:00:02
all workers fail again
all workers retry at 10:00:06
Jitter spreads those retries across a time window and reduces synchronized load.
Classify Errors Before Retrying
Your retry policy should be based on error semantics, not only on an HTTP status code.
Usually retryable
These failures may be temporary:
- Rate-limit responses
- Temporary upstream overload
- Gateway or service-unavailable responses
- Network connection resets
- Request timeouts, when the operation is safe to retry
A timeout needs special treatment. The provider may have completed the request even though the client did not receive the response. Retrying a non-idempotent operation can create duplicate side effects. For text generation, duplicate work may increase cost even if the application result is eventually correct.
Usually not retryable
Fail fast for:
- Invalid API keys
- Permission failures
- Malformed request bodies
- Invalid model IDs
- Unsupported parameters
- Content-policy failures that will not change on retry
- Account quota exhaustion
- Insufficient prepaid balance
KeyoAPI returns rate limits, exhausted account quota, and insufficient prepaid balance as separate error conditions. Reduce request frequency and check your account usage and balance before retrying. A retry loop should not treat those conditions as interchangeable.
Model availability is a separate concern
Model identifiers can change. If a provider returns a model-not-found error, do not guess a replacement model name. Query the provider's current model catalog and use an exact identifier returned by that API.
For KeyoAPI, the documented model discovery request is:
curl "$KEYOAPI_BASE_URL/models" \
-H "Authorization: Bearer $KEYOAPI_API_KEY"
Set KEYOAPI_BASE_URL to:
https://www.keyoapi.xyz/v1
The response should be treated as live configuration. Applications should not assume that a model remains available indefinitely.
A Provider-Agnostic Retry Algorithm
The following pseudocode illustrates the control flow without depending on a particular SDK or provider response format:
function process(job): deadline = now + job.maximum_duration attempt = 0 while now < deadline and attempt < job.maximum_attempts: response = send_request(job) if response.success: return response.result classification = classify(response) if classification == permanent_failure: record_failure(job, response) return failure if classification == quota_or_balance_failure: alert_and_stop(job, response) return failure delay = retry_after(response) if delay is missing: delay = exponential_backoff_with_jitter(attempt) if now + delay >= deadline: break sleep(delay) attempt += 1 move_to_dead_letter_or_retry_queue(job) return failure
Several details matter:
- The retry budget has both an attempt limit and a time limit.
- Permanent failures do not consume unnecessary retries.
- Quota and balance failures trigger operational action.
- A final failure is recorded with enough context to investigate.
- The queue or dead-letter system preserves the original request ID.
Controlling Concurrency
Retry logic cannot compensate for excessive concurrency. Start with a conservative worker count and measure:
- Requests per second
- Tokens or input volume per minute
- Successful completion rate
- Rate-limit rate
- Queue wait time
- Provider latency
- Retry count
- Cost per completed request
A useful control loop is:
if rate_limit_rate increases: reduce concurrency increase admission delay
elif queue latency is high and rate_limit_rate is low: increase concurrency gradually
Use gradual changes rather than continuously adjusting concurrency on every response. Otherwise, the system can oscillate between overload and underutilization.
You may also need separate limits for:
- Different models
- Different tenants
- Interactive versus batch traffic
- Input-heavy versus output-heavy requests
- High-priority versus low-priority jobs
A single global semaphore is often too coarse for a multi-tenant service.
Authentication and Secret Handling
Provider migration frequently exposes authentication mistakes because each provider may use different credentials, headers, or account scopes.
For KeyoAPI, the documented authentication format is Bearer token authentication against the API base URL. A request to the chat endpoint has this general shape:
curl "$KEYOAPI_BASE_URL/chat/completions" \
-H "Authorization: Bearer $KEYOAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "MODEL_ID_FROM_LIVE_CATALOG", "messages": [ { "role": "user", "content": "Return a short status summary." } ] }'
The model identifier must come from the current /models response. The request schema and supported models should be confirmed in the KeyoAPI documentation before production use.
Protect API keys by:
- Storing them in a secret manager or environment configuration
- Keeping them out of source repositories and logs
- Never embedding them in browser code
- Rotating them when personnel or deployment boundaries change
- Using separate keys or accounts for environments where supported
- Redacting authorization headers from request traces
Do not log full prompts, outputs, or authorization headers by default. These may contain personal, proprietary, or regulated data.
Timeouts, Idempotency, and Duplicate Work
Set a client timeout that reflects the model and request type. A timeout that is too short can create unnecessary retries; one that is too long can consume worker capacity while the queue grows.
Use separate limits for:
- Connection establishment
- Request transmission
- Time to first response
- Total request duration
When retrying after a timeout, ask whether duplicate work is acceptable. If the provider supports idempotency keys, use them according to its documentation. If it does not, assign an internal request ID and deduplicate results in your own system.
A practical pattern is:
request_id -> pending | running | succeeded | failed
Before starting a retry, check whether the request already has a successful result. This is particularly important when the client timed out after the provider may have completed the operation.
Error Handling at the Application Boundary
Do not expose raw provider errors directly to end users. Convert them into stable application-level categories:
| Internal category | Typical client behavior |
|---|---|
temporary_overload |
Retry asynchronously |
rate_limited |
Queue or retry after delay |
quota_exhausted |
Show an operational or billing message |
insufficient_balance |
Stop retries and notify an administrator |
invalid_model |
Refresh model configuration and fail the request |
authentication_failed |
Alert and require credential correction |
request_invalid |
Return a validation error |
deadline_exceeded |
Return a timeout and preserve job state |
Include a correlation ID in the application response and logs. Store the upstream status and a sanitized error category for debugging, but avoid making your application contract depend on provider-specific wording.
Observability for Rate-Limit Incidents
At minimum, measure these metrics:
- Queue depth
- Queue age
- Active worker count
- Request rate
- Success rate
- Rate-limit count
- Retry count by attempt number
- Retry delay
- Timeout count
- Permanent failure count
- Dead-letter count
- Estimated cost
- Usage by tenant and model
Useful alerts include:
- Queue age exceeds the user-facing latency objective
- Rate-limit responses exceed a normal baseline
- Retry volume grows faster than successful completions
- A model suddenly returns unavailable errors
- Balance or quota errors begin appearing
- Dead-letter volume increases
- Authentication failures occur across multiple workers
Log structured fields such as:
request_id
provider
model_id
attempt
error_category
http_status
retry_delay_ms
queue_wait_ms
total_duration_ms
tenant_id
Do not log sensitive request contents merely to make an incident easier to investigate.
Cost Controls
Retries consume resources. A badly tuned retry loop can increase cost while reducing reliability.
Use these controls:
- Limit maximum attempts
- Set an overall deadline
- Avoid retrying permanent failures
- Cap queue retention
- Set per-tenant budgets
- Separate interactive and batch workloads
- Track usage by model and request type
- Prefer smaller or faster models only when quality requirements permit
- Alert on unusual retry-related usage
When evaluating a migration, compare total cost per successful result, not only the nominal request price. Include queue delay, failed attempts, repeated generation, operational overhead, and the cost of fallback processing.
For current KeyoAPI pricing, consult the live pricing and model catalog. Pricing and model availability can change, so do not hard-code those values into an article, service, or deployment decision without checking the current page.
Migrating Behind a Provider Adapter
A provider adapter keeps application code independent from provider-specific request formats and error details.
Define an internal interface around your actual business needs:
generate_text( model, messages, timeout, request_id
) -> result
The adapter should handle:
- Authentication headers
- Endpoint construction
- Request serialization
- Response parsing
- Error classification
- Provider-specific retry hints
- Model catalog validation
The queue and retry layer should call the adapter rather than constructing provider requests throughout the application.
A migration workflow can then proceed in stages:
1. Inventory current behavior
Record:
- Models currently used
- Maximum prompt and output sizes
- Streaming requirements
- Tool or structured-output requirements
- Timeout behavior
- Error handling
- Rate and concurrency assumptions
- Data residency and retention requirements
- Current cost per successful result
2. Verify the target provider
Check the live documentation and model catalog for:
- Authentication format
- Base URL
- Supported models
- Request and response schema
- Streaming behavior
- Error format
- Rate limits
- Quota behavior
- Timeout guidance
- Pricing
- Data handling terms
Do not infer Gemini feature parity merely because two services expose a similar-looking endpoint. KeyoAPI documents an OpenAI-compatible gateway and chat completions — confirm any Gemini-class model, endpoint, or behavior in the current documentation and model catalog before production.
3. Implement the adapter
Keep provider-specific details inside one module. The rest of the system should see normalized results and error categories.
4. Replay representative traffic
Use sanitized production-like requests to compare:
- Output quality
- Latency distribution
- Error rates
- Rate-limit behavior
- Token or usage accounting
- Retry frequency
- Total cost
Do not rely on a single successful request as evidence of production compatibility.
5. Run a controlled rollout
Start with a small percentage of traffic. Keep rollback available, and compare the new path against the existing path using the same queue and observability controls.
Production Checklist
Before deploying a rate-limit-aware integration, verify:
- Requests enter a bounded queue or have explicit admission control.
- Worker concurrency is limited and configurable.
- Retries use exponential backoff and jitter.
-
Retry-Afteris honored when present. - Maximum attempts and an overall deadline are enforced.
- Permanent failures are not retried.
- Quota and balance errors trigger operational action.
- Model IDs are validated against the provider's current catalog.
- Authentication headers are stored and transmitted securely.
- API keys are absent from source code, browser code, and logs.
- Timeouts are configured for connection and total request duration.
- Duplicate work after timeouts is addressed.
- Queue depth, age, retries, failures, and costs are observable.
- Dead-letter handling exists for jobs that exceed their retry budget.
- Per-tenant or per-workload limits prevent noisy neighbors.
- Migration tests cover schema, model behavior, latency, errors, and cost.
- Current provider documentation, model availability, and pricing have been checked.
Conclusion
Handling Gemini API rate limits is a traffic-management problem, not just a retry problem. A bounded queue controls admission, a concurrency limit controls pressure, and exponential backoff with jitter prevents synchronized retries. Correct error classification prevents the system from wasting time on invalid credentials, unavailable models, exhausted quotas, or insufficient balance.
The same design supports provider migration because the queue, retry policy, observability, and application error contract can remain stable while provider-specific authentication and request handling live behind an adapter. Verify every target capability against the provider's current documentation and model catalog before treating an alternative as production-compatible.