KeyoAPI

← Blog ·

Gemini API Rate Limit Exceeded: Queue and Backoff Design

Learn how to handle Gemini API rate limits with queues, exponential backoff, jitter, bounded retries, authentication, observability, and provider-agnostic migration patterns.

Rate-limit errors are rarely solved by adding a longer timeout or retrying immediately. A production client needs to control concurrency, delay retries, preserve request ordering where necessary, and distinguish temporary throttling from permanent failures such as invalid credentials, exhausted account quota, or an unavailable model.

This article presents a provider-agnostic design for handling Gemini API rate limits and explains how the same architecture can support an alternative API gateway such as KeyoAPI. The examples avoid assuming compatibility between providers. Before migrating, verify the target provider's current endpoint, model catalog, request schema, limits, and pricing in its live documentation.

What a Rate Limit Error Means

A rate-limit response generally means that the service is temporarily refusing additional work. The cause may be:

These cases do not all have the same remedy.

A temporary rate limit may succeed after a delay. Account quota exhaustion may require an account or billing change. Insufficient balance cannot be fixed by retrying. An invalid model ID or authentication failure should fail immediately rather than enter a retry loop.

Your client should therefore classify failures before deciding whether to retry.

The Core Design: Queue, Limit, Retry

A reliable request path has four stages:

  1. Accept work into a bounded queue.
  2. Control the number of active requests.
  3. Retry only transient failures.
  4. Apply exponential backoff with jitter.

The queue protects the upstream API from traffic bursts. The concurrency limit controls in-flight requests. Backoff spreads retries over time. Jitter prevents many workers from retrying simultaneously after receiving the same error.

A simplified flow looks like this:

application request | v
bounded work queue | v
worker pool with concurrency limit | v
provider request | +--> success | +--> transient error --> delayed retry | +--> permanent error --> fail and report

Why a queue is better than direct retries

Suppose 1,000 application requests arrive at once and every request immediately calls the provider. If the provider rejects most of them, the application may generate another burst when all clients retry together.

A queue changes the traffic shape:

A queue does not increase your quota. It gives your system a way to respect the quota consistently.

Choose the Queue Boundary

The correct queue boundary depends on your workload.

In-process queue

An in-process queue is appropriate for:

Its main limitation is durability. Queued work disappears when the process restarts unless the application can safely recreate it.

Durable queue

Use a durable queue when:

The queue should store enough information to reconstruct the request, including a request ID, model identifier, input reference, attempt count, and deadline. Avoid storing sensitive prompts or outputs unless the data-handling policy permits it.

Bounded admission

Every queue needs a capacity policy. When the queue is full, choose explicitly among:

An unbounded queue only moves the failure from the provider to your own memory and latency budget.

Use Exponential Backoff With Jitter

A typical retry delay is:

delay = min(max_delay, base_delay * 2^attempt) + random(0, jitter_window)

For example, a service might use:

The exact values should be measured against the provider's documented limits and your application's latency requirements. Do not assume that a particular delay is appropriate for every provider.

Honor server guidance

When the response includes a Retry-After value, treat it as the primary delay signal unless your operational policy imposes a stricter maximum. If the value is absent, use your exponential backoff policy.

Do not retry immediately because the previous request failed quickly. A fast failure is still a signal that the upstream service is refusing work.

Avoid synchronized retries

Without jitter, workers often follow this pattern:

request fails at 10:00:00
all workers retry at 10:00:02
all workers fail again
all workers retry at 10:00:06

Jitter spreads those retries across a time window and reduces synchronized load.

Classify Errors Before Retrying

Your retry policy should be based on error semantics, not only on an HTTP status code.

Usually retryable

These failures may be temporary:

A timeout needs special treatment. The provider may have completed the request even though the client did not receive the response. Retrying a non-idempotent operation can create duplicate side effects. For text generation, duplicate work may increase cost even if the application result is eventually correct.

Usually not retryable

Fail fast for:

KeyoAPI returns rate limits, exhausted account quota, and insufficient prepaid balance as separate error conditions. Reduce request frequency and check your account usage and balance before retrying. A retry loop should not treat those conditions as interchangeable.

Model availability is a separate concern

Model identifiers can change. If a provider returns a model-not-found error, do not guess a replacement model name. Query the provider's current model catalog and use an exact identifier returned by that API.

For KeyoAPI, the documented model discovery request is:

curl "$KEYOAPI_BASE_URL/models" \
  -H "Authorization: Bearer $KEYOAPI_API_KEY"

Set KEYOAPI_BASE_URL to:

https://www.keyoapi.xyz/v1

The response should be treated as live configuration. Applications should not assume that a model remains available indefinitely.

A Provider-Agnostic Retry Algorithm

The following pseudocode illustrates the control flow without depending on a particular SDK or provider response format:

function process(job): deadline = now + job.maximum_duration attempt = 0 while now < deadline and attempt < job.maximum_attempts: response = send_request(job) if response.success: return response.result classification = classify(response) if classification == permanent_failure: record_failure(job, response) return failure if classification == quota_or_balance_failure: alert_and_stop(job, response) return failure delay = retry_after(response) if delay is missing: delay = exponential_backoff_with_jitter(attempt) if now + delay >= deadline: break sleep(delay) attempt += 1 move_to_dead_letter_or_retry_queue(job) return failure

Several details matter:

Controlling Concurrency

Retry logic cannot compensate for excessive concurrency. Start with a conservative worker count and measure:

A useful control loop is:

if rate_limit_rate increases: reduce concurrency increase admission delay
elif queue latency is high and rate_limit_rate is low: increase concurrency gradually

Use gradual changes rather than continuously adjusting concurrency on every response. Otherwise, the system can oscillate between overload and underutilization.

You may also need separate limits for:

A single global semaphore is often too coarse for a multi-tenant service.

Authentication and Secret Handling

Provider migration frequently exposes authentication mistakes because each provider may use different credentials, headers, or account scopes.

For KeyoAPI, the documented authentication format is Bearer token authentication against the API base URL. A request to the chat endpoint has this general shape:

curl "$KEYOAPI_BASE_URL/chat/completions" \
  -H "Authorization: Bearer $KEYOAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "MODEL_ID_FROM_LIVE_CATALOG", "messages": [ { "role": "user", "content": "Return a short status summary." } ] }'

The model identifier must come from the current /models response. The request schema and supported models should be confirmed in the KeyoAPI documentation before production use.

Protect API keys by:

Do not log full prompts, outputs, or authorization headers by default. These may contain personal, proprietary, or regulated data.

Timeouts, Idempotency, and Duplicate Work

Set a client timeout that reflects the model and request type. A timeout that is too short can create unnecessary retries; one that is too long can consume worker capacity while the queue grows.

Use separate limits for:

When retrying after a timeout, ask whether duplicate work is acceptable. If the provider supports idempotency keys, use them according to its documentation. If it does not, assign an internal request ID and deduplicate results in your own system.

A practical pattern is:

request_id -> pending | running | succeeded | failed

Before starting a retry, check whether the request already has a successful result. This is particularly important when the client timed out after the provider may have completed the operation.

Error Handling at the Application Boundary

Do not expose raw provider errors directly to end users. Convert them into stable application-level categories:

Internal category Typical client behavior
temporary_overload Retry asynchronously
rate_limited Queue or retry after delay
quota_exhausted Show an operational or billing message
insufficient_balance Stop retries and notify an administrator
invalid_model Refresh model configuration and fail the request
authentication_failed Alert and require credential correction
request_invalid Return a validation error
deadline_exceeded Return a timeout and preserve job state

Include a correlation ID in the application response and logs. Store the upstream status and a sanitized error category for debugging, but avoid making your application contract depend on provider-specific wording.

Observability for Rate-Limit Incidents

At minimum, measure these metrics:

Useful alerts include:

Log structured fields such as:

request_id
provider
model_id
attempt
error_category
http_status
retry_delay_ms
queue_wait_ms
total_duration_ms
tenant_id

Do not log sensitive request contents merely to make an incident easier to investigate.

Cost Controls

Retries consume resources. A badly tuned retry loop can increase cost while reducing reliability.

Use these controls:

When evaluating a migration, compare total cost per successful result, not only the nominal request price. Include queue delay, failed attempts, repeated generation, operational overhead, and the cost of fallback processing.

For current KeyoAPI pricing, consult the live pricing and model catalog. Pricing and model availability can change, so do not hard-code those values into an article, service, or deployment decision without checking the current page.

Migrating Behind a Provider Adapter

A provider adapter keeps application code independent from provider-specific request formats and error details.

Define an internal interface around your actual business needs:

generate_text( model, messages, timeout, request_id
) -> result

The adapter should handle:

The queue and retry layer should call the adapter rather than constructing provider requests throughout the application.

A migration workflow can then proceed in stages:

1. Inventory current behavior

Record:

2. Verify the target provider

Check the live documentation and model catalog for:

Do not infer Gemini feature parity merely because two services expose a similar-looking endpoint. KeyoAPI documents an OpenAI-compatible gateway and chat completions — confirm any Gemini-class model, endpoint, or behavior in the current documentation and model catalog before production.

3. Implement the adapter

Keep provider-specific details inside one module. The rest of the system should see normalized results and error categories.

4. Replay representative traffic

Use sanitized production-like requests to compare:

Do not rely on a single successful request as evidence of production compatibility.

5. Run a controlled rollout

Start with a small percentage of traffic. Keep rollback available, and compare the new path against the existing path using the same queue and observability controls.

Production Checklist

Before deploying a rate-limit-aware integration, verify:

Conclusion

Handling Gemini API rate limits is a traffic-management problem, not just a retry problem. A bounded queue controls admission, a concurrency limit controls pressure, and exponential backoff with jitter prevents synchronized retries. Correct error classification prevents the system from wasting time on invalid credentials, unavailable models, exhausted quotas, or insufficient balance.

The same design supports provider migration because the queue, retry policy, observability, and application error contract can remain stable while provider-specific authentication and request handling live behind an adapter. Verify every target capability against the provider's current documentation and model catalog before treating an alternative as production-compatible.

← Blog · Home · Docs