KeyoAPI

← Blog ·

OpenAI API Rate Limits per Minute: Planning Concurrency Safely

Learn how to estimate safe API concurrency from request and token limits, handle throttling, and evaluate OpenAI-compatible alternatives without assuming their quotas match.

Rate limits are not a concurrency setting. They are throughput constraints that interact with request size, response length, latency, and traffic bursts. A system that works at low volume can start returning throttling errors when several workers send large requests at once.

Plan against the limits of the specific account, model, and endpoint you will use. Treat published limits as ceilings, not targets, and measure real traffic before increasing concurrency.

Understand the limits that shape throughput

Rate limits commonly include requests per minute (RPM) and tokens per minute (TPM). Some services also apply daily quotas, concurrent-request limits, or separate limits for particular models and endpoints. The exact limits and how tokens are counted depend on the provider and account.

A request can fit within the RPM limit but exceed the TPM budget. For example, a workload with long prompts and large completions may reach its token limit before it sends many requests. Conversely, short requests can hit the RPM limit while using relatively few tokens.

Use the current limits shown for your account and the documentation for the endpoint you call. Do not assume that limits are shared across models or that switching models preserves the same quota.

Estimate a safe starting concurrency

A rough throughput estimate starts with the smallest relevant limit:

request_budget_per_second = RPM / 60
token_budget_per_second = TPM / 60 estimated_token_rate = request_budget_per_second * average_tokens_per_request estimated_request_rate = token_budget_per_second / average_tokens_per_request safe_request_rate = min(request_budget_per_second, estimated_request_rate)

This is an average-rate estimate, not a guarantee. Providers may count input and output tokens differently, and actual token use varies by request. Use measured usage where possible, and leave headroom for variation and bursts.

Concurrency also depends on latency. A useful initial estimate is:

in_flight_requests = target_requests_per_second * average_latency_seconds

For instance, if a service is sustaining two requests per second and average latency is three seconds, roughly six requests may be in flight on average. This estimate does not mean six is a safe fixed limit: latency spikes, uneven request sizes, and provider-side burst rules can change the result. Start below the estimated ceiling, observe behavior, and adjust gradually.

Account for output size and traffic shape

Estimate tokens per request using representative traffic, not only short test prompts. Include expected prompt length, retrieved context, tool or structured-output overhead where applicable, and typical completion length. Track high-percentile request sizes as well as averages; a small number of very large requests can consume a disproportionate share of capacity.

Also consider whether traffic arrives steadily or in bursts. A queue can smooth bursts by releasing work at a controlled rate. Without one, a sudden batch of requests may exceed a per-minute limit even when the daily average looks modest.

Implement a measured rollout

A reliable rollout turns rate limits into explicit application controls:

  1. Inspect the current limits. Check the provider’s live account and model documentation. Record which dimensions apply to the endpoint and model you intend to use.
  2. Measure representative requests. Capture latency, input and output token usage, status codes, and retry counts. Avoid logging prompts or secrets by default.
  3. Set a starting rate below the estimated ceiling. Use a bounded worker pool and a queue or rate limiter. Limit both requests and token volume if your workload varies significantly in size.
  4. Run a gradual load test. Increase traffic in steps. Watch for throttling, rising latency, growing queue depth, and unexpected cost.
  5. Tune from production data. Re-estimate after prompt changes, model changes, or traffic growth. Rate limits and model availability can change, so treat configuration as something to validate continuously.

Keep limits configurable rather than embedding assumptions throughout the application. Separate per-model settings when the provider applies different quotas, and ensure a model change does not silently inherit an unsuitable concurrency target.

Handle throttling and other failures safely

A rate-limit response should reduce pressure on the service, not trigger an immediate retry storm. When a request is throttled, honor any retry guidance returned by the service. Otherwise, use exponential backoff with random jitter and a maximum retry count.

for attempt in 1.maximum_attempts: result = send_request() if result succeeded: return result if result is throttled or transient: wait using bounded exponential backoff with jitter continue fail without retrying return a retryable failure to the caller or queue

Retry only when the error is plausibly temporary. Authentication failures, invalid parameters, unsupported models, and other client errors generally need correction, not repetition. For requests that may have side effects, check whether the operation is safe to repeat or whether the provider supports an idempotency mechanism before retrying.

Set timeouts, cap retries, and use a circuit breaker or temporary load shedding when failures persist. Otherwise, retries can consume more capacity while delaying healthy work. For queued jobs, preserve the original request context and avoid retrying indefinitely.

Evaluate an OpenAI-compatible alternative carefully

An OpenAI-compatible interface can reduce the amount of client integration work when evaluating another API gateway, but compatible request shapes do not establish identical limits, model behavior, error handling, or availability. Validate the parts your application depends on.

KeyoAPI is an OpenAI-compatible gateway with Bearer authentication, a /v1 base URL, and a live model list at GET /v1/models. Check live docs and account info for current limits, supported models, endpoint behavior, and pricing on /pricing-list before planning production capacity.

A migration evaluation should include:

For a KeyoAPI integration, docs show Bearer authentication and examples using the OpenAI Python and JavaScript SDKs with the KeyoAPI base URL. Use the live model catalog rather than relying on a model ID copied from an old example. Store API keys in environment variables or a secret manager; never expose them in browser code, public repositories, screenshots, or logs.

Monitor cost and availability alongside throughput

Higher concurrency can improve throughput only while downstream capacity and rate limits allow it. It can also raise spend more quickly by allowing more requests to run at once. Track request volume, token usage, estimated cost, latency percentiles, throttling, retries, and queue depth together.

Keep a fallback plan for unavailable models or providers, but do not switch automatically unless the alternative has been tested for quality, output compatibility, and cost. Recheck the live model catalog and current pricing during deployment planning; neither availability nor pricing should be inferred from an old integration example.

Practical checklist

Safe concurrency comes from measured limits, controlled traffic, and feedback from production behavior. Revisit the estimate whenever the workload, model, or provider changes.

← Blog · Home · Docs