KeyoAPI

← Blog ·

AI Avatar API Python Example: Server-Side Job Polling and Errors

Learn how to structure a Python backend for an AI avatar API: submit jobs, poll safely, handle errors and retries, and protect credentials and costs.

AI avatar workflows often take longer than a typical web request. Generating a talking avatar or lip-synced video may involve submitting work, waiting for a job to finish, and retrieving its result. A reliable integration therefore needs more than a request example: it needs a server-side job lifecycle.

KeyoAPI documents a multi-model gateway and document model discovery and text chat. Confirm in live docs an avatar-video endpoint, job format, SDK method, or support for digital-human or lip-sync generation. The example below is therefore provider-neutral pseudocode. Before implementing it, confirm that your chosen API supports asynchronous jobs and check its live documentation for the actual endpoints, authentication rules, job states, and result format.

Confirm the API Contract First

Before writing the integration, verify these details in the provider’s current documentation:

Do not assume that a general-purpose model endpoint accepts video-generation requests. A model catalog can show which models are currently available, but it does not establish that a particular model supports avatar generation or asynchronous job polling.

Keep Provider Calls on the Server

The browser should send a request to your application, not directly to the avatar provider. Your backend can then authenticate the user, validate inputs, apply usage limits, and keep provider credentials out of client code.

A minimal application flow is:

  1. The client submits an avatar-generation request to your backend.
  2. The backend validates the request and creates an application job record.
  3. A worker submits the job to the provider and stores the provider’s job identifier.
  4. The worker polls for completion, or handles a provider callback if one is documented and verified.
  5. The backend stores the result reference and exposes an authorized status endpoint to the client.

Return your application’s job ID to the client. Avoid exposing provider credentials or relying on a provider job ID as the only record of the request.

Python Job-Polling Structure

The following is language-neutral pseudocode written in a Python-like style. Names such as provider.submit_job and provider.get_job are placeholders, not documented SDK methods or endpoints. Replace them only with operations confirmed by the provider’s live documentation.

def create_avatar_job(user, request): validate_request(request) enforce_user_limits(user) job = store.create_job( user_id=user.id, status="queued" request_summary=safe_summary(request), ) queue.enqueue("submit_avatar_job" job.id) return {"job_id": job.id, "status": job.status} def submit_avatar_job(job_id): job = store.get_job(job_id) try: result = provider.submit_job( validated_payload=load_payload(job_id), idempotency_key=job.id, ) except ProviderError as error: handle_submission_error(job, error) return store.update_job( job_id, status="submitted" provider_job_id=result.job_id, ) queue.enqueue("poll_avatar_job" job_id, delay_seconds=2) def poll_avatar_job(job_id): job = store.get_job(job_id) try: result = provider.get_job(job.provider_job_id) except ProviderError as error: handle_poll_error(job, error) return if result.status in documented_success_states(): store.mark_succeeded(job_id, result_reference=result.output) elif result.status in documented_failure_states(): store.mark_failed(job_id, safe_error(result)) else: queue.enqueue( "poll_avatar_job" job_id, delay_seconds=next_poll_delay(job.poll_attempts), )

This structure separates user-facing requests from long-running provider calls. It also gives your application a place to track retries, enforce deadlines, and return a stable status even if the provider’s response format changes.

Polling, Timeouts, and Retries

Use a background worker or durable task queue for polling. Holding a web request open while a video is generated ties up application capacity and can fail when a client, proxy, or server times out.

Choose a polling interval that respects provider guidance. If no interval is specified, use a modest interval that grows after repeated pending responses, and set a maximum job deadline. Avoid tight polling loops: they can waste resources, increase request volume, and contribute to rate limits or cost where requests are billable.

Retry only errors that may be temporary, such as documented rate-limit responses or transient network failures. Use bounded exponential backoff with jitter, and honor any documented retry guidance. Do not repeatedly retry permanent input errors, authentication failures, or unsupported-model responses.

Submission retries need special care. If the provider accepted a job but your worker lost the response, blindly submitting again could create a duplicate generation. Use a documented idempotency mechanism if one exists. Otherwise, record the uncertainty and follow the provider’s documented reconciliation process rather than inventing a deduplication guarantee.

Handle Errors as Part of the Job Lifecycle

Represent errors in your own job record with a clear category, such as validation failure, provider rejection, rate limit, timeout, or provider-side processing failure. Keep user-facing messages concise and avoid returning raw provider responses, request headers, or internal exception details.

A job should reach a terminal application state when it succeeds, fails, is cancelled where supported, or exceeds your defined deadline. Do not leave expired jobs polling indefinitely. Make status reads safe to repeat, and ensure that users can access only jobs belonging to their account.

Log enough information to investigate failures: your application job ID, provider job ID when available, attempt count, timestamps, and a sanitized error category. Redact credentials, sensitive prompt content, uploaded media URLs, and any other personal data from logs.

Authentication and Security

Store provider API keys in a server-side secret manager or protected environment configuration. Never place them in browser code, public repositories, screenshots, or client-visible error messages. Use separate credentials for development and production when the provider supports them, and rotate exposed or stale keys.

Validate uploaded media and request fields before submission. Apply file-size and duration limits based on the provider’s documented constraints, and use controlled storage for source media and generated output. Treat result URLs as sensitive: confirm whether they are public, signed, or temporary, and avoid exposing them to unauthorized users.

Also review the provider’s data retention and deletion terms before sending face images, voice recordings, or other personal data. Apply your own retention policy to stored inputs, outputs, and job metadata.

Cost and Model Availability

Generation workloads can consume more resources than ordinary text requests. Add per-user quotas, concurrency limits, request-size limits, and a maximum number of active jobs. Track submitted jobs, failures, retries, completion time, and usage data available from the provider. Alert on sudden increases in submissions or repeated polling failures.

Model and feature availability can change. Check the live model catalog and documentation before deployment, and verify that the selected model supports the exact avatar workflow you need. Keep model selection configurable, and handle unavailable or rejected models without repeatedly resubmitting the same request.

KeyoAPI docs cover a live model-list endpoint and advise using returned model IDs; confirm in live docs an avatar-generation API or its job-polling contract. Do not adapt a text-chat example into an avatar integration unless the live documentation explicitly documents that capability.

Practical Checklist

A production-ready avatar integration is a job-management system as much as an API call. Confirm the provider’s current contract first, then build polling, error handling, access control, and cost limits around the operations it actually supports.

← Blog · Home · Docs