KeyoAPI

← Blog ·

AI Avatar Video Generator API: Queue, Retry, and Webhook Design

Learn how to build a reliable AI avatar video generator API workflow with asynchronous queues, retries, signed webhooks, idempotency, error handling, and production cost controls.

AI avatar video generation is rarely a good fit for a synchronous HTTP request. Rendering a talking avatar, synchronizing speech with facial movement, and producing a downloadable video can take much longer than a normal API timeout.

A production-ready integration should therefore treat video generation as an asynchronous job:

  1. Accept and validate a request.
  2. Create a durable job record.
  3. Enqueue the work.
  4. Submit the job to an avatar or video provider.
  5. Track provider status.
  6. Retry only when safe.
  7. Process a webhook or poll for completion.
  8. Store the result and notify the client.

This article presents a provider-neutral architecture. It does not assume that a particular API supports avatar video, lip-sync, or talking-head generation. Always verify the provider’s current documentation, model catalog, API limits, and webhook behavior before implementation.

Why avatar video generation should be asynchronous

A typical avatar-video request may involve:

These stages introduce variable latency. A short script may complete quickly, while a longer script, high-resolution output, or busy provider queue may take substantially longer.

A synchronous endpoint creates several problems:

Instead, return a job identifier immediately:

POST /avatar-videos
→ 202 Accepted
{ "job_id": "job_123", "status": "queued"
}

The client can then retrieve status or receive a webhook when the job changes state.

Reference architecture

A robust design separates your public API from the provider integration.

Client │ ▼
Your API ├── Validate request ├── Authenticate caller ├── Create idempotent job └── Publish queue message │ ▼ Worker ├── Submit generation request ├── Store provider job ID ├── Poll or receive webhook ├── Download result └── Update job state │ ▼ Database + Object Storage │ ▼ Client webhook/status API

Recommended components

API service

Database

Stores the durable state of each generation job, including:

Queue

Decouples request intake from expensive work. The queue should support delayed messages or scheduled retries.

Worker

Submits jobs to the provider, handles status transitions, downloads outputs, and performs cleanup.

Object storage

Stores the final video and, if required, intermediate audio or subtitle files. Prefer private objects with time-limited download URLs.

Webhook receiver

Accepts provider callbacks, verifies authenticity, and updates jobs idempotently.

Design a durable job state machine

Avoid representing a job with only pending and complete. More precise states make retries and operations easier.

A practical state model is:

queued
submitted
processing
succeeded
failed
cancelled
expired

You may also use a separate retry state, but it is often simpler to keep the business state as queued or processing and store retry metadata separately.

Important state transitions

queued → submitted
submitted → processing
processing → succeeded
processing → failed
queued → cancelled
processing → cancelled
processing → expired

Every transition should be validated. For example, a late webhook must not move a succeeded job back to processing.

Use an optimistic concurrency check or a database transaction:

UPDATE avatar_jobs
SET status = "succeeded", output_url = ?
WHERE id = ? AND status IN ("submitted", "processing")

If the update affects zero rows, the event may be stale or already processed.

Accept requests safely

The public API should validate inputs before placing work on the queue.

Validate at least:

Do not trust client-supplied URLs. If remote assets are allowed, protect the downloader against server-side request forgery by restricting schemes, private IP ranges, redirects, and destination hosts.

A successful submission should return a stable internal job ID:

{ "job_id": "job_123", "status": "queued", "created_at": "2025-01-01T12:00:00Z", "status_url": "/avatar-videos/job_123"
}

The API should normally return 202 Accepted, not 200 OK, because the video is not ready yet.

Use idempotency to prevent duplicate videos

Network failures are common. A client may submit a request successfully, lose the response, and retry. Without idempotency, the retry can create and charge for a second video.

Require an idempotency key for creation requests:

Idempotency-Key: customer-unique-request-456

Store the key together with a request fingerprint and the resulting job ID.

Recommended behavior:

Do not rely solely on an in-memory cache. Idempotency records must survive process restarts and deployments.

Submit work to the provider

Use a provider adapter rather than spreading provider-specific logic throughout your application.

provider.submit(request)
provider.get_status(provider_job_id)
provider.cancel(provider_job_id)
provider.download_result(provider_job_id)
provider.verify_webhook(request)

The adapter should normalize provider responses into your internal format:

{ provider_job_id, status, output_url, retryable, error_code, error_message
}

This design makes it easier to change providers, compare services, or support different avatar and lip-sync engines without changing your public API.

Do not hard-code unverified endpoints, model IDs, SDK methods, or capabilities. Before selecting an integration, confirm:

Webhook design

Webhooks are usually more efficient than frequent polling, but they must be treated as untrusted, repeatable network input.

Verify webhook authenticity

If the provider supports signed webhooks:

  1. Read the raw request body.
  2. Retrieve the signature header.
  3. Compute the expected signature using the shared secret.
  4. Compare signatures using a constant-time comparison.
  5. Validate the timestamp to prevent replay attacks.
  6. Reject invalid or stale requests.

Do not parse and reserialize JSON before signature verification if the provider signs the raw body.

If a provider does not offer signatures, use additional controls:

Acknowledge quickly

The webhook endpoint should do minimal work before responding:

receive event
→ verify signature
→ validate schema
→ store event or update job
→ return 2xx

Do not download a large video or run post-processing inside the webhook request. Queue that work instead.

Handle duplicate and out-of-order events

Providers may deliver the same event more than once. They may also deliver events out of order.

Store an event ID if one is provided. Otherwise, derive a deduplication key from stable fields such as:

provider + provider_job_id + event_type + event_timestamp

Apply only valid state transitions. A terminal state such as succeeded, failed, or cancelled should generally not be overwritten by a later non-terminal event.

Polling as a fallback

Polling is useful when:

Use scheduled polling rather than a tight loop:

poll after 15 seconds
then 30 seconds
then 60 seconds
then 2 minutes

Stop polling when:

A periodic reconciliation task should find jobs stuck in submitted or processing and query the provider again. This protects against worker crashes and missed callbacks.

Retry policy

Retries should distinguish temporary failures from permanent failures.

Usually retryable

Usually not retryable

Use exponential backoff with jitter:

delay = min(max_delay, base_delay × 2^attempt) + random_jitter

Set a maximum attempt count and maximum job age. Infinite retries can create uncontrolled cost and queue growth.

The ambiguous timeout problem

The most dangerous case is a timeout after the provider may have accepted the request. Retrying the submission immediately can create duplicate work.

Use this sequence:

  1. Generate and persist an internal submission attempt ID.
  2. Submit with a provider-supported idempotency key if available.
  3. If the request times out, query the provider or reconcile using the attempt ID.
  4. Submit again only when you have evidence that the first request was not accepted.

If the provider has no idempotency mechanism and no lookup method, document the risk and use conservative reconciliation rules.

Authentication and authorization

Protect both your API and your provider credentials.

Client authentication

Use one of:

Authorize every job lookup. A user must not be able to access another user’s job by guessing its ID.

Provider authentication

Keep provider keys on the server. Never expose them in:

Planning an adjacent text, image, speech, OCR, or multimodal workflow on the same account? Authentication is a Bearer token against the published base URL https://www.keyoapi.xyz/v1, and avatar-video and talking-avatar models (Duix-Avatar, InfiniteTalk) run on the same key and balance. Check the live documentation and model catalog before designing an avatar integration around it. The documented model discovery endpoint is:

GET https://www.keyoapi.xyz/v1/models

Use only model identifiers returned by the live catalog. Do not infer avatar-video support from the presence of general multimodal or speech models.

Error handling and observability

Expose stable, application-level errors rather than leaking provider internals.

Example error categories:

invalid_request
authentication_failed
quota_exceeded
provider_rate_limited
provider_unavailable
content_rejected
asset_unavailable
generation_failed
generation_expired

Store the detailed provider response internally, subject to privacy requirements.

Track metrics such as:

Include a correlation ID in API responses, logs, queue messages, and provider requests where supported. Never log API keys, signed webhook secrets, private asset URLs, or complete user scripts unless necessary and protected.

Cost controls

Avatar video can be expensive because cost may depend on duration, resolution, rendering time, voice usage, or failed attempts.

Control costs with:

Do not assume that a failed HTTP request means no charge was incurred. Record provider usage when available and reconcile it against your internal billing system.

Cache only when the request is semantically identical and the content is safe to reuse. A cache key may include:

avatar + voice + normalized_script + output_settings + provider_version

Avoid caching private or personalized videos across users.

Model and provider availability

Availability changes over time. A model or feature that exists during development may later be renamed, restricted, or removed.

At deployment time, verify:

Run a small health check that confirms configuration without generating unnecessary billable video. Treat an unavailable model as a configuration or capability error, not as a reason to retry indefinitely.

Practical implementation workflow

A reliable implementation can be built in this order:

  1. Define the internal job schema and state machine.
  2. Implement authenticated job creation with validation.
  3. Add durable idempotency records.
  4. Add a queue and worker.
  5. Build a provider adapter with normalized statuses.
  6. Add bounded retries and delayed queue messages.
  7. Implement webhook verification and deduplication.
  8. Add polling reconciliation.
  9. Store completed videos in private object storage.
  10. Add quotas, metrics, alerts, and cost reconciliation.
  11. Test worker crashes, duplicate requests, duplicate webhooks, timeouts, and stale events.
  12. Verify the provider’s live capabilities before production launch.

Production checklist

API and jobs

Queue and retries

Webhooks

Security

Operations and cost

Conclusion

A dependable AI avatar video generator integration is primarily a workflow and reliability problem, not just an API-call problem. Durable jobs, idempotency, bounded retries, verified webhooks, reconciliation, and strict cost controls prevent the most common production failures.

Keep the provider-specific implementation behind an adapter, verify capabilities against live documentation, and design your public API around asynchronous completion. That approach lets you support talking avatars, lip-sync, and other video-generation workflows without coupling your application to undocumented endpoints or fragile assumptions.

Editorial scope and verification

An avatar-video feature is a media job system, not a single synchronous request. Verify the provider’s current contract for avatar inputs, voices, languages, duration, output format, webhooks, storage, pricing, and consent requirements. A text or speech API does not automatically provide avatar rendering or lip-sync.

Production acceptance checklist

Choose a provider only after a representative test set measures completion rate, latency, quality, retry behavior, and cost per successful video. Keep provider-specific calls behind an adapter so the workflow remains replaceable.

← Blog · Home · Docs