Answer box: For long-running AI tasks, accept the request, save a durable job record and let a queue worker do the processing. Give each job a stable identity, coordinate worker timeouts with queue redelivery, and retry only recoverable failures within explicit limits. Assume a job can arrive more than once: protect saved results and other side effects against duplicates. Queue delivery alone does not provide that protection. (laravel.com)
The design below is for website owners and developers adding document summaries or similar AI workflows. It is a planning and staging-test guide, not a tested implementation for a particular host.
1. Separate the web request from durable processing
Use this proposed flow:
- Authenticate the user and validate the document reference.
- Save a job record and arrange durable queue delivery.
- Return the job ID without waiting for the summary.
- Let a worker claim the job, call the AI provider and save the result.
- Let the browser check an access-controlled status endpoint.
Laravel’s queue documentation describes moving time-consuming work out of web requests into background processing. AWS’s transactional outbox guidance addresses another important boundary: updating a database and sending a queue message are separate operations that can fail independently. (laravel.com)
If your database-backed queue can insert the job and queue entry in the same transaction, verify that arrangement. Otherwise, consider an outbox: save the job and a pending-dispatch record together, then have a dispatcher publish pending records. Make that dispatcher safe to repeat because duplicate messages remain possible. (docs.aws.amazon.com)
Common mistake: saving “queued” in the database, then losing the process before sending the message.
Practical check: interrupt dispatch in staging. The saved job should remain discoverable and eventually dispatch—not sit indefinitely in a misleading state.
Put document references and identifiers in messages rather than copying entire documents into every queue entry. Keep the original input available until the job’s retention policy permits deletion.
2. Confirm that your host can run the worker
Before choosing a queue framework, ask the host:
- Are persistent worker processes supported?
- What stops or restarts them?
- What execution limits apply to queue handlers?
- Is there enough shutdown time to finish or checkpoint active work?
- Can workers reach the database, queue and AI API?
- Where are durable documents and results stored?
Treat these as verification questions, not assumed hosting features.
For a concrete example, Vercel’s documentation checked on October 3, 2026 lists a 300-second maximum for Hobby functions with Fluid compute. Pro and Enterprise have an 800-second generally available maximum and a 1,800-second beta extended maximum for supported runtimes, with configuration and deployment restrictions. These are invocation limits, not a guarantee that a background workflow will finish. (vercel.com)
A queue does not remove the execution limit of the function processing its messages. Design each processing step to fit its runtime, or use an execution environment or durable workflow that supports the required lifecycle. (vercel.com)
For Laravel workers, verify PHP’s PCNTL extension when relying on job timeouts, and configure outgoing HTTP timeouts separately: blocking I/O may not respect the worker timeout as expected. (laravel.com)
Practical check: restart the worker in staging and confirm that unfinished work becomes recoverable under your actual queue and acknowledgement settings.

3. Make job state visible and persistent
Use an application-level job record rather than asking readers to interpret queue internals. Here is a suggested state model:
| State | Meaning shown to the user |
|---|---|
queued |
Accepted and waiting for processing |
running |
A worker holds the current claim |
retry_wait |
A recoverable failure occurred; another attempt is scheduled |
succeeded |
The final result is durably saved |
failed |
Automatic processing has stopped |
cancel_requested |
Cancellation was requested but has not finished |
cancelled |
Further processing and publication have stopped |
Recommended fields include the owner, job ID, input version, attempt count, timestamps, next-attempt time, cancellation flag and result reference. Record a sanitized error category separately from the user-facing explanation.
For recovery, distinguish a worker claim from permanent completion. SQS, for example, makes an unprocessed message available again when its visibility timeout expires. Its visibility timeout is not an absolute guarantee against duplicate delivery. (docs.aws.amazon.com)
In your application design, use a current claim token or generation number. Require that token when saving progress or completing the job, so an outdated worker cannot overwrite a newer attempt.
Show useful status such as “waiting to retry” or “summary saved.” Do not display an invented completion percentage when you cannot measure it.
4. Budget timeouts before configuring retries
Coordinate these separate budgets:
- Provider call: connection, read and overall request limits.
- Worker attempt: time allowed for one processing attempt.
- Queue reservation: when unfinished work becomes eligible for redelivery.
- Runtime: the host’s termination deadline.
- Whole job: total elapsed time allowed across attempts and waiting.
Laravel specifically recommends keeping the worker timeout several seconds shorter than retry_after; otherwise, another attempt may start while the original is still processing. SQS uses its visibility timeout instead of Laravel’s retry_after. (laravel.com)
For a hypothetical staging configuration, suppose you choose:
| Budget | Example value |
|---|---|
| Provider-call wall-clock budget | 60 seconds |
| Worker-attempt limit | 90 seconds |
| Queue reservation | 120 seconds |
| Supported runtime allowance | At least 150 seconds |
| Whole-job deadline | 10 minutes |
These are illustrative settings—not measured AI latency or recommended defaults. Validate document retrieval, preprocessing, provider waiting, result validation and database writes within the attempt budget.
Check SDK behavior too. The OpenAI Python SDK documentation checked for this article describes two automatic retries for eligible failures and a default request timeout of ten minutes. Leaving those defaults unchanged could conflict with a much shorter worker budget. (github.com)
Practical check: deliberately stall the provider fixture. Confirm which timeout fires first and whether enough time remains to record the outcome.
5. Retry selectively, with backoff and a stopping point
Inspect the error details, not just the HTTP status. OpenAI’s guidance distinguishes temporary throttling and overload from authentication, permission, invalid-request and quota problems. Waiting alone does not resolve every failure. (developers.openai.com)
Use a policy like this, adapted to the endpoint:
| Failure | Proposed response |
|---|---|
| Temporary rate limit or overload | Schedule a bounded retry |
| Connection failure or timeout | Reconcile uncertain completion, then decide |
| Invalid credentials or permissions | Stop and request an operator fix |
| Invalid document or request | Stop and explain the input problem |
| Quota or billing restriction | Pause or fail pending account action |
For temporary errors, honor Retry-After when present. If absent, use exponential backoff with jitter—a small random variation that helps avoid synchronized retries. Provider guidance also warns that unsuccessful requests can consume rate-limit capacity. (developers.openai.com)
An illustrative policy might allow three total attempts, using increasing delays and a whole-job deadline. If the provider’s minimum delay extends beyond that deadline, stop rather than retry early.
Choose which layer owns retries. Three worker attempts, each allowing an initial SDK call plus two SDK retries, could produce nine request attempts. That is a hypothetical upper bound, not nine guaranteed completed or billed generations. (github.com)
Schedule delayed retries where supported instead of occupying a worker solely to wait. Keep concurrency bounded alongside retries.
6. Prevent duplicate effects: a document-summary example
AWS documents at-least-once delivery for SQS standard queues and recommends idempotent processing. Idempotency means repeated attempts preserve the intended effect rather than creating it again. (docs.aws.amazon.com)
For this proposed summary workflow, create a stable operation identity incorporating:
- The requesting account.
- The document version or content fingerprint.
- The summarization configuration version.
- An explicit request identity when the user intentionally asks for a fresh run.
Reusing an identity should mean reusing the same intent. Do not silently reuse it for changed input. (aws.amazon.com)
Enforce uniqueness when creating the job and saving its final result. Use transactional or conditional writes rather than an unprotected “check, then insert.”
Worked interruption scenario
- The worker claims job
summary-42. - It checks that no final result exists.
- It calls the provider.
- It saves the validated summary and marks the job successful together.
- It crashes before acknowledging the queue message.
- The queue redelivers the message.
- The next worker finds the successful result and acknowledges without generating another summary.
The interruption and redelivery behavior is consistent with SQS’s documented recovery model; the completion checks are the proposed application safeguards. (docs.aws.amazon.com)
Now move the crash to after the provider finishes but before the result is saved. Your database has no completed summary. A retry may repeat the provider operation unless the endpoint offers a documented way to identify or retrieve the original result. This is the uncertainty that idempotent API designs aim to address. (aws.amazon.com)
Do not promise exactly-once AI execution. Local duplicate protection can prevent duplicate saved results without proving that only one external generation occurred. Treat notifications, publication and other side effects as separate operations needing their own repeat-safe design.
7. Handle failed jobs and cancellation explicitly
After the attempt limit or deadline, retain enough information for review: job ID, input reference, attempt history, sanitized error and provider request ID when available. OpenAI’s Python SDK exposes request IDs for troubleshooting. (github.com)
For this workflow, require an operator to identify the cause before replaying failed work. Preserve the operation identity for a retry of the same intent, and check whether a result already exists.
Implement cancellation cooperatively:
- Record
cancel_requesteddurably. - Check it before claiming work and before expensive steps.
- Check again before publishing the result.
- Define an atomic completion rule: either cancellation or success wins.
Closing the browser should not be your cancellation mechanism. Nor should a Cancel button imply that an already accepted provider operation has stopped.
Framework controls need care. Celery documents that revocation normally skips execution but does not stop a running task; its process-termination option is an administrative last resort, not something to call programmatically. (docs.celeryq.dev)
8. Test failure paths and preserve rollback options
Use synthetic documents and a fake provider in an isolated staging queue. Suggested fixtures:
| Fixture | Expected check |
|---|---|
| Same message delivered twice | One final saved result |
| Worker interrupted after result commit | Redelivery skips generation |
| Worker interrupted before result commit | Outcome is reconciled or uncertainty recorded |
| Temporary rate limit with a delay header | Retry respects the minimum delay |
| Invalid credentials | No endless retry loop |
| Cancel request during processing | Completion follows the defined atomic rule |
| Old worker resumes after losing its claim | Its completion write is rejected |
Monitor queue age, active claims, retries, timeouts, terminal failures and completed results. Where possible, correlate provider attempts and reported usage with job IDs. These are suggested operational checks, not claims about a particular host’s dashboard.
Before changing queues or payload formats, back up job records, retain configuration, version messages and confirm that rollback code can read jobs already written by the new release.
For rollback:
- Stop new dispatch to the affected queue.
- Stop new claims and handle active workers deliberately.
- Restore compatible code and configuration.
- Reconcile unfinished and uncertain jobs.
- Resume with limited concurrency.
Do not purge the queue as a routine rollback step. First determine which accepted jobs it contains and whether they can be reconstructed.
Read next: How to Stop a Staging Website from Sending Emails, Taking Payments or Running AI Jobs