The outcome
Deliver generated results reliably without blocking requests or losing short-lived output files.
Step by step
A workflow you can repeat.
- 01
Create predictions server-side, store the prediction ID and owner, and return a durable application job ID immediately.
- 02
Choose polling, server-sent events, or a webhook according to duration and scale, and authenticate access to job status.
- 03
Make webhook handling idempotent, validate the event and expected prediction, and tolerate repeated or out-of-order updates.
- 04
When a job succeeds, copy required output files to controlled persistent storage before Replicate's API retention window ends.
- 05
Handle cancellation, failure, timeout, retry, moderation, and cleanup states, then test each path with observable metrics.
Working standard
What good use looks like.
- Persist prediction IDs.
- Make callbacks idempotent.
- Copy required files before expiry.
Official references