The outcome

Deliver generated results reliably without blocking requests or losing short-lived output files.

Step by step

A workflow you can repeat.

  1. 01

    Create predictions server-side, store the prediction ID and owner, and return a durable application job ID immediately.

  2. 02

    Choose polling, server-sent events, or a webhook according to duration and scale, and authenticate access to job status.

  3. 03

    Make webhook handling idempotent, validate the event and expected prediction, and tolerate repeated or out-of-order updates.

  4. 04

    When a job succeeds, copy required output files to controlled persistent storage before Replicate's API retention window ends.

  5. 05

    Handle cancellation, failure, timeout, retry, moderation, and cleanup states, then test each path with observable metrics.

Working standard

What good use looks like.

  • Persist prediction IDs.
  • Make callbacks idempotent.
  • Copy required files before expiry.

Official references

Check the current product documentation.