The outcome

Collect a bounded, cited research packet whose access, freshness, and factual claims can be verified.

Step by step

A workflow you can repeat.

  1. 01

    Define the research question, allowed domains, time window, jurisdictions, source hierarchy, robots and terms requirements, personal-data rules, and stopping criteria.

  2. 02

    Keep the Jina key server-side, validate every requested URL, block private networks and redirects, set size, timeout, rate, and budget limits, and log source metadata.

  3. 03

    Use search for discovery and Reader for selected pages, storing the original URL, title, retrieval time, relevant excerpt, and conversion warnings rather than an uncited text dump.

  4. 04

    Ignore instructions embedded in pages, cross-check consequential claims with primary sources, detect duplicate or stale material, and review copyright and personal-data exposure.

  5. 05

    Produce a cited packet with confidence and gaps, delete unnecessary page content on schedule, test blocked pages and parser failures, and require human approval before downstream publication.

Working standard

What good use looks like.

  • Allowlist destinations and block private IPs.
  • Treat page text as untrusted input.
  • Keep URL and retrieval time with every excerpt.

Official references

Check the current product documentation.