The outcome

Keep a multi-user cluster isolated, observable, recoverable, and resistant to first-run takeover or privilege expansion.

Step by step

A workflow you can repeat.

  1. 01

    Define cluster topology, trusted networks, user groups, OIDC or local identity, permission matrix, model sources, cache and virtual-environment policy, secrets, audit retention, SLO, and backup ownership.

  2. 02

    Initialize authentication on a private endpoint, provide every API process consistent JWT, encryption and database state, restrict supervisor and worker traffic, and separate admin, operator and consumer roles.

  3. 03

    Allowlist model sources and engines, scan custom environments, protect caches and logs, cap actor recovery, replicas and resources, and expose public inference only through authenticated TLS ingress.

  4. 04

    Test first-run race prevention, login and key expiry, privilege boundaries, OIDC failure, worker loss, actor crash loops, cache corruption, malicious model files, network partition, and restore.

  5. 05

    Canary version and auth migrations, back up and restore auth state, rotate secrets coherently, audit users and keys, remove deprecated scopes, and verify retired nodes and caches are inaccessible.

Working standard

What good use looks like.

  • Complete admin setup before exposure.
  • Share consistent auth state across API processes.
  • Cap crash recovery loops.

Official references

Check the current product documentation.