Skip to main content

Worker least-privilege cutover — runbook

One ordered cutover for the GSEC Q21 worker release. Owner decisions are already made: Q21=A (least privilege ships) and Q22=B — everything ships in one worker release, together with the Q20 Docker socket proxy (setting worker.docker_proxy_mode: off → report → enforce).

The release carries three changes plus the socket proxy:

  • Per-worker Vault AppRole — each worker gets its own AppRole worker-<worker-id> bound to one shared minimal policy worker-minimal (the two KV reads the worker actually performs). Per-worker state lives in Vault KV at secret/data/workers/<worker-id>/approle. Worker tokens never see tenants/* again — tenant secrets are brokered by the API.
  • Scoped, short-lived storage creds — per-task creds file /opt/mod-worker-state/creds/task-<task-id>.json (0600, bind-mounted ro into the task container), refreshed at half-TTL. BYO-S3 connections with StorageConnection.sts_enabled get STS session creds scoped to that tenant's bucket + run prefix.
  • Non-root worker process — mod-worker system user, hardened unit (NoNewPrivileges, ProtectSystem=strict, AmbientCapabilities=CAP_CHOWN), per-worker Vault creds rendered server-side at install time.

Two numbers that govern the whole window:

  • Rotate the shared core-api secret_id ONLY after every live worker reports a per-worker AppRole. Rotating earlier breaks every worker — old workers still log in with the shared pair.
  • A healthy task is never killed by this release. On any creds-refresh failure (API down, Vault sealed, 5xx) the worker keeps the last good creds, backs off 5m → 1h, and raises a heartbeat warning. Only an explicit stop (task cancel/fail, operator revoke) ends a task. Some tasks run for weeks — the cutover must not interrupt them.

Steps marked [OWNER-GATED] need the owner's sign-off before acting.

0. Prerequisites and gates (checklist — do these first)​

  • Worker release built and staged — the single Q20+Q21 release is on the stage worker channel; old workers keep working (the /credentials response is additive-only; old Go decoders ignore the new keys).

  • Landed API pieces are deployed (already merged, verify live):

    • StorageConnection.sts_enabled flag (per connection, default false).
    • CORE/API/services/worker_approle.py — per-worker AppRole broker (ensure_approle / rotate_approle / revoke_approle).
    • Worker per-task creds file + refresh loop (half-TTL, keep-last-good, 5m→1h backoff, heartbeat warning, never-kill).
  • [OWNER-GATED] Core-api Vault policy widened. The live core-api policy must gain, before the cutover can run:

    path "sys/policies/acl/worker-minimal" { capabilities = ["read"] }
    path "auth/approle/role/worker-*" { capabilities = ["create","read","update","delete","list"] }
    path "secret/data/workers/*" { capabilities = ["create","read","update","delete","list"] }

    This widens the API's Vault rights (it could read every worker's AppRole state) — the owner signs off on this explicitly. Verify after:

    vault policy read core-api | grep -E 'worker-minimal|approle/role/worker-\*|secret/data/workers'
  • Platform settings exist (registry, admin UI; no env vars): worker.approle_per_worker_enforced (off|report|enforce), worker.sts_session_ttl_s (default 7200), worker.creds_refresh_interval (default 300 s, delivered to workers via the heartbeat tunables map), worker.docker_proxy_mode (off|report|enforce).

  • Vault audit device enabled — the enforce step is verified from the audit log (vault audit list must show an enabled device).

  • Baseline snapshot of the live fleet, taken at window start and recorded — this is the count you compare against later:

    SELECT id, name, agent_version, last_heartbeat
    FROM worker_licenses
    WHERE last_heartbeat > now() - interval '10 minutes'
    ORDER BY agent_version NULLS FIRST;
  • Weeks-long tasks inventoried — list running tasks that will outlive the window (SELECT id, status FROM api_task_queue WHERE status IN ('claimed','running'); or the Runs page). These are watched, not stopped.

  • Known stale local backup file under CORE/startup (not in git) contains the live shared secret_id. Do not open it, do not print it. The owner decides delete-or-move after the rotation step (§4); nothing in this runbook touches it before then.

1. Order of operations​

Run the steps in this order. Do not skip the report-mode watch; the enforce flip and the rotation are gated on what it shows.

  1. Report mode — set worker.approle_per_worker_enforced = report and worker.docker_proxy_mode = report. The API still serves static creds and still accepts the shared AppRole; new workers get their per-worker AppRole at enrollment, socket-proxy requests are logged not blocked.
  2. Watch 24–48 h — heartbeats, warnings, version coverage, proxy allowlist hits (section 2).
  3. Enforce — flip worker.approle_per_worker_enforced = enforce and worker.docker_proxy_mode = enforce, but only after every live worker reports the new release version (section 3).
  4. Rotate — the shared core-api secret_id rotation, gated on full fleet coverage (section 4). This is the point of no easy return.
  5. Cleanup — shared-role worker secret_ids deleted, backup file decision, stale per-worker state (section 5).

2. Report mode and the 24–48 h watch​

Set the two settings from the admin UI (platform settings, GS.PSR group) or via the settings API. Then watch for 24–48 hours. What "healthy" looks like:

a) Version coverage. Every worker whose last_heartbeat is within the heartbeat TTL must report the new release version on agent_version. Old workers are allowed during report mode — they keep static creds and the shared AppRole — but the enforce flip (step 3) and the rotation (step 4) are blocked until coverage is 100% of live workers. Re-run the baseline query from section 0 and compare; anything still on an old agent_version is a worker that must be upgraded or retired before the flip.

b) Per-worker AppRole state. For each live worker, confirm the broker created its role and KV state:

vault read auth/approle/role/worker-<worker-id>
vault read secret/data/workers/<worker-id>/approle

Both must exist for every live worker before step 4. worker-minimal must be attached to each role.

c) Heartbeat warnings, not kills. The creds refresher raises a heartbeat warning on any refresh failure and keeps the last good creds. During the watch you want warnings to be transient (API blip, Vault seal window) and to clear on the next successful refresh at half-TTL. A warning that persists past the backoff cap (1 h) on a task means the task is running on stale creds — that is acceptable behaviour, not a failure; investigate the cause (API down, Vault sealed, tenant STS endpoint not actually supporting STS — the API falls back to the static key with a warning and the task keeps running).

d) Socket-proxy report hits. worker.docker_proxy_mode = report logs every proxied Docker API call against the GPU-passthrough allowlist. Any hit outside the allowlist is a worker image that needs an allowlist entry before enforce — fix the allowlist during the watch, not after the flip.

Rollback for this step: set both settings back to off. The API serves static creds and accepts the shared AppRole again; per-worker roles that were already created are harmless (they can stay or be revoked individually). No running task is affected by flipping back — the refresh loop keeps whatever creds it has.

3. Enforce (after 100% version coverage)​

Gate: the baseline query shows every live worker on the new version. Then:

  1. Flip worker.approle_per_worker_enforced = enforce. New enrollments and install-script rendering now require the per-worker AppRole; the shared role is no longer used by workers.
  2. Flip worker.docker_proxy_mode = enforce (allowlist verified clean during the watch).
  3. Verify zero shared-role logins from workers:
vault audit list | grep -E 'auth/approle/role/(core-api|worker-)' | tail -50

Worker logins must show worker-<worker-id> roles only. A shared-role login after the flip means a straggler worker — upgrade it, do not re-open the shared role.

Rollback for this step: flip both settings back to report (or off). The shared core-api role is still alive at this point (rotation has not happened yet), so old-style workers reconnect on their existing env files without re-enrollment. This is why rotation comes AFTER enforce, never before.

4. Rotate the shared core-api secret_id (hard gate)​

Gate — do not start until every live worker reports a per-worker AppRole (section 2b: role + KV state exist for every live worker, and the baseline query shows 100% new-version coverage). Rotating earlier breaks every worker that still logs in with the shared pair.

  1. Rotate the shared core-api secret_id in Vault (new value; the leaked one dies with this step). The API picks it up through its own rotation path; workers no longer use it.

  2. Verify the old shared secret_id is rejected:

    vault audit list | grep 'auth/approle' | tail -20 # no worker logins on core-api
  3. Confirm the fleet is healthy: every worker heartbeats on its worker-<worker-id> role, no new warnings in the heartbeat alert field.

Rollback for this step: rotation is non-destructive by construction for per-worker roles (each worker validates a new secret before discarding the old; delivery failure leaves the old valid). If the release must be pulled after rotation, flip worker.approle_per_worker_enforced back to off and re-create the shared role on demand with a fresh secret_id — the leaked one stays dead. Do not restore the old secret_id under any circumstance.

5. Cleanup​

  • Delete the shared core-api role's worker-facing secret_ids — the API keeps its own. Only workers are cut off; the API's broker role is untouched.

  • [OWNER-GATED] The stale local backup file under CORE/startup (not in git, contains the live shared secret_id): after rotation the owner decides delete or move. Do not open or print its contents; do not act before the rotation is verified.

  • Per-worker state hygiene: worker deactivation/removal revokes all tokens of that role and deletes the AppRole automatically — spot-check one removed worker:

    vault read auth/approle/role/worker-<worker-id> # must 404 after removal
  • Repo grep for secret_id remnants must show nothing beyond placeholder examples.

6. What to watch for weeks-long tasks​

Tasks that started before the cutover keep running through it:

  • The per-task creds file (/opt/mod-worker-state/creds/task-<task-id>.json) is refreshed at half-TTL; MINIO_* env vars remain the bootstrap value, so existing node images keep working byte-for-byte. Opt-in images re-read the file on S3 errors or on a schedule.
  • BYO-STS sessions (TTL worker.sts_session_ttl_s, default 7200 s) are re-minted per refresh, scoped to that tenant's bucket + run prefix. House versitygw creds have no TTL — the refresh is a cheap re-validation returning the same stable key (a no-op swap).
  • Never kill a healthy task. API down or Vault sealed → the task runs on last-good creds indefinitely with a heartbeat warning; when the API recovers the next refresh re-aligns the file. Only explicit stop ends a task (cancel/fail 409/410, task no longer claimed/running, operator revoke).
  • Watch the inventoried long tasks (section 0) for the full 24–48 h: warnings must clear on the next successful refresh; a warning surviving past the 1 h backoff cap on a long task is tolerated behaviour, not a cutover failure — investigate the cause, do not restart the task.
  • Non-root workers on WSL hosts: verify the /usr/lib/wsl/lib read-bind and /dev/dxg GPU passthrough on the first long task after the release; native tasks fall back to running as the worker user if CAP_CHOWN is dropped.

7. Cutover sign-off​

Done when: 100% live-worker version coverage, per-worker AppRole for every live worker, zero shared-role worker logins in the audit log after enforce, old shared secret_id rejected, fleet healthy 24 h, backup-file decision recorded by the owner, and every inventoried weeks-long task still running.