Worker least-privilege cutover — runbook
One ordered cutover for the GSEC Q21 worker release. Owner decisions are
already made: Q21=A (least privilege ships) and Q22=B — everything
ships in one worker release, together with the Q20 Docker socket proxy
(setting worker.docker_proxy_mode: off → report → enforce).
The release carries three changes plus the socket proxy:
- Per-worker Vault AppRole — each worker gets its own AppRole
worker-<worker-id>bound to one shared minimal policyworker-minimal(the two KV reads the worker actually performs). Per-worker state lives in Vault KV atsecret/data/workers/<worker-id>/approle. Worker tokens never seetenants/*again — tenant secrets are brokered by the API. - Scoped, short-lived storage creds — per-task creds file
/opt/mod-worker-state/creds/task-<task-id>.json(0600, bind-mounted ro into the task container), refreshed at half-TTL. BYO-S3 connections withStorageConnection.sts_enabledget STS session creds scoped to that tenant's bucket + run prefix. - Non-root worker process —
mod-workersystem user, hardened unit (NoNewPrivileges,ProtectSystem=strict,AmbientCapabilities=CAP_CHOWN), per-worker Vault creds rendered server-side at install time.
Two numbers that govern the whole window:
- Rotate the shared core-api
secret_idONLY after every live worker reports a per-worker AppRole. Rotating earlier breaks every worker — old workers still log in with the shared pair. - A healthy task is never killed by this release. On any creds-refresh failure (API down, Vault sealed, 5xx) the worker keeps the last good creds, backs off 5m → 1h, and raises a heartbeat warning. Only an explicit stop (task cancel/fail, operator revoke) ends a task. Some tasks run for weeks — the cutover must not interrupt them.
Steps marked [OWNER-GATED] need the owner's sign-off before acting.
0. Prerequisites and gates (checklist — do these first)
-
Worker release built and staged — the single Q20+Q21 release is on the stage worker channel; old workers keep working (the
/credentialsresponse is additive-only; old Go decoders ignore the new keys). -
Landed API pieces are deployed (already merged, verify live):
StorageConnection.sts_enabledflag (per connection, defaultfalse).CORE/API/services/worker_approle.py— per-worker AppRole broker (ensure_approle/rotate_approle/revoke_approle).- Worker per-task creds file + refresh loop (half-TTL, keep-last-good, 5m→1h backoff, heartbeat warning, never-kill).
-
[OWNER-GATED] Core-api Vault policy widened. The live core-api policy must gain, before the cutover can run:
path "sys/policies/acl/worker-minimal" { capabilities = ["read"] }path "auth/approle/role/worker-*" { capabilities = ["create","read","update","delete","list"] }path "secret/data/workers/*" { capabilities = ["create","read","update","delete","list"] }This widens the API's Vault rights (it could read every worker's AppRole state) — the owner signs off on this explicitly. Verify after:
vault policy read core-api | grep -E 'worker-minimal|approle/role/worker-\*|secret/data/workers' -
Platform settings exist (registry, admin UI; no env vars):
worker.approle_per_worker_enforced(off|report|enforce),worker.sts_session_ttl_s(default7200),worker.creds_refresh_interval(default300s, delivered to workers via the heartbeattunablesmap),worker.docker_proxy_mode(off|report|enforce). -
Vault audit device enabled — the enforce step is verified from the audit log (
vault audit listmust show an enabled device). -
Baseline snapshot of the live fleet, taken at window start and recorded — this is the count you compare against later:
SELECT id, name, agent_version, last_heartbeatFROM worker_licensesWHERE last_heartbeat > now() - interval '10 minutes'ORDER BY agent_version NULLS FIRST; -
Weeks-long tasks inventoried — list running tasks that will outlive the window (
SELECT id, status FROM api_task_queue WHERE status IN ('claimed','running');or the Runs page). These are watched, not stopped. -
Known stale local backup file under
CORE/startup(not in git) contains the live sharedsecret_id. Do not open it, do not print it. The owner decides delete-or-move after the rotation step (§4); nothing in this runbook touches it before then.
1. Order of operations
Run the steps in this order. Do not skip the report-mode watch; the enforce flip and the rotation are gated on what it shows.
- Report mode — set
worker.approle_per_worker_enforced = reportandworker.docker_proxy_mode = report. The API still serves static creds and still accepts the shared AppRole; new workers get their per-worker AppRole at enrollment, socket-proxy requests are logged not blocked. - Watch 24–48 h — heartbeats, warnings, version coverage, proxy allowlist hits (section 2).
- Enforce — flip
worker.approle_per_worker_enforced = enforceandworker.docker_proxy_mode = enforce, but only after every live worker reports the new release version (section 3). - Rotate — the shared core-api
secret_idrotation, gated on full fleet coverage (section 4). This is the point of no easy return. - Cleanup — shared-role worker secret_ids deleted, backup file decision, stale per-worker state (section 5).
2. Report mode and the 24–48 h watch
Set the two settings from the admin UI (platform settings, GS.PSR group) or
via the settings API. Then watch for 24–48 hours. What "healthy" looks like:
a) Version coverage. Every worker whose last_heartbeat is within the
heartbeat TTL must report the new release version on agent_version. Old
workers are allowed during report mode — they keep static creds and the
shared AppRole — but the enforce flip (step 3) and the rotation (step 4) are
blocked until coverage is 100% of live workers. Re-run the baseline query from
section 0 and compare; anything still on an old agent_version is a worker
that must be upgraded or retired before the flip.
b) Per-worker AppRole state. For each live worker, confirm the broker created its role and KV state:
vault read auth/approle/role/worker-<worker-id>
vault read secret/data/workers/<worker-id>/approle
Both must exist for every live worker before step 4. worker-minimal must be
attached to each role.
c) Heartbeat warnings, not kills. The creds refresher raises a heartbeat warning on any refresh failure and keeps the last good creds. During the watch you want warnings to be transient (API blip, Vault seal window) and to clear on the next successful refresh at half-TTL. A warning that persists past the backoff cap (1 h) on a task means the task is running on stale creds — that is acceptable behaviour, not a failure; investigate the cause (API down, Vault sealed, tenant STS endpoint not actually supporting STS — the API falls back to the static key with a warning and the task keeps running).
d) Socket-proxy report hits. worker.docker_proxy_mode = report logs
every proxied Docker API call against the GPU-passthrough allowlist. Any hit
outside the allowlist is a worker image that needs an allowlist entry before
enforce — fix the allowlist during the watch, not after the flip.
Rollback for this step: set both settings back to off. The API serves
static creds and accepts the shared AppRole again; per-worker roles that were
already created are harmless (they can stay or be revoked individually). No
running task is affected by flipping back — the refresh loop keeps whatever
creds it has.
3. Enforce (after 100% version coverage)
Gate: the baseline query shows every live worker on the new version. Then:
- Flip
worker.approle_per_worker_enforced = enforce. New enrollments and install-script rendering now require the per-worker AppRole; the shared role is no longer used by workers. - Flip
worker.docker_proxy_mode = enforce(allowlist verified clean during the watch). - Verify zero shared-role logins from workers:
vault audit list | grep -E 'auth/approle/role/(core-api|worker-)' | tail -50
Worker logins must show worker-<worker-id> roles only. A shared-role login
after the flip means a straggler worker — upgrade it, do not re-open the
shared role.
Rollback for this step: flip both settings back to report (or off).
The shared core-api role is still alive at this point (rotation has not
happened yet), so old-style workers reconnect on their existing env files
without re-enrollment. This is why rotation comes AFTER enforce, never before.
4. Rotate the shared core-api secret_id (hard gate)
Gate — do not start until every live worker reports a per-worker AppRole (section 2b: role + KV state exist for every live worker, and the baseline query shows 100% new-version coverage). Rotating earlier breaks every worker that still logs in with the shared pair.
-
Rotate the shared core-api
secret_idin Vault (new value; the leaked one dies with this step). The API picks it up through its own rotation path; workers no longer use it. -
Verify the old shared
secret_idis rejected:vault audit list | grep 'auth/approle' | tail -20 # no worker logins on core-api -
Confirm the fleet is healthy: every worker heartbeats on its
worker-<worker-id>role, no new warnings in the heartbeat alert field.
Rollback for this step: rotation is non-destructive by construction for
per-worker roles (each worker validates a new secret before discarding the
old; delivery failure leaves the old valid). If the release must be pulled
after rotation, flip worker.approle_per_worker_enforced back to off and
re-create the shared role on demand with a fresh secret_id — the leaked
one stays dead. Do not restore the old secret_id under any circumstance.
5. Cleanup
-
Delete the shared core-api role's worker-facing
secret_ids — the API keeps its own. Only workers are cut off; the API's broker role is untouched. -
[OWNER-GATED] The stale local backup file under
CORE/startup(not in git, contains the live sharedsecret_id): after rotation the owner decides delete or move. Do not open or print its contents; do not act before the rotation is verified. -
Per-worker state hygiene: worker deactivation/removal revokes all tokens of that role and deletes the AppRole automatically — spot-check one removed worker:
vault read auth/approle/role/worker-<worker-id> # must 404 after removal -
Repo grep for
secret_idremnants must show nothing beyond placeholder examples.
6. What to watch for weeks-long tasks
Tasks that started before the cutover keep running through it:
- The per-task creds file (
/opt/mod-worker-state/creds/task-<task-id>.json) is refreshed at half-TTL;MINIO_*env vars remain the bootstrap value, so existing node images keep working byte-for-byte. Opt-in images re-read the file on S3 errors or on a schedule. - BYO-STS sessions (TTL
worker.sts_session_ttl_s, default 7200 s) are re-minted per refresh, scoped to that tenant's bucket + run prefix. House versitygw creds have no TTL — the refresh is a cheap re-validation returning the same stable key (a no-op swap). - Never kill a healthy task. API down or Vault sealed → the task runs on last-good creds indefinitely with a heartbeat warning; when the API recovers the next refresh re-aligns the file. Only explicit stop ends a task (cancel/fail 409/410, task no longer claimed/running, operator revoke).
- Watch the inventoried long tasks (section 0) for the full 24–48 h: warnings must clear on the next successful refresh; a warning surviving past the 1 h backoff cap on a long task is tolerated behaviour, not a cutover failure — investigate the cause, do not restart the task.
- Non-root workers on WSL hosts: verify the
/usr/lib/wsl/libread-bind and/dev/dxgGPU passthrough on the first long task after the release; native tasks fall back to running as the worker user ifCAP_CHOWNis dropped.
7. Cutover sign-off
Done when: 100% live-worker version coverage, per-worker AppRole for every
live worker, zero shared-role worker logins in the audit log after enforce,
old shared secret_id rejected, fleet healthy 24 h, backup-file decision
recorded by the owner, and every inventoried weeks-long task still running.