Skip to main content

Owner restart window — full checklist

One drained window, in order. Every step is plain about what it does, what to expect, how to verify, and how to roll it back. The two big safety numbers:

  • API down > 120 s makes every worker fence its external GPU stacks (compose-stops ha_class=external stacks). Keep the API-down span (step 8a→healthy) under 120 s, or expect the fences and repair them in 8l.
  • Drain before you restart, and check before, not after. The in-flight count captured at pause time is the one you wait on.

1. Order of operations + estimated duration​

#StepEst.Tim needed
APre-checks (read-only) — §45 min—
BRollback tags — §51 min—
CCompose edits: FORWARDED_ALLOW_IPS + pin nginx IP — §65 min—
DDrain + wait for in-flight to hit 0 — §70–15 minpause agent lanes
E8a API recreate (image + env) — fence clock starts40–120 s—
F8b Keycloak recreate (new theme)30–90 s—
G8c DB-TLS server side (Postgres + PgBouncer)per runbook—
H8d PgBouncer DNS-stranding check15–30 s—
I8e nginx recreate (CORS map)10–20 s—
J8f Smoke checks10–15 minbrowser: login banner + admin login + PDF ingest
K8g Vault Shamir live rekey (Q9=C)5–10 minowner: shares present, 3-of-5 for verify
L8h mesh.acl_enforced preview + flip (R-012)2 min—
M8i Redeploy stacks 34 → 31 → 30 → 29 (GF.C0b.15; first redeploy pushes the ACL policy)20–40 min—
N8j Q21 setup: core-api policy grant + report-mode settings + baseline5–10 minsign-off on the policy grant (OWNER-GATED)
O8k Setting flips: privacy.erasure_job_enabled, gsec.transport.internal_tls_enforce=report2 min—
P8l Resume dispatch, un-pause lanes, post-window fence check5 min—

2. Tim needed for (everything else is agent-runnable)​

  1. Vault Shamir ceremony (8k… 8g): 5 shares written 0600 per holder; 3 of them needed on hand for the verify seal+unseal. After success the share files are distributed to holders and deleted from this box.
  2. Browser checks (8f): the AC-8 use-notice banner on the login page (proves the new theme), login as a normal user and as platform admin (the admin token is needed for every flip in 8h/8j/8k), and one PDF ingest.
  3. Q21 core-api Vault policy grant (8j): it widens the API's Vault rights (it can then read every worker's AppRole state) — OWNER-GATED sign-off before the command runs.
  4. Stack 29: it is torn_down (not just stopped). Confirm whether POST /v1/stacks/29/start is the intended path or a full POST /v1/stacks/29/deploy is needed (fallback in 8i).

3. Candidate images + freshness verdict​

Built from HEAD 76571ff82. Verified before this checklist was written:

ComponentCandidate tagVerified
APImod-api:window-candidate-76571ff82sentry_sdk import → None (SDK removed in dcd147f14); full import main resolves (stops only at the live-DB create_tables(), as it must)
Keycloakmod-keycloak:window-candidate-76571ff82 (= harbor.modtechlabs.com/mod_core/keycloak:window-candidate-76571ff82)template.ftl inside the image's theme JAR contains modtechUseNotice ×3 (the AC-8 login banner)

Do not reuse mod-api:fx18-candidate or mod-api:latest — they still carry the old requirements.

Freshness (mtime evidence; git log was not available to the prepping agent):

  • CORE/API/requirements.txt — mtime 10-03 23:09 CDT, which is before the candidate build (10-04 02:17 CDT). Requirements are current in the image.
  • CORE/API/Dockerfile — mtime 10-04 17:21 CDT, which is after the build. mtime drift on the 9p /mnt/c install dir is possible, so the lead must confirm what changed in the Dockerfile since 02:17. The app code (COPY . .) does not matter — the API code is bind-mounted live and the image only carries the base image, apt system deps, and pip packages.
  • Rebuild verdict: NOT required unless the post-build Dockerfile change touched an apt-get/pip layer (system deps). If it did, rebuild the API candidate before step 8a; otherwise proceed with the existing candidate.

4. Pre-checks (read-only — do these first, ~5 min)​

  1. Confirm the candidates exist locally (they are local images; do not pull):

    docker image inspect mod-api:window-candidate-76571ff82 --format '{{.Id}} {{.Size}}'
    docker image inspect mod-keycloak:window-candidate-76571ff82 --format '{{.Id}} {{.Size}}'

    Preflight 10-04 20:30 CDT: both present (API bac209b6478d, Keycloak 3299ef41cbf2).

  2. Confirm all containers are up:

    docker ps --format '{{.Names}}\t{{.Status}}' | grep -E "my_fastapi_app|my_keycloak_server|mod_pgbouncer|my_postgres_db|my_nginx_proxy|my_vault"

    Preflight 10-04: all 8 healthy (api, keycloak, pgbouncer, postgres, nginx, vault, vault-agent, vault-unseal).

  3. Confirm there is no in-flight work. Two ways, use both:

    • API (preferred): as a platform admin, GET /v1/system/drain-status → expect "in_flight_runs": 0, "in_flight_tasks": 0.
    • DB (cross-check): the drain endpoint counts WorkflowRun rows in a non-terminal state — QUEUED, RUNNING, PENDING_ACTION (Postgres enum, uppercase) — minus stale ones (no NodeRun progress in 24 h), plus api_task_queue rows in PENDING/ASSIGNED/CLAIMED/ RUNNING.
    • Preflight 10-04: in_flight_runs 0, in_flight_tasks 0 (25 non-terminal run rows exist but are all stale — newest NodeRun progress 09-28, so the drain endpoint excludes them). The stale rows are a separate cleanup problem, not a window blocker; if a stale row ever blocks ready_for_restart, it means a run with recent progress appeared — stop it deliberately and re-check.
  4. Note the current live images so rollback is a single retag + recreate:

    docker inspect my_fastapi_app --format '{{.Config.Image}} {{.Image}}'
    docker inspect my_keycloak_server --format '{{.Config.Image}} {{.Image}}'
  5. Note current platform settings (read-only; the window changes three of them): mesh.acl_enforced (preflight: False), privacy.erasure_job_enabled (preflight: False), gsec.transport.internal_tls_enforce (preflight: off). Also note the current net.trusted_proxy_cidrs (step 6).

  6. Dell gateway baseline (used in 8l):

    curl -s -m 10 http://192.168.1.210:8080/workers | grep -vE "Bearer|authorization|access_token|eyJ|password|secret|token"
    # expect total 2 healthy

5. Rollback tags — create these FIRST (before any swap)​

Retag the current running images so that "go back" is one command each. This must happen before the candidate is re-tagged over the live name, or the live name will point at the new image and you will have no clean rollback.

DATE=$(date +%Y%m%d%H%M)

# API (live container runs image "mod-api" == mod-api:latest)
docker tag mod-api:latest mod-api:rollback-${DATE}

# Keycloak (live container runs the public image)
docker tag keycloak/keycloak:26.3 keycloak/keycloak:rollback-${DATE}

Verify the tags exist before moving on:

docker image inspect mod-api:rollback-${DATE} --format '{{.Id}}'
docker image inspect keycloak/keycloak:rollback-${DATE} --format '{{.Id}}'

6. Compose edits: FORWARDED_ALLOW_IPS (Q5 / R-011) + pin nginx IP​

Where it is set today: CORE/startup/docker-compose.yml, in the api service (line ~1569) and api-background service (line ~1774), both as - FORWARDED_ALLOW_IPS=*.

What it does: the API's entrypoint.sh runs uvicorn with no explicit --forwarded-allow-ips, so with FORWARDED_ALLOW_IPS=* the proxy-headers middleware trusts X-Forwarded-For from any peer — a direct client can spoof the client IP uvicorn records.

Edit (in docker-compose.yml, before step 8a so the recreate picks it up):

  1. Give the my_nginx_proxy service a static ipv4_address under its modtex-network block: 172.24.0.16 — its current dynamic allocation (confirmed 10-04 preflight). Pinning it means the allowed value stays stable across recreates.

  2. Change both FORWARDED_ALLOW_IPS=* lines to:

    - FORWARDED_ALLOW_IPS=127.0.0.1,172.24.0.16

    If you pin a different static IP, use that IP here instead.

Do NOT allow the whole subnet (172.24.0.0/16) — every other container on that bridge (vault, pgbouncer, keycloak, …) could then spoof XFF.

Complementary app-level control (no restart, ~5 s trust-cache expiry):

curl -s -X PATCH "https://<api-host>/v1/platform/settings/net.trusted_proxy_cidrs" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \
-d '{"value":"127.0.0.1/32,172.24.0.16/32"}'

Expected result: after the 8a recreate, a direct (non-nginx) request to the API with a spoofed X-Forwarded-For header no longer changes the recorded client IP; nginx-proxied requests keep the real client IP.

Rollback (one command each): revert the two compose edits (git diff / your editor) + docker compose up -d --force-recreate api api-background; and PATCH net.trusted_proxy_cidrs back to its previous value.

7. Drain (stop dispatching, then wait) — FIRST, before any recreate​

What it does: pauses the dispatch layer so no new runs/tasks start, then waits for the work already in flight to finish. The platform stays drained for the whole window, so users see queueing, not errors.

# 1) Pause dispatch (idempotent), as platform admin:
curl -s -X POST "https://<api-host>/v1/system/drain" -H "Authorization: Bearer $ADMIN_TOKEN"
# 2) Record the in-flight counts NOW (captured at pause time — this is the
# number you wait on):
curl -s "https://<api-host>/v1/system/drain-status" -H "Authorization: Bearer $ADMIN_TOKEN"
# 3) Poll drain-status until in_flight_runs == 0 AND in_flight_tasks == 0.
# Do NOT re-check after you have begun recreating containers.

The 120 s fence, and how we handle it. API down > 120 s makes every worker compose-stop its ha_class=external GPU stacks. Two parts:

  • Before the window: pause the agent lanes (they ride the site tailnet and stack endpoints; Tim does this at ~22:00). Weeks-long tasks on workers keep running through the window (they are not API-dispatched); the drain only stops new dispatch.

  • During the window: keep the span from docker stop my_fastapi_app (inside the 8a recreate) to curl -fsS http://127.0.0.1:8000/health passing under 120 s. The DB-TLS Postgres recreate (8c) can extend the API-down span — order it so the API is recreated last among the DB-dependent services, or accept the fences and use the repair below.

  • After the window (8l): if fences happened, check the Dell gateway:

    curl -s -m 10 http://192.168.1.210:8080/workers | grep -vE "Bearer|authorization|access_token|eyJ|password|secret|token"
    # expect total 2 healthy

    If either worker is down / unhealthy, or a stack shows stopped that was running, repair each via the platform stack task path: POST /v1/stacks/{id}/stop then POST /v1/stacks/{id}/start (the same path as step 8i). Do not hand-poke containers on the worker boxes.

8. The window (ordered)​

All docker compose commands run from CORE/startup. The DB-TLS compose override is db-tls/docker-compose.db-tls.yml. Set:

C="docker compose -f docker-compose.yml -f db-tls/docker-compose.db-tls.yml"

8a. API recreate — image switch + FORWARDED_ALLOW_IPS + DB-TLS client args. The fence clock starts when this container stops.

The api service has no pinned image: — the stack builds it as mod-api and the live container runs mod-api (:latest). Retag the candidate over the live name (rollback tag already saved in §5), then recreate. api-background and mod-migrate share the image and come along.

docker tag mod-api:window-candidate-76571ff82 mod-api:latest
$C --profile all up -d --force-recreate api api-background
# env (FORWARDED_ALLOW_IPS, POSTGRES_SSLMODE, etc.) is injected at
# container CREATE from the compose env_file — the recreate is what applies it.
curl -fsS http://127.0.0.1:8000/health # fence clock stops here

Note the API code is bind-mounted — the image swap only updates system deps/requirements (see §3); the app code served is what is on disk right now.

8b. Keycloak recreate (new login theme).

The keycloak service runs keycloak/keycloak:26.3; the live compose bind-mounts ./keycloak-themes/modtech, which only matters in dev (non-optimized) mode. The candidate bakes the theme as a JAR provider in /opt/keycloak/providers/ (the only theme form read in optimized mode).

docker tag mod-keycloak:window-candidate-76571ff82 keycloak/keycloak:26.3
$C --profile all up -d --force-recreate keycloak

Verify: login page shows the AC-8 use-notice banner (Tim, in 8f).

Theme source — decide one, not both. If you want the baked image to be the source of truth (the intent here), drop the ./keycloak-themes/... volume line when recreating, or run the service in optimized mode. If you leave the bind mount in, it overlays the baked theme with the live dir — same template, harmless, but then the "baked" image is not what is actually serving.

8c. DB-TLS server side (Postgres + PgBouncer).

Do exactly what the DB-TLS runbook says — do not re-do it here. It is the single source of truth for the SCRAM re-hash, MOD_DB_TLS flag, Postgres recreate, PgBouncer recreate, and the DNS-stranding check.

→ DB TLS + SCRAM switch-over window

On the API side you already added the client TLS env (POSTGRES_SSLMODE / POSTGRES_SSLROOTCERT) in 8a's recreate.

8d. PgBouncer DNS-stranding check (mandatory after the Postgres recreate).

Postgres was recreated under the same service name; PgBouncer caches DNS lookups, and a failed lookup cached during the recreation window strands the pooler at the old address (known failure mode — see the pgbouncer DNS-stranding notes). Even if nothing looks wrong, bounce it:

docker ps --filter name=my_postgres_db # confirm healthy FIRST
docker restart mod_pgbouncer
docker ps --filter name=mod_pgbouncer # wait: healthy (its healthcheck
# probes postgres-app — that is the canary)

Also confirm the API itself can reach the DB (the 8e liveness + the login in 8f cover it; a Vault-credential DB round-trip failing after this step = pgbouncer stranding, not Vault).

8e. nginx recreate — pick up the CORS map change.

The CORS map lives in the generated /etc/nginx/nginx.conf: map $http_origin $cors_origin, rendered at container start by generate_cors_map() in CORE/startup/scripts/nginx-entrypoint.sh from CORS_ALLOWED_ORIGINS (auto mode: dashboard, workflows, api + — since the login-page RUM-beacon change — the Keycloak origin). Template source: CORE/startup/nginx.conf.template (map at line ~183, headers at lines ~447/474). Container: my_nginx_proxy (service nginx, compose CORE/startup/docker-compose.yml).

Preflight 10-04: the live map has only dashboard/workflows/api — no Keycloak origin — while the on-disk entrypoint (updated 10-04) emits it in auto mode with KEYCLOAK_HOSTNAME set (it is: dev-keycloak.modtechlabs.com). So the recreate is what fixes the login-page beacon being CORS-blocked.

docker compose up -d --no-deps --force-recreate nginx
# verify the new map rendered:
docker exec my_nginx_proxy sed -n '/map $http_origin \$cors_origin/,/^ }/p' /etc/nginx/nginx.conf
# expect the three domain lines PLUS a dev-keycloak.modtechlabs.com line

Expected result: Access-Control-Allow-Origin is returned for the Keycloak origin on API responses (test: OPTIONS preflight from the login page works; regression guard: bash CORE/startup/tests/test_nginx_cors_map.sh — host-side, no container involved).

Rollback: docker compose up -d --no-deps --force-recreate nginx again after reverting scripts/nginx-entrypoint.sh (the entrypoint is bind-mounted; it re-runs on recreate). No image rebuild is involved.

8f. Smoke checks (before any setting flip or rekey).

Run in order; stop and roll back the failing component if any fails.

# 1) liveness
curl -fsS http://127.0.0.1:8000/health # {"status":"ok"}
# 2) deep health (platform admin) — expect vault_sealed == false
# GET /health?extended=true (or the alias /platform/health/extended)
# 3) (TIM) login works against the NEW Keycloak image — AC-8 banner
# renders on the login page; log in as a normal user AND as a platform
# admin (the admin token is needed for every step below).
# 4) one real workflow run, end to end (POST /v1/workflows/{id}/run →
# watch GET /v1/workflows/{id}/runs to success).
# 5) (TIM) one PDF ingest — lands in object store + index row.
# 6) an audit row is written — psql: latest audit_events row from the
# actions above. The ISO audit trail lives in Postgres, not stdout.

8g. Vault Shamir live rekey (Q9=C) — owner present for the shares.

What it does: rekeys the single-key auto-unseal Vault to 5 shares / threshold 3 with manual unseal. Each share is written to its own 0600 file in the secure state dir (default /var/lib/mod/vault-shares/, fallback $HOME/.local/state/mod/vault-shares/), verified by a seal + unseal ceremony, and the old auto-unseal key is renamed, never deleted (unseal-key.retired-<ts>) — which also stops the auto-unseal watcher.

Script: CORE/startup/scripts/vault-rekey-to-shamir.sh (unit-tested: CORE/startup/tests/test_vault_shamir.sh). It never prints the old key, the new shares, or any token.

9p note (shares are 0600 on Linux only): the /mnt/c install dir is 9p — chmod 600 there is not a real permission boundary and share files must never live in the repo. The script deliberately writes to a native Linux path (MOD_SECURE_STATE_DIR, default /var/lib/mod), so run it from the native-Linux side of the box, not from a /mnt/c shell. The rekey itself runs inside the my_vault container; shares are docker cp'd out to the Linux dir, then the temp file is shredded in both places.

# 1) Plan (safe, default — prints seal type / current shares only):
bash CORE/startup/scripts/vault-rekey-to-shamir.sh
# 2) Live rekey (platform already drained in §7; VAULT_TOKEN enables the
# seal+unseal verify — strongly recommended with the owner present):
VAULT_TOKEN=<root-token> bash CORE/startup/scripts/vault-rekey-to-shamir.sh --apply --confirm

Expected result: 5 share files share-01..share-05 (0600) in the secure dir, vault status shows Seal Type: shamir (5 shares, threshold 3), old key renamed to unseal-key.retired-<ts>, verify step reports "3 new shares unsealed Vault".

Verify: script's own ceremony when VAULT_TOKEN is set. Without it, Tim seals (vault operator seal) and unseals with any 3 of the 5 share files before anyone relies on the retired old key. Distribute the 5 shares to holders, then delete them from this box (holders keep their copies).

Rollback (one command): docker exec my_vault sh -c 'mv /vault/data/unseal-key.retired-<ts> /vault/data/unseal-key' (restore the old auto-unseal key name — auto-unseal resumes; the new shares become useless, the old key is live again. If the rekey is partway and not complete, cancel it instead: vault operator rekey -cancel with a root token.)

8h. Tailnet ACL enforce — mesh.acl_enforced (R-012).

What it does: the platform pushes its generated default-deny HuJSON policy to Headscale (whole tailnet). Generated rules: baseline worker → manager-API:port (so worker heartbeats/claims survive), ingress → each reachable stack, per-stack self/worker/hosts grants (CORE/API/services/mesh_acl.py:build_acl_policy). apply_acl_policy runs with a hard preflight (acl_enforce_preflight): it refuses to push unless the manager node carries the mod-api tag and MOD_MESH_WORKER_API_PORT resolves — otherwise default-deny would sever the live mesh.

⚠ Agent lanes / switchyard: the agent lanes and the switchyard routes run on the host side (Claude Code → switchyard over the site's normal network path, not the worker tailnet). The generated policy only touches tagged tailnet nodes (workers, manager, stacks) — nothing in it references the host or the switchyard carrier. Preflight must still confirm this before the flip (next bullet).

Preflight (read-only, do this before flipping):

# 1) The exact policy that WILL be pushed, plus the gate state:
curl -s "https://<api-host>/v1/platform/mesh-acl/preview" \
-H "Authorization: Bearer $ADMIN_TOKEN" | python3 -m json.tool
# check: worker_api_port_ok == true; the policy's acls include the
# baseline worker->mod-api rule and every reachable stack's rules.
# 2) Confirm the agent-lane path is NOT a tailnet dependency:
# curl -s -m 5 https://switchyard-endpoint (whatever the lanes use) —
# if the lanes reach it via 100.64/10 tailnet addresses instead, STOP
# and resolve that before enforcing — default-deny would drop unlisted
# traffic between tailnet nodes.
# 3) Confirm the manager node is tagged mod-api (the apply-time half of
# preflight is reported in the preview output / on the next push).

Flip (one command):

curl -s -X PATCH "https://<api-host>/v1/platform/settings/mesh.acl_enforced" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \
-d '{"value": true}'

Expected result: the policy is pushed on the next stack lifecycle operation (the first 8i redeploy triggers apply_acl_policy); API logs show mesh ACL: policy pushed (enforcement ON). If preflight fails, the logs show mesh ACL: enforce preflight FAILED — refusing to push and the tailnet stays permissive (fail-safe, no lockout).

Verify: headscale policy get (or the Headscale UI) shows the default-deny policy with the expected accepts; worker heartbeats continue (GET /v1/system/drain-status stays 0-in-flight without errors; docker logs my_fastapi_app | grep -i "mesh ACL" shows the push).

Rollback (one command): PATCH .../v1/platform/settings/mesh.acl_enforced {"value": false} stops future pushes (the setting flip alone does not clear the live policy). To make the live tailnet permissive immediately, PUT an allow-all policy through the Headscale API (key read from the container env, never printed):

curl -s -X PUT "http://127.0.0.1:8090/api/v1/policy" \
-H "Authorization: Bearer $(docker exec mod_headscale printenv HEADSCALE_API_KEY)" \
-H "Content-Type: application/json" \
-d '{"policy":"{\"acls\":[{\"action\":\"accept\",\"src\":[\"*\"],\"dst\":[\"*\"]}]}"'

(Use the headscale container's actual API port if it differs from 8090.)

8i. Redeploy stacks 34 → 31 → 30 → 29 (GF.C0b.15 model-cache HF_TOKEN).

What it does: these four managed-AI stacks (preflight 10-04: 34 consign-ollama stopped, 31 qwen3.8-27b-q8-0-strix stopped, 30 qwen3.8-27b-Dell-w/nginx running, 29 qwen3.8-27b-9700-w/nginx torn_down) were deployed before the model-cache token work. Redeploying them regenerates their env with the minted tenant-scoped model_cache_token as HF_TOKEN when the model cache is the endpoint the stack pulls through — so model pulls authenticate to the cache (CORE/API/services/model_cache_service.py; wiring in managed_ai_stacks.py GF.C0b.15). The first lifecycle operation after the 8h flip also triggers the ACL policy push.

Path: the platform stack task path — stop then start via the API (platform-admin token; the Go worker executes the compose stop/start as a task, exactly like a worker-side redeploy — do not hand-run docker on the worker boxes):

for ID in 34 31 30 29; do
curl -s -X POST "https://<api-host>/v1/stacks/${ID}/stop" \
-H "Authorization: Bearer $ADMIN_TOKEN" | python3 -m json.tool
curl -s -X POST "https://<api-host>/v1/stacks/${ID}/start" \
-H "Authorization: Bearer $ADMIN_TOKEN" | python3 -m json.tool
done

29 is torn_down, not stopped — POST /v1/stacks/29/start may refuse a torn-down stack (Tim decides: intended, or does 29 need POST /v1/stacks/29/deploy instead). If start fails, stop the loop and ask; do not force it.

Verify each one before moving to the next: poll GET /v1/stacks/{id} until status == "running" (the task shows completed; the stack's services are up on the worker — a worker-side docker compose ps spot-check on one stack is enough for the others). For the HF_TOKEN half, after one stack is running: docker inspect <stack-container> --format '{{json .Config.Env}}' | grep -c HF_TOKEN shows a fresh (non-hf_-user-key) minted value, or the GET /v1/stacks/{id} secret/env view shows HF_TOKEN set from the cache.

Rollback: stop+start again (the previous env is regenerated from the same live code — the redeploy is not what changes anything except picking up the minting; there is no image rollback for a stack).

8j. Q21.6/11/12 least-privilege cutover — TONIGHT's scope only.

Full runbook: Worker least-privilege cutover. Worker releases are HELD until the one combined Q20+Q21 release — no worker is upgraded tonight. So the gated later steps do NOT run tonight:

  • Runs tonight (API-side only, safe with old workers):

    1. [OWNER-GATED] Widen the core-api Vault policy — the API broker (services/worker_approle.py) needs these paths before any cutover step can run:

      path "sys/policies/acl/worker-minimal" { capabilities = ["read"] }
      path "auth/approle/role/worker-*" { capabilities = ["create","read","update","delete","list"] }
      path "secret/data/workers/*" { capabilities = ["create","read","update","delete","list"] }

      Apply via the Vault policy edit path (owner signs off — this widens the API's Vault rights), then verify: vault policy read core-api | grep -E 'worker-minimal|approle/role/worker-\*|secret/data/workers'

    2. Baseline snapshot (cutover runbook §0) — record the live fleet: SELECT id, name, agent_version, last_heartbeat FROM worker_licenses WHERE last_heartbeat > now() - interval '10 minutes'; plus the weeks-long task inventory.

    3. Report mode — set worker.approle_per_worker_enforced = report and worker.docker_proxy_mode = report (platform settings). With old workers, report mode is a no-op safety net: the API keeps serving the shared creds and the settings simply exist in their pre-enforce state.

  • MUST WAIT for the combined worker release (never tonight): the 24–48 h watch with new workers, the enforce flips, and — above all — the shared core-api secret_id rotation (runbook §4). The rotation breaks every worker still logging in with the shared pair; it is gated on 100% of live workers running the new release with per-worker AppRoles. Doing it tonight would take the whole fleet down.

Rollback for tonight's part: PATCH both settings back to off; the policy grant is additive (leave it — revoking it just re-blocks the cutover later).

8k. Setting flips (after smoke passes): erasure + internal-TLS report.

Both are platform settings set through the admin API (or Admin UI → Platform Settings). DB-backed and audited; no restart needed to be read, but the erasure job only runs in the current API process — flip it after the new image is up (it is — we just recreated it in 8a).

8k-a. Enable the privacy-erasure sweep (Q3a/b/c):

curl -s -X PATCH "https://<api-host>/v1/platform/settings/privacy.erasure_job_enabled" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \
-d '{"value": true}'
# verify: GET the same key → value true, source = db

8k-b. Turn internal-hop TLS into report mode (Q11). report logs a plain http:// internal hop once per host (soak telemetry) and uses the internal CA bundle when present. It does not refuse plaintext yet — that is the 48 h follow-up.

curl -s -X PATCH "https://<api-host>/v1/platform/settings/gsec.transport.internal_tls_enforce" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \
-d '{"value": "report"}'
# verify: GET the same key → value "report"

Both write an audit row (action=set, actor, key, new value).

8l. Resume dispatch + post-window fence check.

curl -s -X POST "https://<api-host>/v1/system/resume" -H "Authorization: Bearer $ADMIN_TOKEN"
# post-window GPU fence check (from §7):
curl -s -m 10 http://192.168.1.210:8080/workers | grep -vE "Bearer|authorization|access_token|eyJ|password|secret|token"
# expect total 2 healthy — if not, stop+start the fenced stacks via
# POST /v1/stacks/{id}/stop then /start (see §7).

Un-pause the agent lanes now (Tim).

9. 48 h follow-up​

After report has been running ~48 h and the logs show the expected set of internal hops (and no unexpected plain-http host you didn't plan to TLS-ify):

curl -s -X PATCH "https://<api-host>/v1/platform/settings/gsec.transport.internal_tls_enforce" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \
-d '{"value": "require"}'

require makes the internal CA bundle (MOD_INTERNAL_CA_BUNDLE) mandatory and refuses plain http:// internal URLs. Before flipping to require, confirm the bundle is present and trusted (see internal TLS); a missing bundle raises, so you want the bundle in place before require takes effect. Also on the 48 h list: the Q21 report-mode watch (worker least-privilege runbook §2) — its enforce + rotation steps still wait for the combined worker release.

10. Rollback (per component, reverse order)​

No data loss in any rollback: migrations are additive, and the DB-TLS rollback is covered in the linked runbook.

API (and FORWARDED_ALLOW_IPS, client TLS):

docker tag mod-api:rollback-${DATE} mod-api:latest
$C --profile all up -d --force-recreate api api-background
curl -fsS http://127.0.0.1:8000/health

This also reverts FORWARDED_ALLOW_IPS (the compose env reverts with the edit) and the API-side DB-TLS client args. If you had changed the two compose FORWARDED_ALLOW_IPS lines, revert those edits too — they only take effect on recreate, so a compose revert + recreate is the clean path.

Keycloak:

docker tag keycloak/keycloak:rollback-${DATE} keycloak/keycloak:26.3
$C --profile all up -d --force-recreate keycloak

Reverts to the public image + (restored) theme bind mount.

nginx (CORS map): revert scripts/nginx-entrypoint.sh (or CORS_ALLOWED_ORIGINS) and docker compose up -d --no-deps --force-recreate nginx.

DB-TLS: see the ROLLBACK section of DB TLS + SCRAM switch-over window. Reverse order, no re-hash needed (SCRAM hashes still authenticate under the old pg_hba).

Vault Shamir rekey: one rename — restore unseal-key.retired-<ts> to unseal-key inside my_vault (auto-unseal resumes; see 8g).

mesh.acl_enforced: PATCH back to false + PUT the allow-all policy to Headscale (see 8h).

Setting flips (if a flip caused trouble):

# revert privacy-erasure (safe to leave on, but to turn off):
curl -s -X PATCH ".../v1/platform/settings/privacy.erasure_job_enabled" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" -d '{"value": false}'
# revert internal TLS posture to off:
curl -s -X PATCH ".../v1/platform/settings/gsec.transport.internal_tls_enforce" \
-H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" -d '{"value": "off"}'

Both are instant (DB-backed) and write an audit row.

11. Estimated downtime per component​

ComponentTime to recreate + healthyNotes
drain (wait for in-flight)0–15 min, usually < 2no downtime — work already queued
API (+ image switch + FORWARDED_ALLOW_IPS)~40–120 sthe 120 s fence clock is this span
Keycloak~30–90 sJVM cold start + healthcheck (30 s start period)
Postgres (DB-TLS)~20–40 ssee DB-TLS runbook
PgBouncer (DB-TLS)~15–30 ssee DB-TLS runbook
nginx (CORS map)~10–20 sno API dependency — runs after smoke, or any time
Vault Shamir rekey~1–2 minAPI stays up; drain already done
stack redeploys 34/31/30/2920–40 minsequential, API up; these are the LLM stacks — expect them offline during their own redeploy only
API unavailable total~2–4 minoverlaps with Keycloak; keep under 120 s for GPU stacks

The stack stays drained across all of it, so users see queueing, not errors.