Sealed / Air-Gapped Operation
MOD supports a sealed mode for installs that must make no outbound calls to hosts outside the install's own network. Sealed mode is a platform setting, not a compliance product: it supports an operator's egress-control requirements. MOD does not make an install compliant with any framework, and enabling sealed mode does not by itself satisfy any control — the operator remains responsible for their own network policy and assessment.
What sealed mode does
egress.sealed_mode is the master egress switch (default False). When it is
True, the API-side egress policy (CORE/API/utils/egress_policy.py) refuses
every gated outbound category, regardless of each category's own switch. The
gated paths are the ones the egress inventory
(docs/security-audit/EGRESS_INVENTORY.md) identified as capable of carrying
tenant content or install metadata: Hugging Face access, NIM cloud calls,
external-LLM providers, tenant webhooks, the node wizard's internet lookups,
the license phone-home, Stripe server calls, and the HIBP password
lookup.
Two behaviours are worth knowing:
- Local serving is never blocked. Sealed mode only refuses external
targets. A local or stack LLM endpoint (e.g.
http://vllm:8000or a stack alias) keeps working; only provider URLs classified as external are refused. - Reads fail open. If the settings store itself errors, the master read is
treated as NOT sealed — a settings outage must not silently seal a live
install and stop every outbound feature. Per-category reads fall back to
their registered default (
True).
Refusals are reported to the caller as an error/log detail naming the exact
switch that stopped the egress (e.g. egress disabled: egress.sealed_mode is True (sealed mode) — blocks egress.nim_enabled), so an operator can see which
switch caused a failure.
Per-category egress switches
Each category is its own boolean platform setting, default True, so a deploy
changes nothing until an admin turns a switch off. With egress.sealed_mode
True, all of them are overridden to refused.
| Setting | Default | What it stops when False (or when sealed mode is on) |
|---|---|---|
egress.sealed_mode | False | Master switch — refuses every category below, plus Stripe server calls and the HIBP password lookup. |
egress.nim_enabled | True | NIM router initialization and NIM node dispatch (NVIDIA NIM cloud endpoints). |
egress.model_download_enabled | True | Hugging Face downloads and the model-cache upstream fall-through; HF token validation (whoami). Only locally mirrored weights (MinIO-mirrored / cached) remain usable. |
egress.external_llm_enabled | True | Calls to external LLM providers (assistant, wizard/repair LLM transport, vLLM max-model-len probe) — external targets only; local/stack LLM endpoints are unaffected. |
egress.webhooks_enabled | True | Outbound webhooks: trigger dispatch, run-completion callbacks, Slack event notifications, admin SCIM-activity webhook validation probes. Affected calls return a "disabled"/not-verified status instead of posting. |
egress.wizard_internet_enabled | True | Node wizard internet lookups: repo fetches (GitHub) and PyPI classifier lookups. Repo fetches return empty ("couldn't sample"). |
egress.license_phonehome_enabled | True | The license client's server validation call. Skipped when off — a valid cached license must NOT fail; this mirrors the existing network-unreachable behaviour. |
egress.update_check_enabled | True | Update-check / release-feed egress: the worker heartbeat's update_available advertisement and the admin currency-rate fetch. When off, workers are never advertised a newer version and the rate refresh returns a 503 instead of calling out. |
egress.telemetry_enabled | True | Usage-telemetry ingestion: the login-page RUM first-paint beacon. When off the beacon endpoint still acknowledges the POST but records nothing. |
Master-only effects with no per-category key: Stripe client initialization returns the existing 503 shape ("Stripe is disabled on this install"), and the HIBP k-anonymity lookup is skipped so password policy falls back to the local blocklist. Stripe on an on-prem install is normally stopped simply by leaving keys unconfigured.
Verifying egress was actually refused (refusal ledger)
A switch being False (or sealed mode being on) means the gated call is
refused, but an operator auditing a sealed site wants to see the refusal,
not infer it. The update-check and telemetry call sites record each refused
call into a bounded, in-process refusal ledger (category + a stable site
code + the switch that stopped it — never a URL, body, key, or tenant id).
The ledger is surfaced read-only on the compliance egress-status report:
GET /v1/platform/compliance/egress-status (platform admin)
-> "update_check_enabled": bool,
"telemetry_enabled": bool,
"refusals": { "<category>": {"count": n, "recent_reasons": [ ... ]} }
So after a worker heartbeat is refused while egress.update_check_enabled is
False, the report shows an egress.update_check_enabled entry whose
recent_reasons names the worker_heartbeat.update_advertise site. A fresh
install (nothing refused yet) shows refusals: {}. The ledger is
in-process and zeroes on a restart; where a gate also writes a durable audit
row, that row is the authoritative record.
Per-store residency flags (declaration-based)
Sealed / hardened deployments also let an operator declare where each
long-lived data store lives, for the residency evidence a CUI/CJIS/ITAR
customer asks for. This is a declaration: MOD never calls a cloud API to
infer a region — the operator is the source of truth. Two platform settings
(group residency):
| Setting | Default | Meaning |
|---|---|---|
residency.required_region | "" | The region the customer requires stores to live in (us / eu). Empty = no requirement declared (the conformance check is manual). |
residency.store_regions | {} | Operator-declared region per store: postgres, object_storage (versitygw/MinIO default and BYO StorageConnections), valkey, qdrant, backups (off-box shipping target). Values us / eu / unknown; absent = undeclared. |
Both are read-only on the same GET /v1/platform/compliance/egress-status
surface under "residency", alongside the per-store in_boundary flags and
a computed mismatches list (stores whose declared region is not the required
one, when a requirement is set).
The conformance check residency.required_region_match (GY.C12,
services/conformance/checks.py) enforces the match:
residency.required_regionempty → statusmanual(verify by your own evidence; nothing is asserted).- required region set →
passwhen every store declares exactly that region;failotherwise, with per-store mismatch evidence (an undeclared orunknownstore counts as a mismatch).
This is a separate axis from the GY.V3a residency.region_pin (which gates
storage writes against a zone) — the flags here are reporting-only and add
no write gate. MOD supports residency declaration and matching; it does not
certify that data is actually stored in the declared region.
How to turn it on
- Admin UI: Dashboard → Platform Settings, group egress. Set
egress.sealed_modetoTrue(and/or any individual category switch). - API:
PATCH /v1/platform/settings/egress.sealed_modewith the valuetrue(admin auth). The same PATCH path applies to every category key;GET /v1/platform/settingslists the registered keys.
Settings take effect on the next gated call.
What still needs network — and how to pre-stage it
Sealed mode gates the API-side call sites; it does not pre-stage anything. A sealed or air-gapped install still needs its dependencies present locally before the network is cut:
- Container images. Node/stack images are pulled from a registry at deploy
time. Pre-stage them in a local registry (
DOCKER_REGISTRY) or cache them on the workers and setworker.prefer_local_images=True, which skips re-pulling an image that is already cached locally. Image pulls are not gated by the egress switches — they are controlled by registry configuration (EGRESS_INVENTORY row 16), so a sealed install must point at a local registry or rely on cached images. - Model weights. With
egress.model_download_enabledFalse(or sealed), only MinIO-mirrored / already-cached weights are used. Pre-download the models a workflow needs (embedding models, NIM/SDXL weights, etc.) while network is available, so the model-cache and MinIO mirrors are primed. A proven offline run depends on a cache-primed install. - Hugging Face. HF access is gated (
egress.model_download_enabled); HF tokens can still be stored but cannot be validated while sealed. Pre-stage anything fetched from HF in advance. - Frontend CDN loads. The Dashboard loads its third-party libraries (three.js, loaders.gl for point-cloud/LAS preview) from vendored copies served by MOD itself, so the browser makes no CDN calls. Installs running a Dashboard build older than this change should update first.
- Install-time paths. The worker setup script's package-repo fallbacks and build-time package installs are install-time only, not runtime egress; run setup while network is available.
Known not-yet-gated items (stated honestly)
- Worker model-cache HF probe: the worker-side model cache can still reach Hugging Face upstream; sealed mode gates the API-side HF paths, not this probe. Pre-priming the cache is the current mitigation.
- Registry pulls: image pulls are not gated by the egress switches (see
above); they depend on
DOCKER_REGISTRY/worker.prefer_local_images. - SMTP and external executors: SMTP makes no outbound call unless configured; external executors (Kubernetes/Qube/Deadline) and DERP are admin-configured or self-hosted and are not egress-switch targets.
- MOD's own domains: on a sealed on-prem install, MOD's own services are the local network — traffic between the install's containers is not egress.
How to verify a sealed install
- Sealed state is visible to the Dashboard:
GET /v1/platform/featuresreturnssealed_mode: truewhen the master switch is on (falseotherwise). The Dashboard billing page skips loading Stripe.js when sealed. - Per-path checks are listed in
CORE/API/tests/MANUAL_TEST_GSEC_FIXES.md("GY.C8 sealed mode completeness"): refused calls show a log line naming the switch, and a network trace/proxy log should show zero requests to the affected hosts. - Ports/protocols/services: the published surface of the install is
inventoried in
docs/security-audit/PPS_INVENTORY.md, generated bydocs/security-audit/tools/pps_inventory.pyfromdocker-compose.ymland the nginx template — regenerate it after any topology change and confirm the remaining inbound surface matches your policy.
Full path-by-path inventory with which switch covers each row:
docs/security-audit/EGRESS_INVENTORY.md.
Moving models into a sealed site (bundle export/import)
Use the bundle when a sealed (air-gapped) install needs model weights that only exist on a connected install — i.e. the weights were already cached on the connected site, and the sealed site has no egress to fetch them. It is an operator-carried transfer (USB/physical media/escrow), not a network path: nothing in a sealed install ever reaches out for it.
The two ends:
- Export on the connected site. Ask for the repos you want; the cache
streams a single tar (
mod-model-cache-bundle.tar) of the cached file bytes plus amanifest.jsonwritten as the LAST member. Export NEVER fetches upstream — a repo with nothing cached is reportedrepos_not_cached(404), and an empty request isempty_export(400). - Carry the tar to the sealed site (one file, checksum-verifiable).
- Import on the sealed site. The body is the raw tar. The import
streams it into a temp directory, computes every file's sha256, and
matches files to manifest entries by
(size, sha256).
What the verification protects (all of it runs before a single byte is written to the sealed cache):
- Tampering — every manifest entry must verify with exact
sizeandsha256; any mismatch, missing entry, or extra file in the tar rejects the whole bundle (sha256_mismatch/unexpected_bundle_content, 400) and the cache is left untouched (nothing partially imported). - Path traversal — a manifest
pathcontaining a..segment is refused (unsafe path in manifest), and repo ids are validated. - Wrong/foreign bundle — the manifest must declare
format: "mod-model-cache-bundle"andversion: 1; a missing/duplicatemanifest.jsonor non-JSON manifest isinvalid_bundle(400). - Oversize — an incoming bundle over the 64 GiB cap is rejected
413 bundle_too_largebefore it is parsed. - Gated bytes — export is refuse-unless-positively-public: a repo with
any tenant-namespaced (
tenant-*) objects is refusedgated_repo_refused(400), and a manifest naming a tenant-namespaced repo id is refusedgated_bundle_refusedon import. Per-tenant gated bytes never cross a site boundary in a bundle.
Manifest shape (what the receiving end checks):
{"format": "mod-model-cache-bundle", "version": 1,
"created_utc": "...", "cache_version": "...",
"files": [{"repo_id": "...", "revision": "...", "commit": "...",
"path": "...", "size": 0, "sha256": "..."}]}
curl against the system-admin routes (system admin only; the action is
audited as ADMIN_DATA_EXPORT / the matching import action):
# export on the connected site
curl -sS -X POST "https://modintel.modtechlabs.com/v1/system/model-cache/export-bundle" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"repos":[{"repo_id":"stabilityai/sdxl-turbo","revision":"main"}]}' \
-o mod-model-cache-bundle.tar
# import on the sealed site
curl -sS -X POST "https://sealed-site.local/v1/system/model-cache/import-bundle" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/x-tar" \
--data-binary @mod-model-cache-bundle.tar
A successful import returns {"imported": <files>, "bytes": <total>, "repos": [...]}. After import, resolves on the sealed site are zero-upstream
hits for those repos (see the cache-hit proof primitive in
CORE/API/tests/MANUAL_TEST_MODEL_CACHE.md).
Related: SIEM hookup — an external SIEM endpoint is
an egress path, so sealed mode or egress.webhooks_enabled = false stops the
shipper and raises one siem_export_failing alert per outage.