Skip to main content

Sealed / Air-Gapped Operation

MOD supports a sealed mode for installs that must make no outbound calls to hosts outside the install's own network. Sealed mode is a platform setting, not a compliance product: it supports an operator's egress-control requirements. MOD does not make an install compliant with any framework, and enabling sealed mode does not by itself satisfy any control — the operator remains responsible for their own network policy and assessment.

What sealed mode does​

egress.sealed_mode is the master egress switch (default False). When it is True, the API-side egress policy (CORE/API/utils/egress_policy.py) refuses every gated outbound category, regardless of each category's own switch. The gated paths are the ones the egress inventory (docs/security-audit/EGRESS_INVENTORY.md) identified as capable of carrying tenant content or install metadata: Hugging Face access, NIM cloud calls, external-LLM providers, tenant webhooks, the node wizard's internet lookups, the license phone-home, Stripe server calls, and the HIBP password lookup.

Two behaviours are worth knowing:

  • Local serving is never blocked. Sealed mode only refuses external targets. A local or stack LLM endpoint (e.g. http://vllm:8000 or a stack alias) keeps working; only provider URLs classified as external are refused.
  • Reads fail open. If the settings store itself errors, the master read is treated as NOT sealed — a settings outage must not silently seal a live install and stop every outbound feature. Per-category reads fall back to their registered default (True).

Refusals are reported to the caller as an error/log detail naming the exact switch that stopped the egress (e.g. egress disabled: egress.sealed_mode is True (sealed mode) — blocks egress.nim_enabled), so an operator can see which switch caused a failure.

Per-category egress switches​

Each category is its own boolean platform setting, default True, so a deploy changes nothing until an admin turns a switch off. With egress.sealed_mode True, all of them are overridden to refused.

SettingDefaultWhat it stops when False (or when sealed mode is on)
egress.sealed_modeFalseMaster switch — refuses every category below, plus Stripe server calls and the HIBP password lookup.
egress.nim_enabledTrueNIM router initialization and NIM node dispatch (NVIDIA NIM cloud endpoints).
egress.model_download_enabledTrueHugging Face downloads and the model-cache upstream fall-through; HF token validation (whoami). Only locally mirrored weights (MinIO-mirrored / cached) remain usable.
egress.external_llm_enabledTrueCalls to external LLM providers (assistant, wizard/repair LLM transport, vLLM max-model-len probe) — external targets only; local/stack LLM endpoints are unaffected.
egress.webhooks_enabledTrueOutbound webhooks: trigger dispatch, run-completion callbacks, Slack event notifications, admin SCIM-activity webhook validation probes. Affected calls return a "disabled"/not-verified status instead of posting.
egress.wizard_internet_enabledTrueNode wizard internet lookups: repo fetches (GitHub) and PyPI classifier lookups. Repo fetches return empty ("couldn't sample").
egress.license_phonehome_enabledTrueThe license client's server validation call. Skipped when off — a valid cached license must NOT fail; this mirrors the existing network-unreachable behaviour.
egress.update_check_enabledTrueUpdate-check / release-feed egress: the worker heartbeat's update_available advertisement and the admin currency-rate fetch. When off, workers are never advertised a newer version and the rate refresh returns a 503 instead of calling out.
egress.telemetry_enabledTrueUsage-telemetry ingestion: the login-page RUM first-paint beacon. When off the beacon endpoint still acknowledges the POST but records nothing.

Master-only effects with no per-category key: Stripe client initialization returns the existing 503 shape ("Stripe is disabled on this install"), and the HIBP k-anonymity lookup is skipped so password policy falls back to the local blocklist. Stripe on an on-prem install is normally stopped simply by leaving keys unconfigured.

Verifying egress was actually refused (refusal ledger)​

A switch being False (or sealed mode being on) means the gated call is refused, but an operator auditing a sealed site wants to see the refusal, not infer it. The update-check and telemetry call sites record each refused call into a bounded, in-process refusal ledger (category + a stable site code + the switch that stopped it — never a URL, body, key, or tenant id). The ledger is surfaced read-only on the compliance egress-status report:

GET /v1/platform/compliance/egress-status (platform admin)
-> "update_check_enabled": bool,
"telemetry_enabled": bool,
"refusals": { "<category>": {"count": n, "recent_reasons": [ ... ]} }

So after a worker heartbeat is refused while egress.update_check_enabled is False, the report shows an egress.update_check_enabled entry whose recent_reasons names the worker_heartbeat.update_advertise site. A fresh install (nothing refused yet) shows refusals: {}. The ledger is in-process and zeroes on a restart; where a gate also writes a durable audit row, that row is the authoritative record.

Per-store residency flags (declaration-based)​

Sealed / hardened deployments also let an operator declare where each long-lived data store lives, for the residency evidence a CUI/CJIS/ITAR customer asks for. This is a declaration: MOD never calls a cloud API to infer a region — the operator is the source of truth. Two platform settings (group residency):

SettingDefaultMeaning
residency.required_region""The region the customer requires stores to live in (us / eu). Empty = no requirement declared (the conformance check is manual).
residency.store_regions{}Operator-declared region per store: postgres, object_storage (versitygw/MinIO default and BYO StorageConnections), valkey, qdrant, backups (off-box shipping target). Values us / eu / unknown; absent = undeclared.

Both are read-only on the same GET /v1/platform/compliance/egress-status surface under "residency", alongside the per-store in_boundary flags and a computed mismatches list (stores whose declared region is not the required one, when a requirement is set).

The conformance check residency.required_region_match (GY.C12, services/conformance/checks.py) enforces the match:

  • residency.required_region empty → status manual (verify by your own evidence; nothing is asserted).
  • required region set → pass when every store declares exactly that region; fail otherwise, with per-store mismatch evidence (an undeclared or unknown store counts as a mismatch).

This is a separate axis from the GY.V3a residency.region_pin (which gates storage writes against a zone) — the flags here are reporting-only and add no write gate. MOD supports residency declaration and matching; it does not certify that data is actually stored in the declared region.

How to turn it on​

  1. Admin UI: Dashboard → Platform Settings, group egress. Set egress.sealed_mode to True (and/or any individual category switch).
  2. API: PATCH /v1/platform/settings/egress.sealed_mode with the value true (admin auth). The same PATCH path applies to every category key; GET /v1/platform/settings lists the registered keys.

Settings take effect on the next gated call.

What still needs network — and how to pre-stage it​

Sealed mode gates the API-side call sites; it does not pre-stage anything. A sealed or air-gapped install still needs its dependencies present locally before the network is cut:

  • Container images. Node/stack images are pulled from a registry at deploy time. Pre-stage them in a local registry (DOCKER_REGISTRY) or cache them on the workers and set worker.prefer_local_images = True, which skips re-pulling an image that is already cached locally. Image pulls are not gated by the egress switches — they are controlled by registry configuration (EGRESS_INVENTORY row 16), so a sealed install must point at a local registry or rely on cached images.
  • Model weights. With egress.model_download_enabled False (or sealed), only MinIO-mirrored / already-cached weights are used. Pre-download the models a workflow needs (embedding models, NIM/SDXL weights, etc.) while network is available, so the model-cache and MinIO mirrors are primed. A proven offline run depends on a cache-primed install.
  • Hugging Face. HF access is gated (egress.model_download_enabled); HF tokens can still be stored but cannot be validated while sealed. Pre-stage anything fetched from HF in advance.
  • Frontend CDN loads. The Dashboard loads its third-party libraries (three.js, loaders.gl for point-cloud/LAS preview) from vendored copies served by MOD itself, so the browser makes no CDN calls. Installs running a Dashboard build older than this change should update first.
  • Install-time paths. The worker setup script's package-repo fallbacks and build-time package installs are install-time only, not runtime egress; run setup while network is available.

Known not-yet-gated items (stated honestly)​

  • Worker model-cache HF probe: the worker-side model cache can still reach Hugging Face upstream; sealed mode gates the API-side HF paths, not this probe. Pre-priming the cache is the current mitigation.
  • Registry pulls: image pulls are not gated by the egress switches (see above); they depend on DOCKER_REGISTRY / worker.prefer_local_images.
  • SMTP and external executors: SMTP makes no outbound call unless configured; external executors (Kubernetes/Qube/Deadline) and DERP are admin-configured or self-hosted and are not egress-switch targets.
  • MOD's own domains: on a sealed on-prem install, MOD's own services are the local network — traffic between the install's containers is not egress.

How to verify a sealed install​

  1. Sealed state is visible to the Dashboard: GET /v1/platform/features returns sealed_mode: true when the master switch is on (false otherwise). The Dashboard billing page skips loading Stripe.js when sealed.
  2. Per-path checks are listed in CORE/API/tests/MANUAL_TEST_GSEC_FIXES.md ("GY.C8 sealed mode completeness"): refused calls show a log line naming the switch, and a network trace/proxy log should show zero requests to the affected hosts.
  3. Ports/protocols/services: the published surface of the install is inventoried in docs/security-audit/PPS_INVENTORY.md, generated by docs/security-audit/tools/pps_inventory.py from docker-compose.yml and the nginx template — regenerate it after any topology change and confirm the remaining inbound surface matches your policy.

Full path-by-path inventory with which switch covers each row: docs/security-audit/EGRESS_INVENTORY.md.

Moving models into a sealed site (bundle export/import)​

Use the bundle when a sealed (air-gapped) install needs model weights that only exist on a connected install — i.e. the weights were already cached on the connected site, and the sealed site has no egress to fetch them. It is an operator-carried transfer (USB/physical media/escrow), not a network path: nothing in a sealed install ever reaches out for it.

The two ends:

  1. Export on the connected site. Ask for the repos you want; the cache streams a single tar (mod-model-cache-bundle.tar) of the cached file bytes plus a manifest.json written as the LAST member. Export NEVER fetches upstream — a repo with nothing cached is reported repos_not_cached (404), and an empty request is empty_export (400).
  2. Carry the tar to the sealed site (one file, checksum-verifiable).
  3. Import on the sealed site. The body is the raw tar. The import streams it into a temp directory, computes every file's sha256, and matches files to manifest entries by (size, sha256).

What the verification protects (all of it runs before a single byte is written to the sealed cache):

  • Tampering — every manifest entry must verify with exact size and sha256; any mismatch, missing entry, or extra file in the tar rejects the whole bundle (sha256_mismatch / unexpected_bundle_content, 400) and the cache is left untouched (nothing partially imported).
  • Path traversal — a manifest path containing a .. segment is refused (unsafe path in manifest), and repo ids are validated.
  • Wrong/foreign bundle — the manifest must declare format: "mod-model-cache-bundle" and version: 1; a missing/duplicate manifest.json or non-JSON manifest is invalid_bundle (400).
  • Oversize — an incoming bundle over the 64 GiB cap is rejected 413 bundle_too_large before it is parsed.
  • Gated bytes — export is refuse-unless-positively-public: a repo with any tenant-namespaced (tenant-*) objects is refused gated_repo_refused (400), and a manifest naming a tenant-namespaced repo id is refused gated_bundle_refused on import. Per-tenant gated bytes never cross a site boundary in a bundle.

Manifest shape (what the receiving end checks):

{"format": "mod-model-cache-bundle", "version": 1,
"created_utc": "...", "cache_version": "...",
"files": [{"repo_id": "...", "revision": "...", "commit": "...",
"path": "...", "size": 0, "sha256": "..."}]}

curl against the system-admin routes (system admin only; the action is audited as ADMIN_DATA_EXPORT / the matching import action):

# export on the connected site
curl -sS -X POST "https://modintel.modtechlabs.com/v1/system/model-cache/export-bundle" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"repos":[{"repo_id":"stabilityai/sdxl-turbo","revision":"main"}]}' \
-o mod-model-cache-bundle.tar

# import on the sealed site
curl -sS -X POST "https://sealed-site.local/v1/system/model-cache/import-bundle" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/x-tar" \
--data-binary @mod-model-cache-bundle.tar

A successful import returns {"imported": <files>, "bytes": <total>, "repos": [...]}. After import, resolves on the sealed site are zero-upstream hits for those repos (see the cache-hit proof primitive in CORE/API/tests/MANUAL_TEST_MODEL_CACHE.md).

Related: SIEM hookup — an external SIEM endpoint is an egress path, so sealed mode or egress.webhooks_enabled = false stops the shipper and raises one siem_export_failing alert per outage.