Skip to main content

SIEM Hookup

MOD supports sending its audit trail to your own SIEM: a shipper in CORE/API/services/siem_export.py periodically streams rows of permission_audit_logs to an operator-configured endpoint. MOD does not inspect, score or certify what your SIEM does with those events, and turning the shipper on does not by itself satisfy any audit control — the receiver side (storage, retention, alerting) is yours.

Everything below describes what the code actually does. If your receiver is not one of the four formats below, the shipper will not speak to it.

What ships​

One tick (run_export_tick, driven by the scheduler job siem_export in CORE/API/services/scheduler.py) queries permission_audit_logs rows with id > siem.export_checkpoint, ordered ascending, capped at siem.batch_size. Each row is normalized to a metadata-only event: id, ts, tenant_id, actor, action, resource (resource_type:resource_id), outcome, source_ip, integrity_hash. Audit metadata blobs and request/response bodies are never serialized into the export.

outcome is derived from the action name, not a column: an action containing any of failed, failure, denied, blocked, invalid, expired, error, suspicious, unusual is reported as failure. In practice that means events such as auth_login_failed and permission_request_denied arrive as failures, and failures get the loud severity in every format (syslog severity=4, CEF sev=7, OTLP severityNumber=13 / severityText=ERROR).

The checkpoint (siem.export_checkpoint) advances only after a batch is fully sent. An outage therefore loses nothing: worst case the next tick re-sends the same rows, and each carries its row id for SIEM-side dedup.

Formats and transports​

siem.formatsiem.transportWire
rfc5424 (default)tcp_tls (default) or udpRFC 5424 syslog line, facility local4, hostname mod-api, app name mod-core-api, structured-data tag mod@32473 carrying event/tenant/actor/action/resource/outcome/ip/integrity
ceftcp_tls or udpArcSight CEF flat line CEF:0|MOD|MOD Core|audit|<action>|<sev> with extensions rt, src, externalId, outcome, cs1=tenant, cs2=actor, cs3=resource, cs4=integrity hash
splunk_hechttps onlyOne Splunk HEC JSON document per event (sourcetype: _json, source: mod-audit), Authorization: Splunk <token>
otlphttps onlyOne OTLP/HTTP JSON request per batch (resourceLogs → logRecords, service.name = mod-core-api, scope.name = mod-audit, attributes mod.*), Authorization: Bearer <token>

tcp_tls means RFC 5425: one TLS session per batch, each message framed <octet-count> SP <message> SP. TLS uses the Python default context (ssl.create_default_context()), so the receiver's certificate must verify against your system CA bundle — a self-signed receiver without a CA handover will fail the send. No client certificate is presented; auth on the syslog transports is by socket only. udp sends one datagram per line (no authentication, and very long lines can be truncated by the receiver).

rfc5424/cef require a syslog transport; splunk_hec/otlp require https. The pairing is checked at send time, so a bad combination shows up as a send failure (see the alert below) rather than a settings-validation error.

Settings​

All seven keys are registered in siem_export.py itself (group siem) and are set through the ordinary platform-settings surface — GET /v1/platform/settings/{key} to read, PATCH /v1/platform/settings/{key} with {"value": ...} to write (platform admin; the write is audited).

KeyDefaultMeaning
siem.enabledFalseMaster switch; when false a tick is a no-op.
siem.formatrfc5424One of rfc5424, cef, splunk_hec, otlp.
siem.endpoint""host:port for tcp_tls/udp, full URL for https.
siem.transporttcp_tlsOne of tcp_tls, udp, https.
siem.batch_size500Max audit rows shipped per tick.
siem.interval_seconds30Minimum seconds between ticks (the scheduler heartbeat itself is a fixed 10 s; this setting is re-read every pass, so raising it takes effect immediately and lowering it needs no restart either — the interval is enforced in the tick, not the trigger).
siem.export_checkpoint0Last row id fully delivered. Maintained by the shipper.
# enable, syslog over a TLS listener
curl -sS -X PATCH "$BASE/v1/platform/settings/siem.enabled" \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"value": true}'
curl -sS -X PATCH "$BASE/v1/platform/settings/siem.endpoint" \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"value": "siem.internal.local:65140"}'

The token (HEC / OTLP only)​

The splunk_hec and otlp auth token is a secret: it is read from Vault KV v2 at infrastructure/siem/token, key value, using the same AppRole pattern as the HF token and the managed-LLM key (VAULT_APPROLE_ROLE_ID / VAULT_APPROLE_SECRET_ID). It is never a platform setting and never appears in a log or an exception surface. Without a readable token those formats return no_token and raise the outage alert instead of sending.

vault kv write infrastructure/siem/token value=<hec-or-otlp-token>

Egress gate and the AU-5 alert​

An external endpoint (per utils/egress_policy.url_is_external) is refused when the webhooks egress category is off — i.e. egress.sealed_mode is True or egress.webhooks_enabled is False (see Sealed / Air-Gapped Operation). Nothing is sent, and the refusal is treated as an outage. A local/alias endpoint is not gated.

On any outage — send failure, missing token, egress refusal — exactly one SIEM_EXPORT_FAILING security audit event is written per outage (models.AuditAction.SIEM_EXPORT_FAILING, value siem_export_failing, tenant 0), with metadata reason, endpoint (host only, never a path or credential), format, transport. Retries keep happening every tick; the alert stops repeating until a send succeeds, and a new outage alerts again. That is the AU-5 signal: the audit trail itself records that the audit trail stopped shipping.

The counter half of the alert: a lost audit row leaves no trace in the trail, so CORE/API/audit.py also keeps a Prometheus counter audit_write_failures_total (label site, one increment per failed audit write). It is best-effort by design — a failing counter degrades to just the WARNING log line and never raises into the request path. Increments happen in _record_audit_write_failure() at the audit write sites in audit.py (log_security_event, log_admin_action, and log_resource_lifecycle:<action> for resource lifecycle events) and in CORE/API/services/secrets_service.py (secret_accessed_audit, secret_accessed_by_name_audit, secret_access_denied_audit). Counters are exposed at GET /metrics on the API (CORE/API/main.py, via prometheus-fastapi-instrumentator, include_in_schema=False) and that endpoint is gated by the platform setting api.metrics_require_auth (group api, default True: authenticated platform admins, or localhost/docker-internal source IPs — the in-cluster mod_prometheus reaches the API as a docker-bridge peer, so its scrape keeps working). Set the setting to False only for an internal/debug posture. Alert on the counter in your own pipeline:

curl -sS "$BASE/metrics" -H "Authorization: Bearer $TOKEN" \
| grep audit_write_failures_total

MOD supports that alerting hook; how your SIEM or monitoring layer scores, alerts on, or acts on the counter is yours.

Receiver examples​

rsyslog as the endpoint (UDP) — one RFC 5424 line per datagram, parsed as a normal syslog record with the mod@32473 structured data attached:

# /etc/rsyslog.d/mod-audit.conf
input(type="imudp" port="15140" name="mod-audit-in")
template(name="ModAuditRaw" type="string" string="%RAWLINE%\n")
if :name is "mod-audit-in" then {
:output: "/var/log/mod-audit.log"
:format: "ModAuditRaw"
}

Set siem.transport=udp, siem.endpoint=127.0.0.1:15140, siem.format=rfc5424. For tcp_tls the receiver must implement RFC 5425 octet-count framing over a TLS endpoint whose certificate verifies against the system CA bundle; the shipper will not negotiate a client certificate.

Splunk HEC — siem.format=splunk_hec, siem.transport=https, siem.endpoint=https://<hec-receiver>/services/collector/event, token in the Vault slot above. The shipper posts one event document per audit row, so a collector that batches differently is fine but expect one request per row.

OTLP/HTTP — siem.format=otlp, siem.transport=https, siem.endpoint = your OTel collector's /v1/logs (or equivalent) URL, token in the Vault slot. Records arrive in one request per batch with attributes prefixed mod..

There is no Elasticsearch ingest, no Postmark-style form, and no plain unauthenticated HTTPS POST: those are not transports the code supports.

How to test​

  1. Unit level (no network; transports are monkeypatched to fail loudly):

    cd CORE/API && scripts/pytest-sibling.sh tests/test_gy_c5_siem_export.py \
    > /tmp/siem.log 2>&1; grep -E "[0-9]+ (passed|failed)" /tmp/siem.log
  2. Live, local receiver: point siem.endpoint at a UDP listener, set siem.enabled=true, siem.interval_seconds=30, then generate an audit event (any admin setting write is audited). Expect the listener to print one <124>1 ... mod-api mod-core-api - [mod@32473 ...] line per row, and siem.export_checkpoint to have advanced to the highest row id.

  3. Outage path: stop the listener. The next tick returns send_failed, the checkpoint does not advance, and one siem_export_failing row appears:

    curl -sS "$BASE/v1/audit-logs/permissions?action=siem_export_failing" \
    -H "Authorization: Bearer $TOKEN"
  4. Sealed interaction: with egress.sealed_mode=true and an external endpoint, the tick returns egress_refused and alerts once — the shipper never silently bypasses the egress policy.

The integrity hash in the stream​

Each shipped event carries the row's integrity_hash — the HMAC-SHA256 the writer computed over the row's canonical JSON (json.dumps(..., sort_keys=True)) including that row's previous_hash, keyed by AUDIT_HMAC_KEY (generated at install into secrets/audit_hmac_key.env). Chains are per tenant, and the scheduled verifier (audit.chain_verify_interval_minutes, default 60, with audit.chain_verify_checkpoint) re-checks them and reports a break as a SUSPICIOUS_ACTIVITY event rather than repairing anything.

Because the SIEM stream is metadata-only, the hash arriving in your SIEM is a correlation value, not something you can recompute there: recomputation needs the full record (target user, permission ids, metadata, user agent), which the shipper deliberately does not send. To verify hashes yourself, use GET /v1/audit-logs/export, which includes the per-record integrity_hash, previous_hash and integrity_valid result — that export is gated behind the AUDIT_HMAC entitlement (a 402 response when the install's edition does not grant it).

Time-sync status (AU-8)​

Timestamps in the shipped events are only as trustworthy as the clock that wrote them, so the compliance surface exposes a read-only clock report at GET /v1/platform/compliance/time-sync (platform admin or a time-boxed auditor — the same posture as GET /v1/platform/compliance/status).

  • Host clock: on Linux the kernel adjtimex(2) state is read through ctypes (no extra dependency). STA_UNSYNC in the status word means the host clock is not disciplined by ntpd/systemd-timesyncd; maxerror and esterror (microseconds) bound the residual uncertainty. On a kernel that refuses adjtimex (WSL, a sandbox, a non-Linux host) the state is reported unknown — never guessed.
  • Worker skew: skew is server_received_at - worker_sent_at. A worker heartbeat carries its own RFC3339 timestamp, so once the worker release that reports it is deployed, each worker is scored against audit.max_clock_skew_ms; until then every worker is reported not reported.
  • Status word: degraded when the host clock is unsynced or any worker skew exceeds audit.max_clock_skew_ms. The conformance check reads the same report, so a clock that drifts shows up as a failing check and, with conformance.drift_notify on, as a drift alert.

Nothing in this endpoint opens a connection or adjusts a clock — it is a report, not a disciplinarian. Point your SIEM or monitoring layer at it on a schedule and alert on degraded; the time source itself (chrony/ntpd, or systemd-timesyncd on each host) is yours to run.