Reproduction steps:
- Configure a workspace with an egress policy (redaction required):
WS/egress_policy.json = {"version":1,"rules":[{"match":"any_sensitive","action":"redact"}]}
- Call the interactive egress guard with sensitive content while the guard's
three helper modules are unimportable (a broken/partial install, or an
in-process actor that poisoned the import cache, e.g.
sys.modules["primitives.egress_classifier"] = None):
station_llm.guard_egress_messages(
[{"role":"user","content":"Patient Jane Doe, SSN 123-45-6789, dx: melanoma"}],
provider="openai")
- Compare with the classifier RUNTIME-error branch under the SAME policy
(monkeypatch egress_classifier.probe_messages to raise).
- Run repro_egress_guard_import_fail_open_v099.py against a clean v0.99
extraction.
Expected:
Per the function's own docstring — "Fails CLOSED: if a policy is configured and
the guard raises, we deny rather than leak" — both failure modes must DENY when
an egress policy is present. A guard that is "unavailable" is not evidence the
content is safe to send.
Actual:
The two failure branches disagree under an identical configured policy:
BRANCH A (guard modules unimportable): decision='allow', raw PHI forwarded
BRANCH B (classifier raises at runtime): decision='denied' (correct)
The classifier-error branch was already fixed (community #22 / the shweta egress
fix; the fix comment is inline at station_llm.py lines 232-238) to check whether
an egress_policy.json exists and DENY when it does. The ImportError branch
immediately above it (line 227-228) was NOT given the same treatment:
except ImportError:
return ("allow", messages, {}, "egress guard unavailable")
so it returns "allow" unconditionally, forwarding the raw messages — PHI/PII
unredacted — to the model provider even when redaction is policy-mandated. This
re-opens the exact leak #22 closed, on the sibling (import-failure) path.
Honest scope / trigger:
The three modules (egress_classifier, egress_policy, egress_tokens) import only
stdlib at top level, so this does NOT fire in a healthy install. It fires on a
broken/partial deployment, or if an in-process actor poisons the import cache.
(Whether a sandboxed module can reach sys.modules to force this is a separate,
unverified question — the platform's module sandbox is already shown porous in
the filesystem/network findings, but I did not verify import-cache poisoning from
inside the sandbox, so I am not claiming it.) The reportable defect is the
fail-OPEN DIRECTION and the contract/sibling inconsistency: one branch of the
same function leaks what the other denies, contradicting the documented
fail-closed contract.
Suggested fix:
Mirror the already-correct classifier-error branch — on ImportError, deny when a
policy is configured, allow only in the no-config (dev) carve-out:
except ImportError as e:
if os.path.exists(os.path.join(WS, "egress_policy.json")):
return ("denied", messages, {}, f"egress guard unavailable, failing closed: {e}")
return ("allow", messages, {}, f"egress guard unavailable (no egress policy configured): {e}")
(Same shape the docstring already promises and the sibling branch already
implements — the two failure paths should make the same policy-gated decision.)