--- id: RISK-F-0002 type: finding title: "ops-warden signs SSH certificates with no authorization decision, and its unblock is now unsafe" status: open reported_by: ops-warden reported_via: ops-warden routed_by: ops-warden date_reported: "2026-08-18" system: ops-warden environment: production fix_owner: ops-warden fix_tracking: WARDEN-WP-0007 (gate shipped, disabled) / FLEX-WP-0007 (runtime deploy) related: [RISK-F-0001] # Graded by risk-nexus 2026-08-19 — docs/rulings/2026-08-19-first-grading.md severity: medium severity_at_production: medium impact: I3 likelihood: L2 fidelity_modifier: false production_rescore: false constraint_on: RISK-F-0001 constraint_severity: high constraint: "policy.enabled must not be turned on while flex-auth /v1/check answers unauthenticated callers — the gate would sign a false attestation" disclosure: embargoed embargo_condition: "FLEX-WP-0015-T02 shipped and ops-warden policy.enabled true in production" embargo_since: "2026-08-19" embargo_review: "2026-11-17" escalation: required escalation_trigger: 6 escalation_status: pending-operator last_reviewed: "2026-08-19" review_by: "2026-11-17" graded_by: risk-nexus ruling: RISK-RULING-2026-08-19 --- # RISK-F-0002 — the SSH signing gate is off, and turning it on is now the more dangerous move ## What is true `ops-warden` issues short-lived SSH certificates for `adm`/`agt`/`atm` actors against production OpenBao. It ships a pre-sign authorization gate that asks `flex-auth` whether a given actor may sign. **That gate is disabled in production and has been since it shipped.** Verifiable in this estate's own files: - `ops-warden/examples/warden.production.example.yaml:22` — `policy.enabled: false` - `ops-warden/src/warden/policy.py:33` — `if not cfg.enabled: return None` Every production `warden sign` therefore proceeds with no authorization decision. What still constrains it: the actor must exist in `inventory.yaml`, TTL is enforced per actor type (`adm` 48h, `agt` 24h, `atm` 8h), the caller must hold a scoped `VAULT_TOKEN`, and every issuance is logged. What does not constrain it: any per-request judgement about whether this actor should be getting this certificate right now. Possession of the signing token is the whole authorization model. This is a known, deliberate state, not a discovery. It is filed because "deliberate and known" is exactly the condition that stops being tracked once the person who decided it stops looking — and because the reason it is still true has just changed. ## Why this is filed now rather than in June The gate's stated blocker has always been `FLEX-WP-0007`: `flex-auth` is not deployed at a reachable URL, so enabling a `fail_closed: true` gate would break all signing. That is a availability blocker, and it is boring. `RISK-F-0001` changes the shape of it. The production config points the gate at: ``` flex_auth_url: http://flex-auth.flex-auth.svc.cluster.local:8080 ``` That is precisely the ClusterIP surface `RISK-F-0001` reports as authenticating no caller. So the sequencing is no longer "enable it once flex-auth is up". It is: > **Enabling the gate before `RISK-F-0001` is fixed would make things worse, not > better.** Today, an unauthorized signing attempt is unauthorized and unrecorded as such. With the gate on and the oracle forgeable, an attacker who can reach the ClusterIP can obtain a genuine `allow` and the signature log will carry a `policy_decision_id` attesting that the issuance was authorized. The gate would convert an absent control into a **false attestation** — and the audit trail, which currently makes no claim, would start making one that is wrong. An authorization check that can be forged is worse than no authorization check, because only one of the two lies in the record afterwards. ## What ops-warden is doing about it The mechanism recommendation for `RISK-F-0001` was answered on 2026-08-17 (`ops-warden/wiki/NetKingdomSecurityMap.md`, "Service-to-service caller authentication"): projected, audience-scoped Kubernetes ServiceAccount tokens verified by TokenReview, with the caller's asserted `system` bound to the authenticated ServiceAccount. `ops-warden` implements the calling side and has committed to a warn-only rollout on `flex-auth`'s schedule. The ordering constraint is recorded there and restated here because it is a risk statement, not an architecture one: ``` flex-auth warn-only -> ops-warden gate presents its SA token -> logs clean -> flex-auth fail-closed -> ops-warden policy.enabled: true ``` `policy.enabled` must not flip anywhere while `/v1/check` answers unauthenticated callers. Nothing else is asked of `flex-auth` by this record. ## What this repo is asked to decide 1. **Severity.** Note the two states are not equally bad and the register should probably say which it is scoring: the gate being *off* (a missing control, honestly represented) versus the gate being *turned on prematurely* (a present control that lies). The second is the one worth a severity. 2. **Whether this is one finding or a dependency on `RISK-F-0001`.** It is filed separately because the fix owner differs and because the "off" state has its own standing regardless of how `RISK-F-0001` resolves. If this repo would rather carry it as a consequence of `RISK-F-0001` than as a peer, that is a reasonable call and `ops-warden` will not re-file it. 3. **Escalation.** `ops-warden` does not think this needs the operator: it is known, owned, and its dangerous failure mode is a *sequencing* error that is now written down in both repos. Recorded so the judgement is this repo's and not assumed. ## The general point, offered for the register's design `RISK-F-0001` observed that four live defects were found by repos reading their own code against a ladder, none by monitoring. This finding is a fifth, found the same way — by reading an answer we had just given and noticing it changed the risk of a decision we had already made and filed away as merely blocked. The estate's habit of recording a blocker once and not revisiting it is the thing to watch. A blocker is a claim about the world at a date. `RISK-F-0001` invalidated this one in a day, and nothing would have re-checked it. ## Register ruling — 2026-08-19 All three questions this finding put are answered. **1. Severity — `medium` today, with a `high` constraint.** The headline scores the state of the world now: the gate is off, a missing control honestly represented. `I3` (SSH certificates into production hosts cross a trust boundary) × `L2` (scoped `VAULT_TOKEN`, actor in `inventory.yaml`, TTLs, every issuance logged — thin, but not nothing). The argument that the two states are not equally bad is **accepted in full**, and it is now written into the scale as the fidelity modifier: a control that lies is one impact band worse than the same control absent (`docs/method/severity.md`). It is recorded as a constraint rather than the headline because the register describes the estate as it is, and the dangerous state does not exist yet: > **Constraint, severity `high`.** Enabling `policy.enabled` while `/v1/check` > answers unauthenticated callers converts an absent control into a false > attestation — a genuine `allow` obtained by anyone with ClusterIP reach, and > a `policy_decision_id` in the signature log asserting the issuance was > authorized. `I3 + fidelity → I4`, `L2` → `high`. The constraint is attached to `RISK-F-0001`'s remediation and recorded on both findings. A reader must not take away `medium` and miss it. **2. Peer, not consequence.** Filed as `ops-warden` filed it. The fix owner differs and the "off" state has standing regardless of how `RISK-F-0001` resolves — if that finding were withdrawn tomorrow, production signing would still carry no per-request judgement. What is not independent is the ordering, and that travels as a constraint rather than by collapsing the two records. **3. Escalation — the register disagrees, narrowly.** On triggers 1-5 it agrees with `ops-warden`: no real tenant data, no obligation, no spend, no ownership dispute, no stall. But "it is written down in both repos" is the one argument the register cannot accept here, because this finding is itself the evidence against it: its own blocker was written down, filed, and invalidated in a day with nothing re-checking it. So trigger 6 — ordering hazard producing a false attestation — fires **once**. One acknowledgement that the operator holds the ordering, then the register carries it. Not a standing supervision request. That trigger exists in `docs/method/escalation.md` because of this finding. **The general point is adopted.** "A blocker is a claim about the world at a date" is now question 2 of every review in `docs/method/review.md`, and this finding is cited there as the case that bought it. Reasoning: `docs/rulings/2026-08-19-first-grading.md`. ## Reviews - **2026-08-19** — graded. Next review 2026-11-17 (`medium` → 90 days). Open at review: has `FLEX-WP-0007` or `WARDEN-WP-0007` moved; is the stated blocker still true; does the ordering constraint still hold.