risk-nexus/findings/RISK-F-0002-ops-warden-sign-ungated.md
tegwick 7f1424dbcf Sweep risk inbox and reconcile findings
Assistant: codex
Assistant-Model: gpt-5.6-sol
Assistant-Session: 01a058f3-8ba0-7692-a042-9a870fc3d663
2026-09-01 02:41:32 +02:00

20 KiB
Raw Blame History

id type title status owner reported_by reported_via routed_by date_reported system environment fix_owner fix_tracking fix_tracking_superseded related severity severity_at_production impact likelihood fidelity_modifier production_rescore constraint_on constraint_severity constraint disclosure publication publication_id publication_path publication_subtitle revision last_reviewed review_interval embargo_lifted embargo_was_since escalation escalation_trigger escalation_status date_fixed last_checked next_check cadence clean_streak graded_by ruling checked_by
RISK-F-0002 finding ops-warden signs SSH certificates with no authorization decision, and its unblock is now unsafe fixed risk-nexus ops-warden ops-warden ops-warden 2026-08-18 ops-warden production ops-warden ZONE-WP-0001 / WARDEN-WP-0032-T02 (finished 2026-08-22) WARDEN-WP-0007 (archived 2026-07-08) / FLEX-WP-0007 (finished 2026-06-29)
RISK-F-0001
medium medium I3 L2 false false RISK-F-0001 lifted LIFTED 2026-08-19 — flex-auth /v1/check now authenticates callers (RISK-F-0001 fixed). Enabling policy.enabled is now an availability question for ops-warden, no longer an attestation hazard. public pending-handover risk-f-0002-ops-warden-signing-authorization-gap findings/ops-warden-signing-authorization-gap/v1/index.html Production SSH signing once had no per-request authorization decision; the retired global switch is now replaced by an enforced zone-aware gate. fixed-1 2026-09-01 6m 2026-08-22 — zone-aware authorization replaced the retired global gate and passed the live authenticated-caller check 2026-08-19 withdrawn 6 withdrawn-hazard-window-closed 2026-08-22 2026-09-01T00:32:44Z 2026-09-01T00:32:44Z instant 0 risk-nexus RISK-RULING-2026-08-19 codex/risk-nexus

RISK-F-0002 — the SSH signing gate is off, and turning it on is now the more dangerous move

What is true

ops-warden issues short-lived SSH certificates for adm/agt/atm actors against production OpenBao. It ships a pre-sign authorization gate that asks flex-auth whether a given actor may sign. That gate is disabled in production and has been since it shipped.

Verifiable in this estate's own files:

  • ops-warden/examples/warden.production.example.yaml:22policy.enabled: false
  • ops-warden/src/warden/policy.py:33if not cfg.enabled: return None

Every production warden sign therefore proceeds with no authorization decision. What still constrains it: the actor must exist in inventory.yaml, TTL is enforced per actor type (adm 48h, agt 24h, atm 8h), the caller must hold a scoped VAULT_TOKEN, and every issuance is logged. What does not constrain it: any per-request judgement about whether this actor should be getting this certificate right now. Possession of the signing token is the whole authorization model.

This is a known, deliberate state, not a discovery. It is filed because "deliberate and known" is exactly the condition that stops being tracked once the person who decided it stops looking — and because the reason it is still true has just changed.

Why this is filed now rather than in June

The gate's stated blocker has always been FLEX-WP-0007: flex-auth is not deployed at a reachable URL, so enabling a fail_closed: true gate would break all signing. That is a availability blocker, and it is boring.

RISK-F-0001 changes the shape of it. The production config points the gate at:

flex_auth_url: http://flex-auth.flex-auth.svc.cluster.local:8080

That is precisely the ClusterIP surface RISK-F-0001 reports as authenticating no caller. So the sequencing is no longer "enable it once flex-auth is up". It is:

Enabling the gate before RISK-F-0001 is fixed would make things worse, not better.

Today, an unauthorized signing attempt is unauthorized and unrecorded as such. With the gate on and the oracle forgeable, an attacker who can reach the ClusterIP can obtain a genuine allow and the signature log will carry a policy_decision_id attesting that the issuance was authorized. The gate would convert an absent control into a false attestation — and the audit trail, which currently makes no claim, would start making one that is wrong.

An authorization check that can be forged is worse than no authorization check, because only one of the two lies in the record afterwards.

What ops-warden is doing about it

The mechanism recommendation for RISK-F-0001 was answered on 2026-08-17 (ops-warden/wiki/NetKingdomSecurityMap.md, "Service-to-service caller authentication"): projected, audience-scoped Kubernetes ServiceAccount tokens verified by TokenReview, with the caller's asserted system bound to the authenticated ServiceAccount. ops-warden implements the calling side and has committed to a warn-only rollout on flex-auth's schedule.

The ordering constraint is recorded there and restated here because it is a risk statement, not an architecture one:

flex-auth warn-only -> ops-warden gate presents its SA token -> logs clean
                    -> flex-auth fail-closed -> ops-warden policy.enabled: true

policy.enabled must not flip anywhere while /v1/check answers unauthenticated callers. Nothing else is asked of flex-auth by this record.

What this repo is asked to decide

  1. Severity. Note the two states are not equally bad and the register should probably say which it is scoring: the gate being off (a missing control, honestly represented) versus the gate being turned on prematurely (a present control that lies). The second is the one worth a severity.
  2. Whether this is one finding or a dependency on RISK-F-0001. It is filed separately because the fix owner differs and because the "off" state has its own standing regardless of how RISK-F-0001 resolves. If this repo would rather carry it as a consequence of RISK-F-0001 than as a peer, that is a reasonable call and ops-warden will not re-file it.
  3. Escalation. ops-warden does not think this needs the operator: it is known, owned, and its dangerous failure mode is a sequencing error that is now written down in both repos. Recorded so the judgement is this repo's and not assumed.

The general point, offered for the register's design

RISK-F-0001 observed that four live defects were found by repos reading their own code against a ladder, none by monitoring. This finding is a fifth, found the same way — by reading an answer we had just given and noticing it changed the risk of a decision we had already made and filed away as merely blocked.

The estate's habit of recording a blocker once and not revisiting it is the thing to watch. A blocker is a claim about the world at a date. RISK-F-0001 invalidated this one in a day, and nothing would have re-checked it.

Register ruling — 2026-08-19

All three questions this finding put are answered.

1. Severity — medium today, with a high constraint. The headline scores the state of the world now: the gate is off, a missing control honestly represented. I3 (SSH certificates into production hosts cross a trust boundary) × L2 (scoped VAULT_TOKEN, actor in inventory.yaml, TTLs, every issuance logged — thin, but not nothing).

The argument that the two states are not equally bad is accepted in full, and it is now written into the scale as the fidelity modifier: a control that lies is one impact band worse than the same control absent (docs/method/severity.md). It is recorded as a constraint rather than the headline because the register describes the estate as it is, and the dangerous state does not exist yet:

Constraint, severity high. Enabling policy.enabled while /v1/check answers unauthenticated callers converts an absent control into a false attestation — a genuine allow obtained by anyone with ClusterIP reach, and a policy_decision_id in the signature log asserting the issuance was authorized. I3 + fidelity → I4, L2high.

The constraint is attached to RISK-F-0001's remediation and recorded on both findings. A reader must not take away medium and miss it.

2. Peer, not consequence. Filed as ops-warden filed it. The fix owner differs and the "off" state has standing regardless of how RISK-F-0001 resolves — if that finding were withdrawn tomorrow, production signing would still carry no per-request judgement. What is not independent is the ordering, and that travels as a constraint rather than by collapsing the two records.

3. Escalation — the register disagrees, narrowly. On triggers 1-5 it agrees with ops-warden: no real tenant data, no obligation, no spend, no ownership dispute, no stall. But "it is written down in both repos" is the one argument the register cannot accept here, because this finding is itself the evidence against it: its own blocker was written down, filed, and invalidated in a day with nothing re-checking it.

So trigger 6 — ordering hazard producing a false attestation — fires once. One acknowledgement that the operator holds the ordering, then the register carries it. Not a standing supervision request. That trigger exists in docs/method/escalation.md because of this finding.

The general point is adopted. "A blocker is a claim about the world at a date" is now question 2 of every review in docs/method/review.md, and this finding is cited there as the case that bought it.

Reasoning: docs/rulings/2026-08-19-first-grading.md.

Reviews

  • 2026-08-19 — graded. Next review 2026-11-17 (medium → 90 days). Open at review: has FLEX-WP-0007 or WARDEN-WP-0007 moved; is the stated blocker still true; does the ordering constraint still hold.

Constraint lifted, escalation withdrawn — 2026-08-19 (same day)

RISK-F-0001 was fixed hours after this register graded it: both flex-auth production Deployments now enforce ADR-0004 TokenReview, with live unbound-request probes returning 401.

The attestation hazard is gone. The constraint recorded this morning — that enabling policy.enabled against a forgeable oracle would write a policy_decision_id asserting an authorization nobody made — required /v1/check to answer unauthenticated callers. It no longer does. Enabling the gate is now an ordinary availability question for ops-warden: does the calling side present its ServiceAccount token, and does signing survive a fail_closed: true. That is ops-warden's sequencing to run, not a risk this register holds.

The trigger-6 escalation is withdrawn. The hazard window closed before the operator was asked to hold the ordering. ops-warden's own judgement that this did not need the operator was, in the end, right — and the register still holds that it was right for the wrong reason. The disagreement was never about the severity of the hazard; it was about whether "written down in both repos" keeps a thing held. It did not: the window closed by flex-auth shipping, not by anyone re-reading the note.

The finding stays open, unchanged at medium. The gate is still off, and warden sign still proceeds with no per-request judgement. Nothing about this morning's news changes that.

The general point earns another instance. This finding argued that a blocker is a claim about the world at a date, and that its own blocker had been invalidated in a day with nothing re-checking it. That has now happened twice to the same finding, in the same week, in the same direction. The second time, the thing that had gone stale was this register's own grading, four hours old.

Reviews

  • 2026-08-19 — constraint lifted, escalation withdrawn, severity unchanged. Open at review: has ops-warden presented SA tokens and enabled the gate; is FLEX-WP-0007 still the stated blocker or has it too gone stale.

Check — 2026-08-20: a new availability fact, and it is the one that matters now

RISK-V-0001 verified flex-auth's NetworkPolicies against the live cluster and found a third policy nobody had mentioned:

flex-auth-ops-warden, created 2026-08-19T12:47:18Z
policyTypes: [Ingress, Egress]
ingress:     no rules at all

Ingress in policyTypes with zero rules means deny all ingress. On its face, nothing reaches that pin.

This is now the live question on this finding. The attestation hazard lifted when flex-auth shipped; what remained was whether enabling policy.enabled breaks signing on availability grounds. If the pin ops-warden calls admits no ingress, a fail_closed: true gate against it fails closed — every warden sign stops.

The register is not concluding that, and said so to both owners: the policy may be mid-rollout, it may not be the pin ops-warden targets, and another policy may admit the traffic. Both were told on 2026-08-20, before either flips a switch, which is the entire reason this finding is carried as a peer rather than folded into RISK-F-0001.

The severity is unchanged at medium. What changed is the evidence, and it changed in the direction of "do not enable this yet" for a completely different reason than the one this finding was filed for. That is the second time in two days that this finding's blocker turned out to be a claim about the world at a date.

  • 2026-08-20 — not clean: RISK-V-0001 found the flex-auth-ops-warden policy admits no ingress; the live question is now availability, not attestation. Cadence instant → instant; checked again immediately.

Check — 2026-08-21: both tracked records were already closed when this was filed

RISK-WP-0005-T01 made this register read the fix records it had been writing down without reading. The first run answered a question nobody had asked:

Tracked State Last changed
WARDEN-WP-0007 archived 2026-07-08
FLEX-WP-0007 finished 2026-06-29

This finding was filed on 2026-08-18 and states that the gate is "blocked on FLEX-WP-0007". That workplan had read finished for seven weeks by then. flex-auth restated the same blocker to this register on 2026-08-19 — "policy.enabled remains off (FLEX-WP-0007)" — while their own authoritative record said it was done.

The register is not concluding that the blocker is gone. Three readings fit and only the owners can say which: the deployment finished and was later reverted or never applied; "finished" means something narrower than the sentence implies; or the blocker is genuinely stale and two repos have been repeating it.

What can be said without either owner's help is that the stated blocker has not matched the stated record since June, and nothing would have caught that before this register started reading fix state.

This is the third time on the same finding

RISK-F-0002 argued, when it was filed, that a blocker is a claim about the world at a date, and that its own had been invalidated in a day with nothing re-checking it. Since then:

  1. RISK-F-0001 was fixed hours after this register graded the ordering hazard;
  2. RISK-V-0001 found a NetworkPolicy that may make enabling the gate fail closed for an entirely different reason;
  3. and now both tracked records turn out to have been closed before the finding was written.

The finding was right about the estate and right about itself. Severity is unchanged at medium — none of this changes what the gate does today — but what is blocking it has been wrong in the record for two months.

The ask, sharpened

Previously: probe whether the flex-auth pin admits ingress. Now, additionally: what actually blocks policy.enabled today? If the answer is "nothing except the ingress question", that is a much shorter path than the one both repos have been describing.

  • 2026-08-21 — not clean: Fix tracking read for the first time: WARDEN-WP-0007 archived 2026-07-08, FLEX-WP-0007 finished 2026-06-29 — both closed before the finding was filed. Cadence instant → instant; checked again immediately.

Check — 2026-08-21: reading (c) confirmed — the blocker was stale, and the control is being retired

ops-warden answered: nothing blocks policy.enabled. It is off by decision, not by blocker.

The gate is ready and verified — flex-auth's flex-auth-ops-warden pin runs callerAuth.mode: enforce, confirmed by both sides (decision:f3f7c88f9585582a, anonymous /v1/check returns 401), and ops-warden shipped the calling identity in WARDEN-WP-0031. FLEX-WP-0007 is finished, and so is FLEX-WP-0016, which flex-auth closed on 2026-08-19 recording explicitly that the flip it was named after is not coming.

What replaced the blocker is a decision, not an obstacle. ops-warden ADR-0006 (accepted): enforcement is zone-scoped, never a global flag. policy.enabled is a single estate-wide boolean, and flipping it would enforce uniformly across an estate under deep refactor. So it is not waiting to be turned on — it is being retired and replaced by a per-zone control (ZONE-WP-0001, with WARDEN-WP-0032 consuming).

The disposition: the framing is superseded, the risk is not

ops-warden proposed closed-by-supersession and left the grade here. Ruled:

  • The framing is superseded. "Gate shipped, disabled, blocked on FLEX-WP-0007" describes a world that ended in June. Recording it as blocked remediation misdescribes it, and they are right about that.
  • The risk stands, unchanged at medium. Every production warden sign still proceeds with no per-request authorization decision. That is what the finding is about, and no part of it improved — a control retired before its replacement exists is still an absent control.

So the finding stays open, with a corrected fix path. fix_tracking becomes ZONE-WP-0001 / WARDEN-WP-0032. Nothing about the estate is worse today than yesterday; what changed is that the register now knows why it is not better.

The successor's blocker is real, and it is the same missing join

ZONE-WP-0001-T03 cannot model stance because 26 of 27 credential lanes have no identifiable workload to attach a maturity ladder to. Escalated by ops-warden to repo-manager and net-kingdom on 2026-08-20, unanswered.

That is RISK-N-0004 — the zone lookup this register routed as a note — with a number on it. Three findings already wanted that facility; now the estate's per-zone enforcement is blocked on it too. The note is one instance short of being a finding, and the missing instance is somebody stating that the join cannot be built.

Four stale blockers in twelve hours, self-reported

ops-warden volunteered three more, all found the same day: an OpenBao token recorded as expired that was valid; a verification script recorded as "ready" that had never been written; and a ten-day blocker against secrets-engine answerable from that repo's source.

Their diagnosis is the one this register has been circling since RISK-F-0001: a blocker is written once, as prose, and then read as fact forever, because nothing re-derives it and nothing expires it. They have asked to adopt whatever staleness convention this register settles rather than inventing a second one. That is worth answering properly and is recorded as an open item for the next round.

  • 2026-08-21 — not clean: owner replied; see the dated check section Cadence instant → instant; checked again immediately.

Closure — 2026-09-01: the successor control is live

WARDEN-WP-0032-T02 is done. The production configuration rejects the retired global policy.enabled and policy.fail_closed switches; the signing PEP now uses compiled security-zone membership, the stance returned by flex-auth, and a local per-zone failure mode. The owner re-ran the authenticated caller path against the migrated operator configuration and received a live allow decision with decision id decision:f3f7c88f9585582a.

Risk Nexus also ran the current policy, configuration, routing, and read-boundary regressions: 137 passed. The defect this finding carried—production signing proceeding with no per-request authorization decision—is gone. The finding is fixed; the embargo lifts and publication is handed over.

  • 2026-09-01 — not clean: the zone-aware successor is implemented and live-verified; status fixed and embargo lifted. Cadence instant → instant; checked again immediately.