net-kingdom/workplans/NK-WP-0034-verification-that-verifies.md
tegwick 1da6e5457a
All checks were successful
CI Smoke / host-smoke (push) Successful in 2s
CI Smoke / container-smoke (push) Successful in 5s
Close NK-WP-0034: verify-t06 proven FAIL-then-PASS on production
Production lldap-coulomb lacked TIMEOUT/CACHE_TIMEOUT/SIZELIMIT since the
2026-08-27 unrepaired apply; the new verifier failed at resolver-tuning,
the repaired attended --apply restored them (DB readback confirmed), and
verify-t06 now passes all four proofs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 299762@bnt-lap001
Assistant-Session: d3d3cea1-869c-44f1-be2a-3d6d3550e72e
2026-09-23 21:20:19 +02:00

8.6 KiB

id type title domain repo status flavor owner topic_slug created updated related state_hub_workstream_id
NK-WP-0034 workplan Make the SSO/MFA verification actually verify infotech net-kingdom finished implementation codex infotech 2026-08-28 2026-09-23
NK-WP-0033
CUST-ADR-012
c87e142c-4eda-5c1b-84a9-ee5a848f3f63

Make the SSO/MFA verification actually verify

Goal

Close the blind spot NK-WP-0033 exposed: verification that passes while the thing it verifies cannot work.

Why this exists

reconcile-lldap-resolver-live.sh was approved, documented, revision-pinned and referenced by an incident workplan as the supported attended operation. Its first execution on 2026-08-27 found four defects, two of which meant it could never have completed a run under any circumstances.

verify-t06.sh reported success throughout.

It checked that a resolver and a realm existed. It did not check that GET /user/?realm=coulomb&username=<user> returns a user — which is the entire purpose of a resolver, and would have caught two of the four on the day they landed. The resolver was also created without TIMEOUT, CACHE_TIMEOUT and SIZELIMIT, so the WebUI had refused to save or test it for as long as it had existed, and nothing noticed.

A verification that asserts its objects exist rather than that its property holds is not a weaker check. It is a check that reports green while the system is broken, which is worse than no check, because it is believed.

Assert the property, not the objects

id: NK-WP-0034-T01
status: done
priority: high
state_hub_task_id: "9b5d8034-0ef1-53f9-ad4d-37de536bdc62"

Rewrite verify-t06.sh so it proves what the realm and resolver are for: that a known user resolves through the realm, and that a privacyIDEA MFA validation against that user succeeds. Existence checks may stay as diagnostics; they may not stand in for the outcome.

Include the resolver's numeric parameters in what is asserted — a resolver the WebUI cannot save or test is misconfigured whether or not it currently answers.

Acceptance: with the tuning fields cleared by hand, verify-t06.sh fails. With them restored, it passes. Demonstrated in both directions, not argued.

Audit the other verify scripts for the same shape

id: NK-WP-0034-T02
status: done
priority: medium
state_hub_task_id: "0bd09a40-8b4d-5f9d-8d4c-81c55aa0c565"

verify-t02.sh through verify-t08.sh were written the same way by the same hand. Read each and record, per script, the property it is supposed to establish and whether it establishes it.

The output is a list, not a rewrite. Fixing them all is a larger decision than this task should make on its own.

Acceptance: one table — script, intended property, actual assertion, verdict.

Label every attended procedure with its exercise status

id: NK-WP-0034-T03
status: done
priority: medium
state_hub_task_id: "becb4dce-4971-5755-888f-5276c4a56fe7"

docs/attended-procedure-standard.md requires an Exercise status: line in each runbook. There are 32 shell scripts under sso-mfa/k8s/; the reconciliation runbook is the only one carrying the line.

Work through them and record, honestly, which have ever been run. "Unknown" is a permitted answer and a more useful one than a guess — the point is that an operator learns before running whether they are the first.

Acceptance: every attended procedure's runbook carries the line; scripts with no runbook are listed as a separate finding.

One helper per API

id: NK-WP-0034-T04
status: done
priority: low
state_hub_task_id: "d43ae8e8-eda6-5c1d-9e1b-01966b0e92da"

The Content-Type-on-bodyless-GET defect was found and fixed in bootstrap-realm.sh's pi_api, with a comment explaining it, and then reintroduced verbatim in a script written later in the same directory.

Extract one privacyIDEA request helper and have both scripts use it. A fix applied to a copy is a fix that expires the next time someone writes a script.

Acceptance: one implementation of the privacyIDEA request path in sso-mfa/k8s/privacyidea/; both callers use it; the GET behaviour has a test that fails if the header returns.

Review and implementation — 2026-09-05

Reviewed the proposal against current code and implemented all repository-local changes. T01 waits only for the attended scratch/provider exercise required by its acceptance. No live credentials were requested and no production resolver was changed.

  • T01: replaced warning-only T06 success with an attended known-user and TOTP/HOTP verifier, requiring realm binding, numeric tuning, exact user and resolver, and token-backed success. Credentials are prompted in process; missing attendance fails. Updated realm-repair and DR callers. Eighteen tests exercise the actual shared HTTP transport, including clearing and restoring every tuning field and rejecting passthrough/static passwords. See docs/verify-t06.md for the remaining attended acceptance and limits.
  • T02: completed the per-script property/assertion/verdict table in docs/verification-audit-2026-09-05.md. Other verify scripts retain their implementations; their gaps are explicitly recorded for follow-up.
  • T03: inventoried every shell script under sso-mfa/k8s/, added exercise headers to component runbooks, recorded unknown history and scripts without runbooks in docs/attended-procedure-inventory.md. Corrected the reconciliation runbook's attempted-run claim: no successful completion receipt exists.
  • T04: privacyidea/pi_api.py now supplies the request transport to both realm bootstrap and resolver reconciliation (and T06). Authentication also uses the shared transport in bootstrap. Its shell adapter keeps the token/body off argv. Fixed bootstrap's first-success set -e counter exit found during the review. The HTTP fixture rejects JSON Content-Type on GET and verifies the real Python and shell adapter requests.

Remaining gate: an attended operator runs the scratch target failure/restoration procedure and records receipts. An automated fixture is not reported as an exercised provider run. NK-WP-0033 still needs its own incident receipt and predecessor-disposition ruling.

Validation: python3 -m pytest tests tools -q passed 106 tests, including 18 privacyIDEA and 15 cadence cases. Ruff lint/format, shell syntax, embedded Python compilation and git diff --check passed.

State Hub reconciliation was attempted with both the installed CLI and current checkout. Full reconciliation remains pending because API queries/writes timed out or returned connection-refused errors. Generated index/intake metadata was reviewed; the source files remain authoritative.

Attended acceptance — 2026-09-23 (T01 done)

T01's acceptance was demonstrated in both directions on production privacyIDEA. The fields were not cleared on purpose: production had been missing them since 2026-08-27. The operator was Bernd Worsch, using railiance01 checkout 6096c395.

  1. FAIL, fields absent. verify-t06.sh --user platform-root failed twice (runs at 18:41:54 and 18:42:22 server time). The privacyIDEA audit shows pi-admin auth, GET /realm/ and GET /resolver/lldap-coulomb succeeding, and no user lookup after them. That places the failure in the resolver-tuning phase. The live lldap-coulomb config had 11 keys and no TIMEOUT, CACHE_TIMEOUT or SIZELIMIT. The unrepaired reconcile-lldap-resolver-live.sh --apply of 2026-08-27 (NK-WP-0033 defect 3) had dropped them, and the old T06 passed regardless.
  2. Restore. With operator approval, the repaired reconcile-lldap-resolver-live.sh --apply --predecessor-unavailable ran at 19:14:45 and made one resolver write. Before the write, every non-secret value was confirmed identical to the declared body, and the bind value came from operators/lldap/admin version 1. Receipt: PASS: resolver lookup, privacyIDEA MFA, predecessor denial=NOT-PROVEN, readiness, health, cleanup=PASS. A database readback shows TIMEOUT=5, CACHE_TIMEOUT=120 and SIZELIMIT=500, 14 keys in total.
  3. PASS, fields restored. verify-t06.sh --user platform-root returned {"phase": "complete", "proofs": ["realm-binding", "resolver-tuning", "known-user-lookup", "token-backed-mfa"], "result": "PASS"}.

The scratch-target exercise planned earlier in the session is no longer needed. The failure direction was observed on the real target, attributed through the audit log, and reversed by the declared remediation.

Two runs in the same session failed at authentication because the wrong pi-admin value was entered (audit: "Wrong credentials"). Neither was a defect in the verifier.