Make the failure matrix an executable harness
Some checks failed
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Has been cancelled

AUDIT-WP-0005-T05 (progress). scripts/failure_matrix.py, make failure-matrix.

Two modes: MODE=local stands up PostgreSQL and the receiver in Docker and runs
all 15 scenarios including infrastructure disruption; MODE=remote targets a
deployed receiver and skips disruption unless DISRUPT=1, since restarting a
production database is not this script's call.

Rehearsed locally: 15 passed, 0 failed. Delivery and reconciliation, rejection
and dead-letter visibility, redaction with per-path counting, correlation
lookup, privilege separation both directions, credential rotation mid-ingestion
with no delivery gap, operator replay and duplicate replay, receiver
unavailability with sender retry, and a database restart mid-ingestion where 5
of 7 attempts were acknowledged and all 5 survived.

Two deliberate choices. Stored-record counts are read straight from the
database rather than through the API, because the assertion is about what is
stored and asking the service to vouch for itself is weaker evidence. The
retry policy retries 503/500 and treats 400/401/403/409 as terminal, which is
the documented response contract - so what is under test is a sender that
follows it.

Harness credibility checked rather than assumed: exit 0 on success, exit 2
against an unreachable receiver rather than passing silently, and the evidence
JSON carries no tokens, credentials or event payloads so it can go to
NK-WP-0024 as-is.

The local rehearsal is not a substitute for the live run: it does not exercise
CNPG failover, NetworkPolicy enforcement, or OpenBao-leased credentials.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-10 17:49:32 +02:00
parent 2b0bbdecb4
commit 4de672aaf0
6 changed files with 564 additions and 3 deletions

View file

@ -273,7 +273,7 @@ empty-source case that is today's actual situation.
```task
id: AUDIT-WP-0005-T05
status: todo
status: progress
priority: high
state_hub_task_id: "1da30fec-b9f1-4be0-b42c-15797a8c4392"
```
@ -294,6 +294,46 @@ Done when the matrix has been run against the deployed system and each
outcome recorded, including any case where behaviour differed from the
documented contract.
Progress 2026-08-10: the matrix is now an executable harness rather than a
checklist — `scripts/failure_matrix.py`, `make failure-matrix`. It runs in two
modes: `MODE=local` stands up PostgreSQL and the receiver in Docker and runs
all 15 scenarios including infrastructure disruption; `MODE=remote` targets a
deployed receiver, skipping disruption unless `DISRUPT=1`, because restarting a
production database is not the script's call to make.
Rehearsed locally 2026-08-10: **15 passed, 0 failed, 0 skipped.**
- S01-S03 delivery: accepted once, resubmission reconciles as duplicate, same
id with a different payload conflicts — one record throughout
- S04-S06 rejection: cross-tenant refused and terminal for the sender (one
attempt, no retry storm), visible as a dead letter, bad credential rejected
- S07-S08 redaction: secret-shaped field masked with the event still stored,
counted by field path
- S09 correlation lookup returns all three related events
- S10-S11 privilege separation both directions
- S12 credential rotation mid-ingestion with no delivery gap
- S13 operator replay and duplicate replay both reconcile, still one record
- S14 receiver unavailable: sender retries to exactly one record after recovery
- S15 database restart mid-ingestion: of 7 attempts, 5 were acknowledged and
all 5 survived the restart; the rest returned 503, which is correct
Two things the harness does deliberately. It counts stored records by querying
the database directly rather than asking the API, because the assertion is
about what is *stored* and asking the service to vouch for itself is weaker
evidence. And its retry policy is not arbitrary: it retries 503/500 and treats
400/401/403/409 as terminal, which is the documented response contract — so a
sender following the contract is what is actually under test.
Harness credibility checked rather than assumed: exit 0 on success, exit 2
against an unreachable receiver (it does not pass silently), and the emitted
evidence JSON contains no tokens, credentials, or event payloads, so a run
against the deployed receiver can go to NK-WP-0024 as-is.
Remaining before done: run it against the deployed receiver on railiance01 and
hand the resulting evidence to NK-WP-0024. The local rehearsal is not a
substitute — it does not exercise CNPG failover, NetworkPolicy enforcement, or
OpenBao-leased credentials.
## T06 - Operational handover
```task