Add audit maintenance, verified recovery and reproducible verification
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e6f1-443f-7783-9920-a16b2ffc467f
This commit is contained in:
parent
820f0ff7b5
commit
3166da1d2e
13 changed files with 413 additions and 5 deletions
71
docs/acknowledgment-operations.md
Normal file
71
docs/acknowledgment-operations.md
Normal file
|
|
@ -0,0 +1,71 @@
|
|||
# Acknowledgment maintenance and recovery
|
||||
|
||||
These operations belong to RTEL-WP-0002-T04. Run with the database owner's OS
|
||||
identity, against a private owned directory (0700) and database (0600). No
|
||||
credential or browser login is needed for filesystem administration. This is
|
||||
not a public HTTP administration interface. Database access confers access to
|
||||
acknowledgment evidence; restrict it accordingly.
|
||||
|
||||
## Inspect and retry audit delivery
|
||||
|
||||
```sh
|
||||
python3 scripts/alert_admin.py --database /state/private/ack.sqlite status
|
||||
python3 scripts/alert_admin.py --database /state/private/ack.sqlite list-blocked
|
||||
python3 scripts/alert_admin.py --database /state/private/ack.sqlite requeue --event-id EVENT_UUID
|
||||
```
|
||||
|
||||
Status reports pending, blocked and delivered counts. Listing returns at most
|
||||
100 blocked event IDs (override with `--limit`, maximum 1000); it never prints
|
||||
actor identities, credentials or event payloads. Repair the receiver schema,
|
||||
authorization or custody fault before explicitly requeuing a particular event.
|
||||
Requeue only changes blocked → pending. The running worker retries within its
|
||||
next pass using the original event ID and exact payload. It cannot rewrite an
|
||||
acknowledgment or declare delivery successful. Exit 2 means the selected event
|
||||
was not blocked (including unknown or already delivered); exit 1 is a refusal.
|
||||
|
||||
Private GET `/metrics` exports:
|
||||
|
||||
- `railiance_ack_audit_events{status="pending|blocked|delivered"}` (all three
|
||||
series, including zero; the label takes one of the three individual values).
|
||||
- `railiance_ack_audit_last_drain_timestamp_seconds` (zero before the first pass).
|
||||
|
||||
No actor, alert or event IDs are metric labels. A completed local drain does not
|
||||
prove an idle audit receiver is reachable. Keep this endpoint on the private
|
||||
Service: the public Ingress remains `/ack/` only. Package integration must admit
|
||||
the Prometheus scraper through network policy and configure its scrape and
|
||||
alerts. Alert on blocked debt, sustained pending debt and a stale worker;
|
||||
independent receiver readback/reconciliation remains an activation gate.
|
||||
|
||||
## Backup and restore
|
||||
|
||||
```sh
|
||||
python3 scripts/alert_admin.py --database /state/private/ack.sqlite backup --destination /backup/private/ack-snapshot.sqlite
|
||||
python3 scripts/alert_admin.py --database /backup/private/ack-snapshot.sqlite restore --destination /state/private/ack-restored.sqlite
|
||||
```
|
||||
|
||||
Create destination directories beforehand with mode 0700 and the service's
|
||||
owner. Backup uses SQLite's online backup API, so it captures committed
|
||||
acknowledgments and their outbox atomically even while the service is running.
|
||||
A successful command validates SQLite integrity, foreign keys, required
|
||||
immutability triggers and acknowledgment/outbox correspondence, then fsyncs the
|
||||
snapshot. Existing destinations, symlinks and unsafe permissions are refused.
|
||||
Treat interrupted commands as unverified; retain and inspect their destination
|
||||
before choosing a new path. These checks detect structural problems, not
|
||||
malicious alteration by an administrator.
|
||||
|
||||
For recovery, stop the runtime and retain the current database first. Restore
|
||||
to a **new path**, change the package configuration to that path, then restart.
|
||||
Do not run original and restored databases as competing writers. Credentials
|
||||
and sessions are not in this database; users sign in again. Acknowledgments made
|
||||
after the selected backup are outside its recovery point: reconcile with
|
||||
independent audit evidence before opening the recovered service. A snapshot
|
||||
alone cannot reconstruct those later acknowledgments.
|
||||
|
||||
The restored store preserves pending, blocked and delivered states, original
|
||||
actors/decisions and event IDs. Already delivered events are not resent; an
|
||||
event accepted by audit-core whose reply was lost is replayed with its same ID
|
||||
and deduplicated by the receiver. Repeated human confirmation preserves the
|
||||
first acknowledgment. Tests exercise both cases, including real audit-core.
|
||||
|
||||
Off-host copying, retention, encryption, backup scheduling and a native restore
|
||||
drill remain package/platform responsibilities under the existing task.
|
||||
|
|
@ -159,9 +159,14 @@ email does not establish those remaining facts.
|
|||
Runtime validation (hash-pinned dependencies in requirements-runtime.lock):
|
||||
|
||||
```bash
|
||||
RTEL_FLEX_AUTH_BINARY=/tmp/rtel-flex-auth /tmp/rtel-ack-venv/bin/python -m unittest discover -s tests_runtime -v
|
||||
bash scripts/verify.sh
|
||||
```
|
||||
|
||||
The candidate workload, container recipe and runtime settings are owned by
|
||||
`rapp-telemetry/acknowledgment`. The policy source and exact client registration
|
||||
are in `integration/`; they are not active registrations.
|
||||
|
||||
See [maintenance and recovery](acknowledgment-operations.md) for aggregate
|
||||
metrics, explicit blocked-event retries, and backup/restore. The full verification
|
||||
command requires the documented sibling owner checkouts; it never silently skips
|
||||
the native contract tests.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue