Prove installed Railiance runtime recovery under worker failures
Some checks failed
Governed runtime contract / contract (push) Failing after 18s
Some checks failed
Governed runtime contract / contract (push) Failing after 18s
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e387-534d-70e3-ad53-4ea05676db8c
This commit is contained in:
parent
73aaa4bcd4
commit
353554eb49
4 changed files with 562 additions and 0 deletions
|
|
@ -170,6 +170,43 @@ every permanent Activity Core close code (including `expired_lease`), and Glas
|
|||
cancellation. It does not replace the isolated expired-row API smoke or the
|
||||
natural-run and sandbox-cleanup observations required from the deployed host.
|
||||
|
||||
### Installed-runtime recovery drill
|
||||
|
||||
Run `scripts/prove-runtime-recovery.py` with the selected artifact's interpreter:
|
||||
|
||||
```bash
|
||||
"${RUNTIME}/bin/python3" -I -B scripts/prove-runtime-recovery.py \
|
||||
--runtime "${RUNTIME}" --sha256 "${RUNTIME_SHA256}" \
|
||||
--profile-ref harness.agent-dev-local@1.1.1
|
||||
```
|
||||
|
||||
The script verifies the complete artifact digest and that all four runtime
|
||||
packages load from that artifact. It uses private temporary repositories,
|
||||
sandbox stores and worker state. The queue and rein authoring are fixtures;
|
||||
the installed worker, Glas lifecycle, bwrap execution, repository acceptance,
|
||||
external metrics and close outbox are real. No provider or production queue
|
||||
client is invoked. The selected profile is made ready only in an in-memory
|
||||
fixture catalog; no installed profile or standing service is changed.
|
||||
|
||||
The receipt covers response-lost close replay with one execution/commit,
|
||||
execution failure, periodic lease rejection, a real SIGTERM to the disposable
|
||||
worker, a killed lock holder, and SIGKILL of a disposable worker after sandbox
|
||||
creation. Cancellation deliberately returns a late sandbox commit to prove
|
||||
that it is not imported into the source. Every case checks lock reacquisition,
|
||||
source state, sandbox/workspace cleanup and applicable durable evidence.
|
||||
|
||||
After abrupt worker death, cleanup is **explicit owner recovery**, not an
|
||||
automatic startup sweep. Recover only the sandbox id attributed to the failed
|
||||
run, using sand-boxer's `get`/`destroy` with the same owner state directory.
|
||||
Verify destruction and workspace removal before resuming. Do not delete lock
|
||||
files to release a kernel lock, erase pending close evidence, or replay the
|
||||
workload to reconcile a close. Use `close-outbox status|replay` for that last
|
||||
step. Activity Core remains responsible for expiring the abandoned queue lease.
|
||||
|
||||
Railiance receipt: `docs/evidence/2026-09-27-installed-runtime-recovery.json`.
|
||||
It proves these host mechanisms against runtime `b6e4e8a4`; it does not prove
|
||||
live Activity Core expiry, credential delivery, or a paid production run.
|
||||
|
||||
### Host access to cluster services (no port-forward)
|
||||
|
||||
On railiance01 (single-node k3s), set **k8s://** pseudo-URLs in
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue