Add a tool-neutral environment orientation for all coding agents.
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s

Collects the verified environment facts and traps from 2026-09-21:
workstation vs railiance01 kube context, tunnels, harness permissions,
change gate and rollout headroom, attended OpenBao login, secret-read
hazards, ArgoCD layout, and vocabulary. Linked from AGENTS.md and the
session protocol.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
codex 2026-09-21 21:33:04 +02:00
parent f2ca7fb883
commit c5f7f2df23
3 changed files with 121 additions and 0 deletions

View file

@ -5,6 +5,9 @@ MCP server name in `~/.claude.json`: `dev-hub`
**Step 1 — Orient**
Before any production, credential or GitOps work, read
`docs/agent-environment-orientation.md` (environment facts and traps, all agents).
Read the offline-safe brief first — it works without a live hub connection:
```bash
cat .custodian-brief.md

View file

@ -116,6 +116,19 @@ curl -s -X PATCH "http://127.0.0.1:8000/tasks/<task_id>" \
---
## Environment orientation (read before touching railiance01, OpenBao or ArgoCD)
`docs/agent-environment-orientation.md` holds the verified facts about this estate's environment:
- where things run, and that the workstation's kube context is NOT railiance01;
- the tunnel addresses;
- how to make live changes safely, including CPU headroom and rollout deadlock;
- the attended OpenBao admin login and its traps;
- secret-reading hazards;
- the ArgoCD layout;
- vocabulary.
It is tool-neutral. Read it before any production, credential or GitOps work.
## Credential and access routing
**Audience:** Codex, Claude Code, Grok, and custodian agents that call **llm-connect**

View file

@ -0,0 +1,105 @@
# Agent environment orientation
**Audience:** every coding agent working in this estate (Claude Code, Codex, Grok, custodian workers). It is tool-neutral.
**Owner:** the-custodian. **Last verified:** 2026-09-21.
These are facts about the *environment*: where things run, how to reach them, and the traps that cost real time. Each section names its owner. When a fact changes, fix it here and in the owner's record.
Read this before touching railiance01, OpenBao, ArgoCD or credentials.
---
## 1. Where you are
- **The workstation is WSL2** (Ubuntu on Windows). It is dev-only; production runs on railiance01. There is no `xdg-open` by default: see §5.
- **railiance01** (`92.205.62.239`) is production. It is a **single-node k3s cluster**, the only control-plane node. `ssh railiance01` reaches it as user `tegwick`.
- **The workstation's kube context is NOT railiance01.** A `kubectl` run on the workstation talks to another cluster. On 2026-09-21 that sent an agent to the wrong conclusion about ArgoCD.
- To read or change railiance01, run commands on the node: `ssh railiance01 'kubectl …'`.
- Before trusting any result, confirm the node IP is `92.205.62.239` (`kubectl get nodes -o wide`).
- On railiance01, `kubectl` works directly. **`helm` needs `export KUBECONFIG=/etc/rancher/k3s/k3s.yaml`**, or it tries `localhost:8080`.
- **"railiance01" in older records sometimes means coulombcore.** That usage predates the 2026-07-02 correction. coulombcore is the frozen old cluster; it still runs a stale ArgoCD.
## 2. Reaching services from the workstation
Services are private by default (railiance-master ADR-0008) and are reached through `ops-bridge` SSH tunnels. `bridge status` shows them; `bridge up <name>` restores one.
| Service | Address on the workstation | Tunnel |
|---|---|---|
| State Hub API | `http://127.0.0.1:8000` (health: `/state/health`) | `state-hub-primary` |
| OpenBao | `http://127.0.0.1:18200` | `openbao-ui-railiance01` |
| k3s API | via `k3s-api-railiance01` | `k3s-api-railiance01` |
**Trap:** the shell default `BAO_ADDR=https://bao.coulomb.social` is **unreachable** from the workstation. Use `BAO_ADDR=http://127.0.0.1:18200` (and `VAULT_ADDR` the same) for any `bao` command.
## 3. Permissions and the agent harness (Claude Code)
- The project allows `Bash(ssh railiance01 *)`. A command that **starts** with `ssh railiance01 '…'` matches that rule. A command that starts with anything else (`timeout …`, `cd … && for …`) goes to the auto-mode classifier. The classifier blocks cluster mutations as "Shared Cluster Mutation".
- **A block is a stop signal, not a puzzle to route around.** Get the founder's go-ahead for the live change first, then run it in the `ssh railiance01 '…'` form. `scp <file> railiance01:/tmp/…` followed by `ssh railiance01 'kubectl apply -f /tmp/…'` is the working pattern for applying a file.
- The harness blocks foreground `sleep`. To wait on the cluster, loop inside the ssh command, e.g. `for i in $(seq 1 20); do <check> && break; sleep 3; done`.
## 4. Changing production (railiance01)
- **Every live change needs the founder's go-ahead.** Record it in Mode of Authority terms (§8): `ADMINISTER @ realm:kubernetes/railiance01`, `activation=APPROVED`.
- **The change gate is tiered by ADR-0006 `readiness_state`**, not by environment, because it is a single cluster. Record: `the-custodian/docs/kubernetes-change-gate-decision.md`.
- Below `production-approved`: direct apply under founder plan approval.
- At `production-approved`: `CONSTRUCT` through git and ArgoCD.
- Emergency: `activation=BREAK_GLASS`, recorded, then reconciled into git.
- Unmapped targets count as production-tier. The `whitehat` namespace is explicitly non-production.
- **Check CPU headroom before any rollout.** The node runs near its limit (~85% of CPU requested).
- Command: `kubectl describe node | grep -A4 "Allocated resources"`.
- A rolling update creates the new pod **before** stopping the old one. On a node with no requestable CPU, even a change that *lowers* requests deadlocks. This happened on 2026-09-21; freeing 50m elsewhere unblocked it.
- Free capacity first, then change requests.
- **Always dry-run and diff before applying:** `kubectl diff` (client-side) and `--dry-run=server`.
- A *server-side* diff can report a field-ownership **conflict** where values are semantically equal, e.g. `1` vs `1000m` CPU owned by an older client-side apply.
- ArgoCD's apply strategy uses client-side apply, so check the client-side diff too.
- **Helm releases:** deploy through the owner's documented path, e.g. state-hub's `PROMOTE.md`, with its headroom preflight and `--atomic`. Pin the image tag the release already runs unless the change is the image.
- A release stuck in `pending-upgrade`: compare the revisions' values and manifests with live first. If identical, `helm rollback <rel> <last-deployed>` clears it without touching objects.
- Declare what you change in the owner's repository in the same session. Otherwise the next upstream re-apply silently reverts it. On 2026-09-21 the Knative installer would have restored the CPU requests that had been starving the backups.
## 5. OpenBao admin actions (attended login)
Owner: railiance-platform (OpenBao) and ops-warden (the `warden access` lane).
- Admin actions go through `warden access openbao-platform-admin-login --exec -- <reviewed-command>`.
- **The founder runs it in their own terminal** (`! …` in Claude Code). KeyCape OIDC/MFA needs a TTY and a browser; it fails from an agent's non-TTY shell.
- Prefix `BAO_ADDR=http://127.0.0.1:18200 VAULT_ADDR=http://127.0.0.1:18200` (§2).
- **Browser on WSL:** `~/.local/bin/xdg-open` is a one-line shim to `/mnt/c/Windows/explorer.exe "$1"`, so `bao login -method=oidc` can open KeyCape. Without an opener, the login URL goes into warden's suppressed output and the login times out.
- **The reviewed command must be completely silent.** warden fails closed on ANY child output, and `bao … write` prints "Success!". Start the child with `exec >/dev/null 2>&1`. Make it idempotent: compare before writing, never overwrite a differing object, and use distinct exit codes for "differs", "write failed" and "verify failed". A working example: `railiance-platform` RPF-WP-0045 execution record.
- **Do not trust warden's exit code; read the printed line.** As of 2026-09-21:
- A failed login can exit 0, and the audit records every attempt as `ok`.
- `attended command completed but session revocation could not be confirmed` (exit 1) means **the child succeeded**; only revocation is unconfirmed.
- `attended login failed closed before command handoff` means the child **never ran**.
- Reported to ops-warden (hub message `6a1ce1bb`).
## 6. Secrets and credentials
- Run the credential-routing check (`warden route find "<need>"`) before requesting anything. See the "Credential and access routing" section in every repo's `AGENTS.md`.
- **Never read a Secret with `-o yaml`, `-o json`, `describe`-style tools, or even `-o jsonpath='{.metadata}'`.** A Secret created with `kubectl apply` carries its full data in the `kubectl.kubernetes.io/last-applied-configuration` annotation, so a metadata read prints the secret. To test for the annotation without printing a value:
`kubectl get secret <n> -o go-template='{{ if index .metadata.annotations "kubectl.kubernetes.io/last-applied-configuration" }}HAS-ANNOTATION{{ else }}clean{{ end }}'`
- ExternalSecrets use ClusterSecretStores. The target pattern is **OpenBao Kubernetes auth**: one ServiceAccount per consumer, 15-minute tokens, and a policy scoped to exact paths. Five stores still use static `*-eso-token` Secrets, which expire. Do not re-run the old `*-eso-token-apply` scripts.
## 7. GitOps (ArgoCD) on railiance01
Owner: railiance-platform (Applications) and railiance-enablement (the ArgoCD install, S4).
- ArgoCD Core v3.5.3 (headless) runs in `argocd`, **installed 2026-09-21**. Adoption is one app at a time, per RPF-WP-0044:
- openbao-secretstore and target-revenue are adopted, with manual sync and automated sync **off**.
- issue-core and external-secrets are not yet adopted.
- **Use `argocd/railiance01/bootstrap/`. NEVER run `make argocd-bootstrap-deploy` against railiance01.** It renders the old root, which has automated sync, prune and self-heal, and would adopt everything at once.
- **`argocd/applications/` on `main` is still read by coulombcore's ArgoCD.** Editing it is a live change on coulombcore.
- Adopted apps: a hand deploy (`make deploy`, `kubectl apply`) now fights ArgoCD. Change git, then sync.
## 8. Vocabulary
- **SecurityCanon Mode of Authority** (`security-canon/infospace/vocabulary/mode-of-authority/`):
- Write authority as `MODE @ SCOPE`. Modes: USE, OBSERVE, OPERATE, ADMINISTER, GOVERN, CONSTRUCT, SOVEREIGN. They are not a privilege ladder.
- `activation`: APPROVED, BREAK_GLASS, … `BREAK_GLASS` is an activation, not a mode.
- `EvidenceBoundary`: `target-audited`, `external-audited`.
- **Never write bare "operator".** `Operator` is a CARING lifecycle role and `OPERATE` is an Auth Mode. Bernd's decisions are "the founder, exercising `GOVERN @ estate`".
- **Security layer model (net-kingdom §11):** a repository declares its own layer in `INTENT.md` frontmatter, which governs. `layer.yaml` is derived. There is no version in the declaration, and case is folded. Rulings: gate-house GH-DEC-2026-017, 020 and 021.
## 9. Hub and workplan hygiene
- The State Hub is a read model. Write workplan files in the repo, then run `statehub fix-consistency --repo <repo>`; its sync step pushes. Never register workplans or tasks by hand.
- Do not run `git push` yourself unless asked. The fix-consistency sync push is fine. Leave other people's uncommitted files alone, and stage by explicit path.
- The hub's needs-human prompt text (`intervention_note`) is not synced by fix-consistency. If you change it in a workplan file, the dashboard keeps showing the old text.