--- id: RAIL-HO-WP-0009 type: workplan title: "Firewall declared-state integrity and k3s API exposure" domain: financials repo: railiance-infra status: active owner: codex topic_slug: railiance created: "2026-08-12" updated: "2026-08-12" related_repos: - railiance-cluster - railiance-platform - railiance-master state_hub_workstream_id: "ebf3c8d1-9066-4f13-a2d5-1da0409f2162" --- # RAIL-HO-WP-0009 - Firewall declared-state integrity and k3s API exposure ## Goal Close a security defect in which this repo's declared firewall configuration was **weaker** than the live host, and remove the recurring exposure that produced it. ## What was found Discovered 2026-08-11 while diagnosing lost cluster access. The lost access was incidental; the defect underneath was not. **The defect.** `ansible/roles/base/tasks/main.yml` declared `6443/tcp` allowed with no source restriction: ```yaml - name: Allow k3s API in UFW ansible.builtin.ufw: rule: allow port: '6443' proto: tcp ``` The live host had **source-restricted** rules, added by hand with operator comments (`k3s-api-operator-current`, `k3s-api-operator-hist`). Security was tightened on the host and never fed back into the source of truth. Consequence: **running this role would have removed the restriction and exposed the Kubernetes API to the internet.** A convergence run intended to harden the host would have de-hardened it, silently, with no failure to notice. **Why it went undetected.** Nothing compares declared UFW state to live UFW state. The divergence was security-relevant, had existed for some time, and surfaced only because an ISP lease rotation (`89.244.90.246` → `.236`) happened to break operator access mid-session. **Contradiction with our own INTENT.** This repo declares "Declarative and Reproducible — no irreproducible, hand-tuned hosts" and "Hardened by Default". The hardening that actually protected the API was precisely the part that was not declarative. **Stale grants.** Two standing allowlist entries pointed at addresses no longer ours: `89.244.90.246` (rotated) and `85.132.220.102` (historic). Dynamic ISP addresses get reassigned, so an abandoned grant becomes a grant to a stranger. ## Boundaries This workplan may: - change firewall declarations and their inventory variables in this repo - converge the base role against managed hosts, with operator approval - propose the tunnel-based API access pattern and implement it here It must not: - change k3s or cluster runtime configuration (that is `railiance-cluster`, S2) - silently widen any listening surface - leave a host unreachable over SSH ## Tasks ```task id: RAIL-HO-WP-0009-T01 status: done priority: high state_hub_task_id: "9a000a00-7cfd-4798-803f-bf59359a09b7" ``` Make the allowlist declarative. Add `k3s_api_allowed_sources` (empty default — 6443 closed to all external sources, the safe failure, SSH unaffected so hosts stay recoverable) and `k3s_api_revoked_sources` so rotated addresses are pruned rather than left standing. Grant approved sources **before** deleting the blanket rule so convergence never opens a window without API access. Declare the current operator address and the two stale grants in `group_vars/all.yaml`; correct `docs/verification.md` to state that 6443 is source-restricted. Committed 2026-08-11. Access restored on the live host by hand in the same session (`89.244.90.236` granted; `kubectl` verified, node Ready v1.35.1+k3s1). ```task id: RAIL-HO-WP-0009-T02 status: progress priority: high state_hub_task_id: "7d91dfc2-481b-4943-90e9-bb1b814a23dc" ``` Converge the base role against `railiance01` and confirm the resulting UFW state matches the declaration exactly: the current operator source allowed, no blanket `Anywhere` rule on 6443, and both revoked addresses gone. **Production action — operator approval required before running.** Until this runs, the live host still carries the two stale grants. Verify SSH remains available throughout, and re-check `kubectl get nodes` after. **Progress 2026-08-12.** The security goal is met surgically: both stale grants (`89.244.90.246`, `85.132.220.102`) are deleted from the live host, and the live 6443 allowlist now matches the declaration exactly. Full convergence was **deliberately not run**. A `--check` against `Railiance01` reported **11 changes**, most unrelated to the firewall: an sshd restart, `MemoryMax=1500M` and `MemorySwapMax=512M` on `user-1000.slice`, PAM `nproc` caps for `tegwick`, swappiness, timezone, and the ops-bridge key injection. That is a substantial and never-applied behavioural change to a production host — notably the user-slice memory cap, which could OOM running agent workloads. It is a separate decision from pruning two firewall grants, and the base role has **no tags**, so convergence cannot currently be scoped to UFW alone. Two follow-ons fall out of this: - adding `tags:` to the base role, so firewall changes can be converged without dragging unrelated drift with them - deciding whether the resource-limit and sshd changes should be applied; they are the declared baseline and have simply never been run **A third operator address was found mid-session.** `89.244.90.255` appeared in the live allowlist between two checks. It is legitimate — SSH pubkey auth as `tegwick` from that address on 2026-08-02, and `[UFW BLOCK]` entries on 6443 at 23:41 on 2026-08-12 immediately before it was granted. It is now declared in `group_vars/all.yaml`. That is worth recording as evidence rather than as a footnote: the allowlist drifted again, by hand, *during the very session that was fixing allowlist drift*. It is the strongest available argument for T04. ```task id: RAIL-HO-WP-0009-T03 status: todo priority: high state_hub_task_id: "3e835b96-736c-458c-91eb-04bfd4c7e0e7" ``` Audit the rest of the base role for the same class of defect: any place where the declared state is weaker than, or simply divergent from, what was applied by hand. UFW was found by accident; sshd config, fail2ban jails, sudoers, and listening services deserve a deliberate pass. Record what is found even where no change is needed — the absence of drift is itself evidence worth having. Note during the same session: port `2224/tcp` is open to Anywhere on `railiance01` (comment `nydus-ex-api dashboard agent`) and does not appear in this role at all. Establish whether it is intended, and either declare it or remove it. **Second finding, 2026-08-12: UFW is entirely inactive on `CoulombCore`.** `ufw status` returns `Status: inactive` — no firewall at all, on a host running ArgoCD, the container registry and databases. The declared baseline (`docs/verification.md`) says "UFW active, default deny inbound". Not an emergency: 6443 is not reachable from outside (k3s appears bound locally or a provider firewall is in front), and 22/443/80 are the expected surface. But it is the same defect class as the k3s finding, in the opposite direction — the declaration is *stronger* than reality here, and equally undetected. Converging `CoulombCore` would **enable UFW on a frozen production host**, which is a real availability risk and must not be done casually. Treat as its own decision. **Third finding: full convergence carries unrelated drift.** `--check` against `Railiance01` reports 11 changes, most of them not firewall-related and none ever applied — an sshd restart, `MemoryMax=1500M` / `MemorySwapMax=512M` on `user-1000.slice`, PAM `nproc` caps, swappiness, timezone. The user-slice memory cap in particular could OOM running agent workloads. The base role has **no tags**, so convergence cannot be scoped. Add tags as part of this task. ```task id: RAIL-HO-WP-0009-T04 status: todo priority: medium state_hub_task_id: "908630e8-e245-47f5-a060-3949250f522c" ``` Remove the k3s API from the public internet. Operator addresses rotate, so an allowlist is a treadmill: it will drift again, and each drift is either an outage or a stale grant. `docs/deploy-stack.md` already documents API access over the ops-bridge SSH tunnel (`k3s-api-coulombcore`, local port 16443) — apply the same pattern to `railiance01` and reduce the public allowlist to nothing. Decide explicitly rather than by default: this trades convenience for exposure, and the tunnel becomes a dependency of every operator action. ```task id: RAIL-HO-WP-0009-T05 status: todo priority: medium state_hub_task_id: "81af8470-286c-440a-bbe7-b47a17ee6750" ``` Propose a declared-vs-live conformance check for firewall state, and route it to whoever owns the conformance loop. This defect is a concrete instance of the unowned **Q7 Governance and Change Management** gap recorded in `railiance-platform/ArchitectureBlueprint.md` §5.3 — a UFW declared-vs-live comparison is a small, sharp first thing for that loop to do. The existing Goss verification suite is the natural home for the check itself; what is missing is the loop that runs it and reacts. ## Outcome Pending. T01 done; the live host is reachable but not yet converged.