railiance-infra/workplans/RAIL-HO-WP-0009-firewall-declared-state-and-api-exposure.md
codex 3d1bd75b6b
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
Record two further drift findings in T03
UFW is entirely inactive on CoulombCore - no firewall on a host running ArgoCD,
the registry and databases - while the declared baseline says UFW active with
default deny. Same defect class as the k3s finding but in the opposite
direction: the declaration is stronger than reality, and equally undetected.
Not an emergency (6443 unreachable from outside, 22/443/80 the expected
surface), but converging that host would enable UFW on a frozen production
system and needs its own decision.

Full convergence of Railiance01 carries 11 changes, most unrelated to the
firewall and none ever applied, including a user-slice memory cap that could OOM
running agent workloads. The base role has no tags, so convergence cannot be
scoped - adding tags folded into this task.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 02:00:17 +02:00

216 lines
8.8 KiB
Markdown

---
id: RAIL-HO-WP-0009
type: workplan
title: "Firewall declared-state integrity and k3s API exposure"
domain: financials
repo: railiance-infra
status: active
owner: codex
topic_slug: railiance
created: "2026-08-12"
updated: "2026-08-12"
related_repos:
- railiance-cluster
- railiance-platform
- railiance-master
state_hub_workstream_id: "ebf3c8d1-9066-4f13-a2d5-1da0409f2162"
---
# RAIL-HO-WP-0009 - Firewall declared-state integrity and k3s API exposure
## Goal
Close a security defect in which this repo's declared firewall configuration was
**weaker** than the live host, and remove the recurring exposure that produced
it.
## What was found
Discovered 2026-08-11 while diagnosing lost cluster access. The lost access was
incidental; the defect underneath was not.
**The defect.** `ansible/roles/base/tasks/main.yml` declared `6443/tcp` allowed
with no source restriction:
```yaml
- name: Allow k3s API in UFW
ansible.builtin.ufw:
rule: allow
port: '6443'
proto: tcp
```
The live host had **source-restricted** rules, added by hand with operator
comments (`k3s-api-operator-current`, `k3s-api-operator-hist`). Security was
tightened on the host and never fed back into the source of truth.
Consequence: **running this role would have removed the restriction and exposed
the Kubernetes API to the internet.** A convergence run intended to harden the
host would have de-hardened it, silently, with no failure to notice.
**Why it went undetected.** Nothing compares declared UFW state to live UFW
state. The divergence was security-relevant, had existed for some time, and
surfaced only because an ISP lease rotation (`89.244.90.246``.236`) happened
to break operator access mid-session.
**Contradiction with our own INTENT.** This repo declares "Declarative and
Reproducible — no irreproducible, hand-tuned hosts" and "Hardened by Default".
The hardening that actually protected the API was precisely the part that was
not declarative.
**Stale grants.** Two standing allowlist entries pointed at addresses no longer
ours: `89.244.90.246` (rotated) and `85.132.220.102` (historic). Dynamic ISP
addresses get reassigned, so an abandoned grant becomes a grant to a stranger.
## Boundaries
This workplan may:
- change firewall declarations and their inventory variables in this repo
- converge the base role against managed hosts, with operator approval
- propose the tunnel-based API access pattern and implement it here
It must not:
- change k3s or cluster runtime configuration (that is `railiance-cluster`, S2)
- silently widen any listening surface
- leave a host unreachable over SSH
## Tasks
```task
id: RAIL-HO-WP-0009-T01
status: done
priority: high
state_hub_task_id: "9a000a00-7cfd-4798-803f-bf59359a09b7"
```
Make the allowlist declarative. Add `k3s_api_allowed_sources` (empty default —
6443 closed to all external sources, the safe failure, SSH unaffected so hosts
stay recoverable) and `k3s_api_revoked_sources` so rotated addresses are pruned
rather than left standing. Grant approved sources **before** deleting the
blanket rule so convergence never opens a window without API access. Declare the
current operator address and the two stale grants in `group_vars/all.yaml`;
correct `docs/verification.md` to state that 6443 is source-restricted.
Committed 2026-08-11. Access restored on the live host by hand in the same
session (`89.244.90.236` granted; `kubectl` verified, node Ready v1.35.1+k3s1).
```task
id: RAIL-HO-WP-0009-T02
status: progress
priority: high
state_hub_task_id: "7d91dfc2-481b-4943-90e9-bb1b814a23dc"
```
Converge the base role against `railiance01` and confirm the resulting UFW state
matches the declaration exactly: the current operator source allowed, no blanket
`Anywhere` rule on 6443, and both revoked addresses gone. **Production action —
operator approval required before running.** Until this runs, the live host
still carries the two stale grants.
Verify SSH remains available throughout, and re-check `kubectl get nodes` after.
**Progress 2026-08-12.** The security goal is met surgically: both stale grants
(`89.244.90.246`, `85.132.220.102`) are deleted from the live host, and the live
6443 allowlist now matches the declaration exactly.
Full convergence was **deliberately not run**. A `--check` against `Railiance01`
reported **11 changes**, most unrelated to the firewall: an sshd restart,
`MemoryMax=1500M` and `MemorySwapMax=512M` on `user-1000.slice`, PAM `nproc`
caps for `tegwick`, swappiness, timezone, and the ops-bridge key injection. That
is a substantial and never-applied behavioural change to a production host —
notably the user-slice memory cap, which could OOM running agent workloads. It
is a separate decision from pruning two firewall grants, and the base role has
**no tags**, so convergence cannot currently be scoped to UFW alone.
Two follow-ons fall out of this:
- adding `tags:` to the base role, so firewall changes can be converged without
dragging unrelated drift with them
- deciding whether the resource-limit and sshd changes should be applied; they
are the declared baseline and have simply never been run
**A third operator address was found mid-session.** `89.244.90.255` appeared in
the live allowlist between two checks. It is legitimate — SSH pubkey auth as
`tegwick` from that address on 2026-08-02, and `[UFW BLOCK]` entries on 6443 at
23:41 on 2026-08-12 immediately before it was granted. It is now declared in
`group_vars/all.yaml`.
That is worth recording as evidence rather than as a footnote: the allowlist
drifted again, by hand, *during the very session that was fixing allowlist
drift*. It is the strongest available argument for T04.
```task
id: RAIL-HO-WP-0009-T03
status: todo
priority: high
state_hub_task_id: "3e835b96-736c-458c-91eb-04bfd4c7e0e7"
```
Audit the rest of the base role for the same class of defect: any place where
the declared state is weaker than, or simply divergent from, what was applied by
hand. UFW was found by accident; sshd config, fail2ban jails, sudoers, and
listening services deserve a deliberate pass. Record what is found even where no
change is needed — the absence of drift is itself evidence worth having.
Note during the same session: port `2224/tcp` is open to Anywhere on
`railiance01` (comment `nydus-ex-api dashboard agent`) and does not appear in
this role at all. Establish whether it is intended, and either declare it or
remove it.
**Second finding, 2026-08-12: UFW is entirely inactive on `CoulombCore`.**
`ufw status` returns `Status: inactive` — no firewall at all, on a host running
ArgoCD, the container registry and databases. The declared baseline
(`docs/verification.md`) says "UFW active, default deny inbound". Not an
emergency: 6443 is not reachable from outside (k3s appears bound locally or a
provider firewall is in front), and 22/443/80 are the expected surface. But it
is the same defect class as the k3s finding, in the opposite direction — the
declaration is *stronger* than reality here, and equally undetected.
Converging `CoulombCore` would **enable UFW on a frozen production host**, which
is a real availability risk and must not be done casually. Treat as its own
decision.
**Third finding: full convergence carries unrelated drift.** `--check` against
`Railiance01` reports 11 changes, most of them not firewall-related and none
ever applied — an sshd restart, `MemoryMax=1500M` / `MemorySwapMax=512M` on
`user-1000.slice`, PAM `nproc` caps, swappiness, timezone. The user-slice memory
cap in particular could OOM running agent workloads. The base role has **no
tags**, so convergence cannot be scoped. Add tags as part of this task.
```task
id: RAIL-HO-WP-0009-T04
status: todo
priority: medium
state_hub_task_id: "908630e8-e245-47f5-a060-3949250f522c"
```
Remove the k3s API from the public internet. Operator addresses rotate, so an
allowlist is a treadmill: it will drift again, and each drift is either an
outage or a stale grant. `docs/deploy-stack.md` already documents API access
over the ops-bridge SSH tunnel (`k3s-api-coulombcore`, local port 16443) — apply
the same pattern to `railiance01` and reduce the public allowlist to nothing.
Decide explicitly rather than by default: this trades convenience for exposure,
and the tunnel becomes a dependency of every operator action.
```task
id: RAIL-HO-WP-0009-T05
status: todo
priority: medium
state_hub_task_id: "81af8470-286c-440a-bbe7-b47a17ee6750"
```
Propose a declared-vs-live conformance check for firewall state, and route it to
whoever owns the conformance loop. This defect is a concrete instance of the
unowned **Q7 Governance and Change Management** gap recorded in
`railiance-platform/ArchitectureBlueprint.md` §5.3 — a UFW declared-vs-live
comparison is a small, sharp first thing for that loop to do.
The existing Goss verification suite is the natural home for the check itself;
what is missing is the loop that runs it and reacts.
## Outcome
Pending. T01 done; the live host is reachable but not yet converged.