Clarify cluster boundary for rail-kubernetes
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 6s

This commit is contained in:
codex 2026-07-25 11:17:51 +02:00
parent 74b490a31e
commit be3b8ff31e
9 changed files with 329 additions and 146 deletions

View file

@ -1,77 +1,30 @@
# Railiance Documentation
# railiance-cluster Docs
Welcome to **Railiance** — an opinionated framework for Infrastructure-as-Code.
Think of it as *Rails for Ops*: convention over configuration, productivity first, reproducible everywhere.
This directory mixes two kinds of material:
---
- stable S2 cluster-substrate docs that belong in `railiance-cluster`
- transitional workload-contract docs that are temporarily hosted here until
`rail-kubernetes` exists
## 🚀 What is Railiance?
Use the sections below to keep those concerns separate.
Railiance is the **self-reliant infrastructure bootstrap** for Coulomb and beyond.
From two bare Linux servers, a Git repo, and valid credentials, you can rebuild a full Kubernetes-based environment with:
## S2 Substrate Docs
- **Automated provisioning** (Ansible, Helm, k3s)
- **Git-driven state** (ArgoCD / Flux)
- **AI-ready automation** with OODA loop (Observe, Orient, Decide, Act)
- **End-to-end reproducibility** — no hidden steps, no snowflake servers
- [Operator runbook](operator-runbook.md)
- [Backup and restore](backup-restore.md)
- [Kubeconfig](kubeconfig.md)
- [First host](first_host.md)
---
## Published Boundary Docs
## 🌱 Core Principles
- [rail-kubernetes extraction map](rail-kubernetes-extraction-map.md)
- [rail-kubernetes substrate profile](rail-kubernetes-substrate-profile.md)
1. **Convention over configuration**
Opinionated defaults remove friction.
## Transitional Workload-Contract Docs
2. **Dont repeat yourself (DRY)**
Single source of truth in Git, reused everywhere.
3. **Automated OODA**
- *Observe*: collect telemetry.
- *Orient*: AI interprets state.
- *Decide*: human reviews & authorizes.
- *Act*: automation applies changes.
4. **Self-Reliance**
Everything can be rebuilt from scratch, with no external dependencies beyond Git + credentials.
5. **Calm Ops**
Systems should be simple, resilient, and predictable.
---
## 📂 Repo Layout
- `ansible/` — bootstrap playbooks (OS setup, k3s install, etc.)
- `helm/` — Helm charts for services
- `k8s/` — raw manifests & overlays
- `tests/` — automated validation of infra steps
- `docs/` — documentation & contributor guides
- `tools/` — helper scripts (e.g. repo bootstrap)
---
## 🛠 Getting Started
1. **Clone the repo**
```bash
git clone <your-gitea-url>/railiance-cluster.git
cd railiance-cluster
2. **Prepare your config**
Edit ~/.railiance_gitea.conf with your Gitea URL, user, and token.
(Never commit this file!)
3. **Run bootstrap**
Use the playbooks in ansible/ to turn a bare Ubuntu server into a Kubernetes host.
4. **Deploy services**
Install baseline services via Helm from the helm/ directory.
## Incidents
- [2026-03-10 — pgpool CrashLoopBackOff on PostgreSQL HA failover](incidents/2026-03-10-pgpool-missing-secret.md)
## Operations
These docs describe the generic workload-on-Kubernetes contract. They are still
hosted in this repo for compatibility, but they are targeted for extraction to
future `rail-kubernetes`.
- [Deployment lifecycle](deployment-lifecycle.md)
- [Railiance app.toml contract](app-toml-contract.md)
@ -81,13 +34,6 @@ From two bare Linux servers, a Git repo, and valid credentials, you can rebuild
- [Promote, rollback, and onboarding](promote-rollback-onboarding.md)
- [Railiance run command](railiance-run-command.md)
## 👥 Contributing
## Incidents
See CONTRIBUTING.md for rules, coding style, and workflow.
## 🌍 Vision
Railiance is more than infra scripts: its the foundation of self-empowering infrastructure,
where humans and AI agents collaborate to manage systems with trust, clarity, and calmness.
From bare metal to resilient clusters, in one repo.
- [2026-03-10 - pgpool CrashLoopBackOff on PostgreSQL HA failover](incidents/2026-03-10-pgpool-missing-secret.md)

View file

@ -10,6 +10,11 @@ dependencies, and secret references are described by stable names. Plaintext
secrets, bearer tokens, kubeconfigs, and private key material must never appear
in `railiance/app.toml`.
Ownership note: this contract is temporarily hosted in `railiance-cluster`
while the wave-1 split prepares a future `rail-kubernetes` repo. Long-term
ownership belongs with the generic workload-on-Kubernetes contract, not with S2
cluster substrate ownership.
The machine-readable schema lives at `schemas/railiance-app.schema.json`. A
minimal example lives at `examples/railiance/app.toml`.

View file

@ -9,16 +9,24 @@ stage must leave evidence, every promotion must have a rollback path, and
critical workloads require explicit human approval before production traffic is
changed.
Ownership note: this specification is temporarily hosted in
`railiance-cluster` during the wave-1 split. Its long-term home is future
`rail-kubernetes`; `railiance-cluster` should retain only substrate
prerequisites and compatibility shims once that repo exists.
## Scope
This specification is owned by `railiance-cluster` because it defines the
cluster runtime contract for promotion gates, canary validation, production
routing, and rollback expectations.
This specification currently lives beside the cluster substrate because the
compatibility tooling still lives here. Long-term ownership is not S2.
Repo boundaries:
- `railiance-cluster` owns the lifecycle semantics, cluster prerequisites,
ingress/routing expectations, and acceptance gates.
- future `rail-kubernetes` owns lifecycle semantics, generic promotion gates,
canary validation, production routing expectations, rollback expectations,
and the generic workload contract.
- `railiance-cluster` owns only the substrate prerequisites published to that
rail: cluster connectivity, namespace behavior, ingress availability, and
operator/runtime expectations.
- `railiance-apps` owns workload-specific Helm values, application release
definitions, and production workload configuration.
- `railiance-platform` owns shared platform services such as databases,
@ -329,4 +337,3 @@ bin/railiance rollback <overlay-dir> --plan # rollback to previous s
The exact command names may change as implementation lands, but the behavior
must preserve the stage gates and evidence requirements in this document.

View file

@ -0,0 +1,85 @@
# rail-kubernetes Extraction Map
This document records the reviewed wave-1 extraction map from
`railiance-cluster` to a future `rail-kubernetes` repo.
The goal is to separate:
- S2 Kubernetes substrate ownership that stays in `railiance-cluster`
- generic workload-on-Kubernetes contract material that moves to
`rail-kubernetes`
- unrelated or workload-specific helpers that must be rehomed elsewhere instead
of being normalized into either repo
## Move To Future `rail-kubernetes`
| Path | Why it moves |
| --- | --- |
| `docs/deployment-lifecycle.md` | Generic Stage 1/2/3 workload lifecycle semantics, not S2 substrate ownership |
| `docs/app-toml-contract.md` | Generic workload declaration contract |
| `docs/overlay-repo-pattern.md` | Generic wrapper/overlay packaging pattern |
| `docs/canary-helm-template.md` | Generic canary chart pattern for workloads |
| `docs/stage2-deploy-observe.md` | Generic Stage 2 command behavior |
| `docs/promote-rollback-onboarding.md` | Generic workload onboarding to the lifecycle |
| `docs/railiance-run-command.md` | Generic Stage 1 command contract |
| `schemas/railiance-app.schema.json` | Schema for the generic workload contract |
| `examples/railiance/app.toml` | Example for the generic workload contract |
| `tools/create_railiance_overlay_repo.sh` | Generic wrapper/overlay scaffolding |
| `tools/cmd/railiance-run` | Generic Stage 1 executor |
| `tools/cmd/railiance-stage2` | Generic Stage 2 canary deploy/observe executor |
| `tools/cmd/railiance-stage3` | Generic Stage 3 promote/rollback executor |
| `bin/railiance` subcommands `create-overlay`, `run`, `deploy`, `observe`, `promote`, `rollback` | Dispatcher surface for the generic workload rail |
| `tools/README_tools.md` sections for `railiance-run`, `railiance-stage2`, `railiance-stage3`, and `create_railiance_overlay_repo.sh` | Tool documentation should move with the tools it describes |
## Retain In `railiance-cluster`
| Path | Why it stays |
| --- | --- |
| `ansible/bootstrap.yml` | k3s and Helm bootstrap are core S2 substrate ownership |
| `docs/backup-restore.md` | Cluster-state backup and restore posture |
| `docs/kubeconfig.md` | Cluster access delivery and handling |
| `docs/operator-runbook.md` | Operator-facing cluster-touching commands, though some rows need separate cleanup |
| `docs/first_host.md` | Host-to-cluster bootstrap guidance tied to S2 setup |
| `tools/cmd/railiance-backup-s2` | S2 cluster-state backup helper |
| `tools/cmd/railiance-restore-s2` | S2 restore guidance helper |
| `tools/cmd/railiance-preflight` | Cluster safety gate before invasive work |
| `tests/smoke_kube.sh` | Cluster health smoke test |
| `tests/test_ha_failover.sh` | Cluster/platform integration smoke for HA failover expectations |
| `Makefile` targets `k3s-install`, `smoke`, `backup`, `restore`, `preflight`, `test-ha-failover` | Cluster substrate operations and safety rails |
| `bin/railiance` subcommands `backup` and `preflight` | Compatibility wrapper for retained S2 operations until a cleaner CLI split exists |
## Rehome Outside Both Repos
These files are real migration debt, but they should not be normalized into
future `rail-kubernetes` or kept as permanent S2 identity.
| Path | Why it needs another home |
| --- | --- |
| `tools/cmd/railiance-verify-activity-core` | Workload-specific reconcile/verify flow for one app |
| `tools/cmd/railiance-reconcile-activity-core-llm-connect` | Cross-repo workload-specific helper, not generic substrate or generic rail |
| `tools/cmd/railiance-deploy-activity-core-triage-robustness` | Workload-specific deployment hardening helper |
| `tools/cmd/railiance-admin-sync-smoke` | Workload-specific app smoke command |
| `Makefile` targets `verify-activity-core`, `reconcile-activity-core-llm-connect`, `deploy-activity-core-triage-robustness`, `admin-sync-smoke` | Operator entry points for workload-specific helpers above |
| `bin/railiance` subcommands `deploy-triage-robustness` and `admin-sync-smoke` | Dispatcher surface for workload-specific helpers |
## Shared Tooling Debt Outside This Wave
The repo also contains bootstrap and shared-tooling commands that belong
neither to S2 substrate ownership nor to the future `rail-kubernetes` rail.
They should get a separate placement review rather than piggybacking on the
current extraction:
- `bin/railiance` subcommands `doctor`, `next`, `plan-host`, `gen-ssh-key`,
`cloudinit`, `init-repo`, `build-spore`, `seed-local`, `checklist`
- `tools/cmd/railiance-doctor`
- `tools/cmd/railiance-plan-host`
- `tools/build_spore.sh`
- `tools/seed_node.sh`
- `tools/create_railiance_repo.sh`
- `tools/furnish_railiance_repo.sh`
## Resulting Rule
Until `rail-kubernetes` exists, `railiance-cluster` may host compatibility
copies of the generic lifecycle and command surface. It must not keep claiming
that those materials are the steady-state S2 ownership boundary.

View file

@ -0,0 +1,113 @@
# rail-kubernetes Substrate Profile
This document records the minimum cluster substrate facts that
`railiance-cluster` publishes for a future `rail-kubernetes` repo.
The intent is narrow: future rail code should depend on explicit substrate
capabilities, not on incidental implementation details scattered across this
repo.
## Runtime Base
- Kubernetes runtime: k3s installed by `ansible/bootstrap.yml`
- Current pinned k3s version in repo: `v1.35.1+k3s1`
- Current pinned Helm version in repo: `v3.17.3`
- Baseline health checks: `make smoke`
- Cluster safety gate before invasive work: `make preflight`
`rail-kubernetes` may assume a k3s-backed Kubernetes API plus Helm-based apply
paths on the current Railiance substrate. It should not assume more than that
without a declared extension.
## Access Model
- Operator access is delivered from `/etc/rancher/k3s/k3s.yaml` to
`~/.kube/config-hosteurope` as documented in `docs/kubeconfig.md`
- The raw k3s kubeconfig uses `127.0.0.1` as the server address and may require
rewrite or tunneling for remote use
- The documented operator model is cluster-admin kubeconfig access, not a
least-privilege workload service-account contract
Future `rail-kubernetes` should assume operator-driven cluster access for now.
It should not assume a published per-workload RBAC template yet.
## Routing And Canary Exposure
- The current generic tooling supports isolated canary deployment by creating a
separate Stage 2 release and namespace
- `docs/canary-helm-template.md` documents a Traefik weighted-routing example,
but `railiance-cluster` does not currently publish weighted routing as a
guaranteed substrate capability
- `tools/cmd/railiance-stage2` can inspect rollout, pods, and ingress/routing
resources for the canary release
Published guarantee for wave 1: isolated canary exposure is the safe default.
Weighted, header-based, path-based, or shadow routing must be treated as
optional future capabilities until S2 documents them explicitly.
## TLS And Cluster Operators
`SCOPE.md` currently records these cluster-level addons/operators on the active
substrate:
- nginx ingress
- cert-manager
- ArgoCD
- CloudNative PG operator
- SSO stack namespaces (`mfa`, `sso`) via Net Kingdom
Future `rail-kubernetes` should depend on capabilities such as ingress
availability or TLS readiness, not on owning these operators directly.
## Observability Expectations
- `tools/cmd/railiance-stage2` live observe checks rollout status, deployment
JSON, canary pod state, ingress/routing resources, and pod metrics
- `docs/stage2-deploy-observe.md` explicitly treats metrics-server absence as a
separate observability note rather than as proof that rollout failed
Published guarantee for wave 1: rollout, pod, and ingress inspection are part
of the substrate contract. Metrics availability is opportunistic unless the
cluster profile is later strengthened.
## Namespace And Release Behavior
- `tools/cmd/railiance-stage2` and `tools/cmd/railiance-stage3` use
`--create-namespace`
- Stage defaults derive namespace and release names from workload contract data
when the workload does not override them
Future `rail-kubernetes` may assume that namespace creation is acceptable in the
current operator flow. It should not assume namespaces or release names are
pre-provisioned by S2.
## Storage And Stateful Constraints
- `tools/cmd/railiance-backup-s2` documents that the current cluster runs k3s
in SQLite mode and backs up `/var/lib/rancher/k3s/server/db/state.db`
- This repo does not yet publish a generic storage-class, PVC, or stateful
workload placement contract for downstream workload rails
That means stateful workload assumptions remain explicit-review territory for
now. `rail-kubernetes` should not invent a storage contract on behalf of S2.
## Safety Rails Published By S2
- age-encrypted cluster backup via `make backup`
- printed restore procedure via `make restore`
- preflight gate via `make preflight`
- smoke validation via `make smoke`
These are substrate safety rails. They are not a substitute for workload-level
promotion or rollback logic.
## Not Yet Published
The future `rail-kubernetes` must currently treat these as absent or
repo-specific until S2 publishes them explicitly:
- guaranteed weighted routing on the active ingress controller
- a generic workload RBAC model
- a generic storage-class or persistent-volume contract
- a secret-delivery contract at the S2 layer
- multi-node HA assumptions beyond the active `RAIL-BS-WP-0007` work