4.4 KiB
rail-kubernetes Substrate Profile
This document records the minimum cluster substrate facts that
railiance-cluster publishes for a future rail-kubernetes repo.
The intent is narrow: future rail code should depend on explicit substrate capabilities, not on incidental implementation details scattered across this repo.
Runtime Base
- Kubernetes runtime: k3s installed by
ansible/bootstrap.yml - Current pinned k3s version in repo:
v1.35.1+k3s1 - Current pinned Helm version in repo:
v3.17.3 - Baseline health checks:
make smoke - Cluster safety gate before invasive work:
make preflight
rail-kubernetes may assume a k3s-backed Kubernetes API plus Helm-based apply
paths on the current Railiance substrate. It should not assume more than that
without a declared extension.
Access Model
- Operator access is delivered from
/etc/rancher/k3s/k3s.yamlto~/.kube/config-hosteuropeas documented indocs/kubeconfig.md - The raw k3s kubeconfig uses
127.0.0.1as the server address and may require rewrite or tunneling for remote use - The documented operator model is cluster-admin kubeconfig access, not a least-privilege workload service-account contract
Future rail-kubernetes should assume operator-driven cluster access for now.
It should not assume a published per-workload RBAC template yet.
Routing And Canary Exposure
- The current generic tooling supports isolated canary deployment by creating a separate Stage 2 release and namespace
docs/canary-helm-template.mddocuments a Traefik weighted-routing example, butrailiance-clusterdoes not currently publish weighted routing as a guaranteed substrate capabilitytools/cmd/railiance-stage2can inspect rollout, pods, and ingress/routing resources for the canary release
Published guarantee for wave 1: isolated canary exposure is the safe default. Weighted, header-based, path-based, or shadow routing must be treated as optional future capabilities until S2 documents them explicitly.
TLS And Cluster Operators
SCOPE.md currently records these cluster-level addons/operators on the active
substrate:
- nginx ingress
- cert-manager
- ArgoCD
- CloudNative PG operator
- SSO stack namespaces (
mfa,sso) via Net Kingdom
Future rail-kubernetes should depend on capabilities such as ingress
availability or TLS readiness, not on owning these operators directly.
Observability Expectations
tools/cmd/railiance-stage2live observe checks rollout status, deployment JSON, canary pod state, ingress/routing resources, and pod metricsdocs/stage2-deploy-observe.mdexplicitly treats metrics-server absence as a separate observability note rather than as proof that rollout failed
Published guarantee for wave 1: rollout, pod, and ingress inspection are part of the substrate contract. Metrics availability is opportunistic unless the cluster profile is later strengthened.
Namespace And Release Behavior
tools/cmd/railiance-stage2andtools/cmd/railiance-stage3use--create-namespace- Stage defaults derive namespace and release names from workload contract data when the workload does not override them
Future rail-kubernetes may assume that namespace creation is acceptable in the
current operator flow. It should not assume namespaces or release names are
pre-provisioned by S2.
Storage And Stateful Constraints
tools/cmd/railiance-backup-s2documents that the current cluster runs k3s in SQLite mode and backs up/var/lib/rancher/k3s/server/db/state.db- This repo does not yet publish a generic storage-class, PVC, or stateful workload placement contract for downstream workload rails
That means stateful workload assumptions remain explicit-review territory for
now. rail-kubernetes should not invent a storage contract on behalf of S2.
Safety Rails Published By S2
- age-encrypted cluster backup via
make backup - printed restore procedure via
make restore - preflight gate via
make preflight - smoke validation via
make smoke
These are substrate safety rails. They are not a substitute for workload-level promotion or rollback logic.
Not Yet Published
The future rail-kubernetes must currently treat these as absent or
repo-specific until S2 publishes them explicitly:
- guaranteed weighted routing on the active ingress controller
- a generic workload RBAC model
- a generic storage-class or persistent-volume contract
- a secret-delivery contract at the S2 layer
- multi-node HA assumptions beyond the active
RAIL-BS-WP-0007work