Compare commits
7 commits
ed75ccc5f3
...
f1c717b791
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f1c717b791 | ||
|
|
433949ab57 | ||
|
|
b48eaea942 | ||
|
|
513407acaa | ||
|
|
e6867e5ca9 | ||
|
|
90c89a4d7d | ||
|
|
0e320c5632 |
24 changed files with 2356 additions and 1 deletions
24
.custodian-brief.md
Normal file
24
.custodian-brief.md
Normal file
|
|
@ -0,0 +1,24 @@
|
||||||
|
<!-- custodian-brief: generated by fix-consistency — do not edit manually -->
|
||||||
|
# Custodian Brief — rail-kubernetes
|
||||||
|
|
||||||
|
**Domain:** financials
|
||||||
|
**Last synced:** 2026-07-25 16:56 UTC
|
||||||
|
**State Hub:** http://127.0.0.1:8000 *(adjust if running on a remote machine)*
|
||||||
|
|
||||||
|
## Active Workstreams
|
||||||
|
|
||||||
|
### Bootstrap rail-kubernetes and import the wave-1 contract
|
||||||
|
Progress: 1/4 done | workplan_id: `1f2254f5-a873-46fc-a77c-bc725eeaaeb8`
|
||||||
|
|
||||||
|
**Open tasks:**
|
||||||
|
- ! T03 - Prepare the compatibility handoff from `railiance-cluster` `18140d90`
|
||||||
|
- ► T02 - Import the reviewed generic contract material `148276a2`
|
||||||
|
- ► T04 - Register `rail-kubernetes` in State Hub and Fabric `d78b5712`
|
||||||
|
|
||||||
|
---
|
||||||
|
## MCP Orientation (when available)
|
||||||
|
|
||||||
|
If the state-hub MCP server is reachable, call:
|
||||||
|
`get_domain_summary("financials")`
|
||||||
|
This provides richer cross-domain context.
|
||||||
|
If the MCP call fails, use this file as your orientation source.
|
||||||
7
.gitignore
vendored
Normal file
7
.gitignore
vendored
Normal file
|
|
@ -0,0 +1,7 @@
|
||||||
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
|
env/
|
||||||
18
.repo-classification.yaml
Normal file
18
.repo-classification.yaml
Normal file
|
|
@ -0,0 +1,18 @@
|
||||||
|
repo_classification:
|
||||||
|
standard: Repo Classification Standard
|
||||||
|
version: "1.0"
|
||||||
|
classified_at: "2026-07-25"
|
||||||
|
classified_by: agent
|
||||||
|
category: project
|
||||||
|
domain: financials
|
||||||
|
secondary_domains:
|
||||||
|
- infotech
|
||||||
|
capability_tags:
|
||||||
|
- platform
|
||||||
|
- operations
|
||||||
|
business_stake:
|
||||||
|
- technology
|
||||||
|
- operations
|
||||||
|
business_mechanics:
|
||||||
|
- coordination
|
||||||
|
- operation
|
||||||
41
AGENTS.md
Normal file
41
AGENTS.md
Normal file
|
|
@ -0,0 +1,41 @@
|
||||||
|
# rail-kubernetes — Agent Instructions
|
||||||
|
|
||||||
|
## Repo Identity
|
||||||
|
|
||||||
|
**Purpose:** Default Railiance execution rail for generic
|
||||||
|
workload-on-Kubernetes semantics.
|
||||||
|
|
||||||
|
**Domain:** financials
|
||||||
|
**Repo slug:** rail-kubernetes
|
||||||
|
**Topic slug:** railiance
|
||||||
|
**Workplan prefix:** `RAIL-K8S-WP-`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Session Protocol
|
||||||
|
|
||||||
|
**Start:**
|
||||||
|
1. Read `README.md`, `INTENT.md`, and `SCOPE.md`
|
||||||
|
2. Review active work in `workplans/`
|
||||||
|
3. Re-read the current rail declaration in `declarations/rail.yaml`
|
||||||
|
|
||||||
|
**During work:**
|
||||||
|
- Keep the boundary with `railiance-cluster` explicit
|
||||||
|
- Keep workload packaging concerns out of this repo
|
||||||
|
- Prefer source-controlled declarations over implicit conventions
|
||||||
|
|
||||||
|
**Close:**
|
||||||
|
1. Update workplan files to reflect progress
|
||||||
|
2. Register or refresh the repo from `.repo-classification.yaml` when needed
|
||||||
|
3. Run `statehub fix-consistency`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Boundary Reminders
|
||||||
|
|
||||||
|
- Do not move substrate bootstrap into this repo
|
||||||
|
- Do not normalize workload-specific helper commands here
|
||||||
|
- Do not let `rapp-*` packaging or `reef-*` substrate inventory drift into this repo
|
||||||
|
|
||||||
|
The canonical framework boundary lives in
|
||||||
|
`railiance-master/docs/rail-kubernetes-boundary.md`.
|
||||||
4
CLAUDE.md
Normal file
4
CLAUDE.md
Normal file
|
|
@ -0,0 +1,4 @@
|
||||||
|
# rail-kubernetes — Claude Code Instructions
|
||||||
|
|
||||||
|
@SCOPE.md
|
||||||
|
@AGENTS.md
|
||||||
55
INTENT.md
Normal file
55
INTENT.md
Normal file
|
|
@ -0,0 +1,55 @@
|
||||||
|
# INTENT
|
||||||
|
|
||||||
|
## Why This Repo Exists
|
||||||
|
|
||||||
|
`rail-kubernetes` exists so Railiance has a first-class home for the default
|
||||||
|
execution rail used by platform services and managed workloads.
|
||||||
|
|
||||||
|
Before this repo, generic workload lifecycle semantics, workload declarations,
|
||||||
|
and Kubernetes rollout helpers were mixed into `railiance-cluster`. That was
|
||||||
|
pragmatic while Railiance had only one runtime path, but it is the wrong steady
|
||||||
|
state once Railiance starts introducing multiple rails.
|
||||||
|
|
||||||
|
This repo establishes the durable answer to:
|
||||||
|
|
||||||
|
- how a Railiance-managed workload targets Kubernetes as a rail
|
||||||
|
- which workload lifecycle semantics are generic to Kubernetes-managed
|
||||||
|
workloads
|
||||||
|
- which substrate guarantees the rail expects from S2
|
||||||
|
- where future packages such as `rapp-openbao` attach when they run on the
|
||||||
|
default rail
|
||||||
|
|
||||||
|
## What This Repo Must Protect
|
||||||
|
|
||||||
|
This repo must keep the wave-1 Kubernetes rail narrow and reusable.
|
||||||
|
|
||||||
|
That means:
|
||||||
|
|
||||||
|
- protect the distinction between rail semantics and cluster ownership
|
||||||
|
- protect the distinction between rail semantics and workload packaging
|
||||||
|
- keep the compatibility path from `railiance-cluster` explicit while the
|
||||||
|
extraction is incomplete
|
||||||
|
- avoid encoding assumptions that only work for one workload or one substrate
|
||||||
|
|
||||||
|
## What This Repo Is Not
|
||||||
|
|
||||||
|
This repo is not:
|
||||||
|
|
||||||
|
- the ownership home for the Kubernetes substrate
|
||||||
|
- the ownership home for OpenBao, Forgejo, or any other single workload
|
||||||
|
- the ownership home for one specific server or cluster substrate
|
||||||
|
- the place to normalize workload-specific operator commands
|
||||||
|
|
||||||
|
## Initial Operating Context
|
||||||
|
|
||||||
|
Wave 1 is intentionally centered on Kubernetes.
|
||||||
|
|
||||||
|
Railiance should treat this repo as the default rail path until a later
|
||||||
|
workload presents a strong enough reason to justify a different rail such as
|
||||||
|
`rail-knative`.
|
||||||
|
|
||||||
|
The first version of this repo therefore prioritizes:
|
||||||
|
|
||||||
|
- a clean boundary with `railiance-cluster`
|
||||||
|
- a stable attachment point for future `rapp-*` repos
|
||||||
|
- compatibility with the currently published S2 substrate profile
|
||||||
46
README.md
46
README.md
|
|
@ -1,3 +1,47 @@
|
||||||
# rail-kubernetes
|
# rail-kubernetes
|
||||||
|
|
||||||
Railiance support for kubernetes workloads
|
Wave-1 execution rail for Railiance-managed workloads on Kubernetes.
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
`rail-kubernetes` is the default base rail for the current Railiance
|
||||||
|
ecosystem. It owns the generic workload-on-Kubernetes execution contract that
|
||||||
|
should no longer stay mixed into `railiance-cluster`.
|
||||||
|
|
||||||
|
This repo exists to hold:
|
||||||
|
|
||||||
|
- the Kubernetes rail declaration
|
||||||
|
- the generic workload lifecycle contract
|
||||||
|
- the migration path from the current `railiance-cluster` compatibility surface
|
||||||
|
- future rail-local tooling and docs for Railiance-managed workloads on Kubernetes
|
||||||
|
|
||||||
|
## Boundary
|
||||||
|
|
||||||
|
`rail-kubernetes` does not own:
|
||||||
|
|
||||||
|
- k3s bootstrap
|
||||||
|
- cluster-level operator ownership
|
||||||
|
- cluster backup/restore posture
|
||||||
|
- workload-specific wrappers that belong in `rapp-*` repos
|
||||||
|
- substrate inventories that belong in `reef-*` repos
|
||||||
|
|
||||||
|
Those concerns stay with `railiance-cluster`, the relevant `rapp-*` repo, or
|
||||||
|
the relevant `reef-*` repo.
|
||||||
|
|
||||||
|
## Current Status
|
||||||
|
|
||||||
|
This repo is in first-wave bootstrap.
|
||||||
|
|
||||||
|
The boundary and source material come from:
|
||||||
|
|
||||||
|
- `railiance-master/docs/rail-kubernetes-boundary.md`
|
||||||
|
- `railiance-cluster/docs/rail-kubernetes-extraction-map.md`
|
||||||
|
- `railiance-cluster/docs/rail-kubernetes-substrate-profile.md`
|
||||||
|
|
||||||
|
## Initial Layout
|
||||||
|
|
||||||
|
- `declarations/rail.yaml` — source-controlled rail contract
|
||||||
|
- `docs/` — wave-1 contract and import plan
|
||||||
|
- `schemas/` — machine-readable workload contract schema
|
||||||
|
- `examples/` — canonical example workload declarations for the rail
|
||||||
|
- `workplans/` — repo-local execution work
|
||||||
|
|
|
||||||
118
SCOPE.md
Normal file
118
SCOPE.md
Normal file
|
|
@ -0,0 +1,118 @@
|
||||||
|
# SCOPE
|
||||||
|
|
||||||
|
> This file helps you quickly understand what this repository is about,
|
||||||
|
> when it is relevant, and when it is not.
|
||||||
|
> It is intentionally lightweight and may be incomplete.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## One-liner
|
||||||
|
|
||||||
|
Default Railiance execution rail for generic workload-on-Kubernetes semantics.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Core Idea
|
||||||
|
|
||||||
|
`rail-kubernetes` is the first concrete `rail-*` repo in Railiance.
|
||||||
|
It exists to hold the generic execution contract for Railiance-managed
|
||||||
|
workloads that run on Kubernetes, without keeping those semantics mixed into
|
||||||
|
the S2 ownership repo.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## In Scope
|
||||||
|
|
||||||
|
- Kubernetes rail declaration and compatibility notes
|
||||||
|
- Generic workload lifecycle semantics for Kubernetes-managed workloads
|
||||||
|
- Generic workload contract surface such as `railiance/app.toml`
|
||||||
|
- Generic rollout, observe, promote, and rollback semantics for the rail
|
||||||
|
- Migration compatibility guidance while the extraction from `railiance-cluster`
|
||||||
|
is still in progress
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Out of Scope
|
||||||
|
|
||||||
|
- k3s bootstrap and cluster installation
|
||||||
|
- cluster-level ingress, operator, and certificate ownership
|
||||||
|
- cluster backup, restore, and preflight safety rails
|
||||||
|
- workload-specific wrappers and app packaging
|
||||||
|
- substrate inventories and machine grouping
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Relevant When
|
||||||
|
|
||||||
|
- Defining how a workload targets the default Kubernetes rail
|
||||||
|
- Moving generic workload lifecycle material out of `railiance-cluster`
|
||||||
|
- Attaching a `rapp-*` package to its Kubernetes execution contract
|
||||||
|
- Clarifying which S2 substrate facts the rail may depend on
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Not Relevant When
|
||||||
|
|
||||||
|
- The work is purely cluster substrate ownership
|
||||||
|
- The work is workload-specific packaging
|
||||||
|
- The work is substrate inventory or server grouping
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Current State
|
||||||
|
|
||||||
|
- Status: active bootstrap
|
||||||
|
- Implementation: baseline repo, declaration, and first repo-local workplan are present
|
||||||
|
- Stability: evolving
|
||||||
|
- Usage: wave-1 home for the default Kubernetes rail
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## How It Fits
|
||||||
|
|
||||||
|
- Upstream dependencies: `railiance-cluster`, `railiance-master`
|
||||||
|
- Downstream consumers: future `rapp-*` repos and reefs that host the Kubernetes rail
|
||||||
|
- Often used with: `railiance-fabric`, `railiance-platform`, `railiance-infra`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Terminology
|
||||||
|
|
||||||
|
- Preferred terms: rail, execution contract, workload lifecycle, substrate profile
|
||||||
|
- Also known as: default base rail
|
||||||
|
- Potentially confusing terms: this repo defines rail semantics; it does not own the Kubernetes substrate itself
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Related / Overlapping Repositories
|
||||||
|
|
||||||
|
- `railiance-cluster` — owns the Kubernetes substrate; `rail-kubernetes` owns the generic workload-on-Kubernetes contract
|
||||||
|
- `railiance-master` — defines the framework architecture and boundary
|
||||||
|
- future `rapp-*` repos — own workload packaging and scaffolding on top of this rail
|
||||||
|
- future `reef-*` repos — own the substrate boundaries that host this rail
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Getting Oriented
|
||||||
|
|
||||||
|
- Start with: `README.md`, `INTENT.md`
|
||||||
|
- Key files / directories: `declarations/`, `docs/`, `workplans/`
|
||||||
|
- Entry points: `declarations/rail.yaml`, `docs/wave-1-contract.md`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Provided Capabilities
|
||||||
|
|
||||||
|
```capability
|
||||||
|
type: governance
|
||||||
|
title: Kubernetes execution rail contract
|
||||||
|
description: Defines the generic execution contract for Railiance-managed workloads that run on Kubernetes.
|
||||||
|
keywords: [railiance, rail, kubernetes, workload, execution]
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
Keep this repo generic. If a behavior is specific to one workload or one
|
||||||
|
substrate, it belongs somewhere else.
|
||||||
15
WORK-RECORDS.md
Normal file
15
WORK-RECORDS.md
Normal file
|
|
@ -0,0 +1,15 @@
|
||||||
|
# Work Records — rail-kubernetes
|
||||||
|
|
||||||
|
> Generated by `statehub fix-consistency` (CUST-WP-0061-T04, work-record
|
||||||
|
> stage 3). Do not edit by hand — edit the source file/block listed for
|
||||||
|
> each record and re-run fix-consistency to refresh this index. Archived
|
||||||
|
> workplans are omitted; closed decisions/intakes/engagements stay listed
|
||||||
|
> so recently-resolved work is still visible. [auto]
|
||||||
|
|
||||||
|
| Kind | ID | Status | Lane | Source |
|
||||||
|
| --- | --- | --- | --- | --- |
|
||||||
|
| workplan | RAIL-K8S-WP-0001 | active | — | workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md |
|
||||||
|
| task | RAIL-K8S-WP-0001-T01 | done | — | workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md |
|
||||||
|
| task | RAIL-K8S-WP-0001-T02 | progress | — | workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md |
|
||||||
|
| task | RAIL-K8S-WP-0001-T03 | wait | — | workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md |
|
||||||
|
| task | RAIL-K8S-WP-0001-T04 | progress | — | workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md |
|
||||||
26
declarations/rail.yaml
Normal file
26
declarations/rail.yaml
Normal file
|
|
@ -0,0 +1,26 @@
|
||||||
|
kind: execution-rail
|
||||||
|
rail_id: rail-kubernetes
|
||||||
|
repo: rail-kubernetes
|
||||||
|
ownership_repo: railiance-cluster
|
||||||
|
execution_architecture: kubernetes
|
||||||
|
status: wave-1-bootstrap
|
||||||
|
supported_rollout_modes:
|
||||||
|
- stage1-run
|
||||||
|
- stage2-isolated-canary
|
||||||
|
- stage3-promote-rollback
|
||||||
|
substrate_prerequisites:
|
||||||
|
- kubernetes-api
|
||||||
|
- helm-apply-path
|
||||||
|
- operator-driven-kubeconfig-access
|
||||||
|
- namespace-creation-supported
|
||||||
|
compatibility_notes:
|
||||||
|
- Weighted routing is optional and must not be assumed in wave 1.
|
||||||
|
- Metrics availability is opportunistic and not a rollout success criterion.
|
||||||
|
- Storage guarantees remain explicit-review territory until S2 publishes them.
|
||||||
|
source_documents:
|
||||||
|
- repo: railiance-master
|
||||||
|
path: docs/rail-kubernetes-boundary.md
|
||||||
|
- repo: railiance-cluster
|
||||||
|
path: docs/rail-kubernetes-extraction-map.md
|
||||||
|
- repo: railiance-cluster
|
||||||
|
path: docs/rail-kubernetes-substrate-profile.md
|
||||||
11
docs/README.md
Normal file
11
docs/README.md
Normal file
|
|
@ -0,0 +1,11 @@
|
||||||
|
# rail-kubernetes Docs
|
||||||
|
|
||||||
|
- `app-toml-contract.md` — canonical workload declaration contract for the rail
|
||||||
|
- `overlay-repo-pattern.md` — compatibility-era overlay pattern on the path to `rapp-*`
|
||||||
|
- `canary-helm-template.md` — stage-aware canary/stable chart pattern
|
||||||
|
- `deployment-lifecycle.md` — canonical stage 1 / 2 / 3 lifecycle contract
|
||||||
|
- `stage2-deploy-observe.md` — stage 2 command behavior contract
|
||||||
|
- `promote-rollback-onboarding.md` — representative lifecycle command path
|
||||||
|
- `railiance-run-command.md` — stage 1 command behavior contract
|
||||||
|
- `wave-1-contract.md` — current rail contract and boundary summary
|
||||||
|
- `source-import-plan.md` — source material and migration direction from existing repos
|
||||||
240
docs/app-toml-contract.md
Normal file
240
docs/app-toml-contract.md
Normal file
|
|
@ -0,0 +1,240 @@
|
||||||
|
# Railiance app.toml Contract
|
||||||
|
|
||||||
|
This document defines the repository-local `railiance/app.toml` contract used by
|
||||||
|
Railiance staged promotion tooling. The file tells Railiance how a workload
|
||||||
|
moves through Stage 1 local validation, Stage 2 production canary, and Stage 3
|
||||||
|
production promotion without relying on bespoke operator notes.
|
||||||
|
|
||||||
|
The contract is intentionally declarative. Commands, health checks, platform
|
||||||
|
dependencies, and secret references are described by stable names. Plaintext
|
||||||
|
secrets, bearer tokens, kubeconfigs, and private key material must never appear
|
||||||
|
in `railiance/app.toml`.
|
||||||
|
|
||||||
|
Ownership note: this contract now lives in `rail-kubernetes` as part of the
|
||||||
|
wave-1 split. `railiance-cluster` may keep compatibility copies during the
|
||||||
|
migration window, but long-term ownership belongs with the generic
|
||||||
|
workload-on-Kubernetes contract here rather than with S2 cluster substrate
|
||||||
|
ownership.
|
||||||
|
|
||||||
|
The machine-readable schema lives at `schemas/railiance-app.schema.json`. A
|
||||||
|
minimal example lives at `examples/railiance/app.toml`.
|
||||||
|
|
||||||
|
## File Location
|
||||||
|
|
||||||
|
Participating workload repositories declare the contract at:
|
||||||
|
|
||||||
|
```text
|
||||||
|
railiance/app.toml
|
||||||
|
```
|
||||||
|
|
||||||
|
Overlay repositories for third-party applications use the same path in the
|
||||||
|
overlay repo, not in the upstream source repository.
|
||||||
|
|
||||||
|
## Versioning
|
||||||
|
|
||||||
|
Every file must include:
|
||||||
|
|
||||||
|
```toml
|
||||||
|
schema_version = "railiance.app.v1"
|
||||||
|
```
|
||||||
|
|
||||||
|
Breaking contract changes require a new schema version. Tooling must fail closed
|
||||||
|
when it sees an unsupported `schema_version`.
|
||||||
|
|
||||||
|
## Top-Level Sections
|
||||||
|
|
||||||
|
### app
|
||||||
|
|
||||||
|
Identifies the workload and its ownership boundary.
|
||||||
|
|
||||||
|
Required fields:
|
||||||
|
|
||||||
|
- `id`: stable lowercase id using letters, numbers, and hyphens.
|
||||||
|
- `name`: human-readable workload name.
|
||||||
|
- `repo`: owning source or overlay repository slug.
|
||||||
|
- `owner`: owning team, domain, or operator group.
|
||||||
|
- `criticality`: one of `low`, `medium`, `high`, or `critical`.
|
||||||
|
- `description`: short purpose statement.
|
||||||
|
|
||||||
|
Production-critical workloads include source forge, identity, State Hub,
|
||||||
|
Inter-Hub, databases, object stores, backup systems, ingress, and cluster-wide
|
||||||
|
policy controllers. For those workloads, `criticality = "critical"` requires
|
||||||
|
explicit human approval before Stage 2 traffic exposure and Stage 3 promotion.
|
||||||
|
|
||||||
|
### source
|
||||||
|
|
||||||
|
Identifies the candidate under promotion.
|
||||||
|
|
||||||
|
Required fields:
|
||||||
|
|
||||||
|
- `revision`: commit id, tag, or immutable source revision expression.
|
||||||
|
- `artifact`: artifact kind, normally `image`, `helm-chart`, or `bundle`.
|
||||||
|
- `digest_policy`: one of `required`, `preferred`, or `not-applicable`.
|
||||||
|
|
||||||
|
If an image is promoted, Stage 2 and Stage 3 tooling should prefer immutable
|
||||||
|
image digests over mutable tags.
|
||||||
|
|
||||||
|
### platform.dependencies
|
||||||
|
|
||||||
|
Declares platform services required before canary or production promotion.
|
||||||
|
|
||||||
|
Each dependency has:
|
||||||
|
|
||||||
|
- `name`: stable service name.
|
||||||
|
- `kind`: dependency kind such as `postgres`, `redis`, `object-store`,
|
||||||
|
`identity`, `state-hub`, `inter-hub`, `network`, or `other`.
|
||||||
|
- `required`: boolean.
|
||||||
|
- `stage`: earliest stage that needs it, one of `stage1`, `stage2`, `stage3`.
|
||||||
|
- `evidence`: non-secret evidence expected before promotion, such as a health
|
||||||
|
endpoint result, Kubernetes Ready condition, or State Hub progress id.
|
||||||
|
|
||||||
|
### secrets.references
|
||||||
|
|
||||||
|
Declares required secret references without secret values.
|
||||||
|
|
||||||
|
Each reference has:
|
||||||
|
|
||||||
|
- `name`: workload-local secret name.
|
||||||
|
- `route`: approved credential route id, for example `openbao-api-key`,
|
||||||
|
`key-cape-oidc-login`, or `activity-core-issue-sink`.
|
||||||
|
- `target`: non-secret target reference such as a Kubernetes Secret name,
|
||||||
|
ExternalSecret name, OpenBao path, or environment variable name.
|
||||||
|
- `stage`: earliest stage that needs the secret.
|
||||||
|
- `required`: boolean.
|
||||||
|
|
||||||
|
Forbidden fields include plaintext values, tokens, passwords, kubeconfigs, or
|
||||||
|
private keys. Tooling must reject suspicious field names such as `value`,
|
||||||
|
`token`, `password`, `secret`, `private_key`, or `kubeconfig` inside secret
|
||||||
|
reference objects unless they are part of the approved non-secret `target` text.
|
||||||
|
|
||||||
|
### observability
|
||||||
|
|
||||||
|
Defines how promotion tooling proves the workload is alive and observable.
|
||||||
|
|
||||||
|
Required fields:
|
||||||
|
|
||||||
|
- `health_endpoints`: one or more HTTP health endpoint declarations.
|
||||||
|
- `metrics`: optional metrics endpoint or query references.
|
||||||
|
- `logs`: optional log selectors or query references.
|
||||||
|
|
||||||
|
Health endpoint declarations include `name`, `url`, `stage`, and expected
|
||||||
|
status code. URLs may be internal service URLs for Stage 2/3; they must not
|
||||||
|
embed credentials.
|
||||||
|
|
||||||
|
### rollback
|
||||||
|
|
||||||
|
Defines how the workload returns to a previous stable state.
|
||||||
|
|
||||||
|
Required fields:
|
||||||
|
|
||||||
|
- `strategy`: one of `helm-revision`, `image-digest`, `traffic-shift`,
|
||||||
|
`manual-runbook`, or `none`.
|
||||||
|
- `command`: command name or runbook path. This may be a placeholder before
|
||||||
|
T07 implements automation, but it must tell the operator where rollback lives.
|
||||||
|
- `verification`: non-secret check to confirm rollback succeeded.
|
||||||
|
|
||||||
|
`strategy = "none"` is allowed only for Stage 1-only workloads and must not be
|
||||||
|
used for production-critical workloads.
|
||||||
|
|
||||||
|
## Stage Sections
|
||||||
|
|
||||||
|
The contract has one table for each stage:
|
||||||
|
|
||||||
|
```toml
|
||||||
|
[stages.stage1]
|
||||||
|
[stages.stage2]
|
||||||
|
[stages.stage3]
|
||||||
|
```
|
||||||
|
|
||||||
|
Each stage includes:
|
||||||
|
|
||||||
|
- `enabled`: boolean.
|
||||||
|
- `namespace`: target Kubernetes namespace, or a local namespace for Stage 1.
|
||||||
|
- `release`: release identity.
|
||||||
|
- `commands`: ordered command aliases or shell commands that tooling may run.
|
||||||
|
- `checks`: ordered check ids to evaluate.
|
||||||
|
- `evidence`: expected non-secret evidence outputs.
|
||||||
|
- `requires_approval`: boolean.
|
||||||
|
|
||||||
|
Stage 2 additionally includes `canary_mode`, one of `weighted`, `header`,
|
||||||
|
`path`, `shadow`, or `isolated`, plus `observation_minutes` and optional
|
||||||
|
`traffic_percent` when weighted routing is used.
|
||||||
|
|
||||||
|
Stage 3 additionally includes `promotion_mode`, one of `traffic-shift`,
|
||||||
|
`release-replace`, `selector-switch`, or `workflow`, plus `previous_stable`.
|
||||||
|
|
||||||
|
## Check Definitions
|
||||||
|
|
||||||
|
Checks live under `[[checks]]` entries and are referenced by stage `checks`.
|
||||||
|
|
||||||
|
Required fields:
|
||||||
|
|
||||||
|
- `id`: stable check id.
|
||||||
|
- `type`: one of `command`, `http`, `kubernetes`, `helm`, `metric`, `log`, or
|
||||||
|
`manual`.
|
||||||
|
- `stage`: earliest stage that may run the check.
|
||||||
|
- `description`: human-readable purpose.
|
||||||
|
- `required`: boolean.
|
||||||
|
|
||||||
|
Type-specific fields:
|
||||||
|
|
||||||
|
- `command`: `run` command string and optional `timeout_seconds`.
|
||||||
|
- `http`: `url`, `expected_status`, and optional `timeout_seconds`.
|
||||||
|
- `kubernetes`: `namespace`, `resource`, and `condition`.
|
||||||
|
- `helm`: `chart`, `values`, and `mode` such as `template` or
|
||||||
|
`server-dry-run`.
|
||||||
|
- `metric`: `query`, `window_minutes`, and `threshold`.
|
||||||
|
- `log`: `selector`, `window_minutes`, and `forbidden_patterns`.
|
||||||
|
- `manual`: `evidence_required` text.
|
||||||
|
|
||||||
|
Checks must not print secrets. If a check needs secret-backed access, the result
|
||||||
|
records only the route, target object, and pass/fail state.
|
||||||
|
|
||||||
|
## Command Semantics
|
||||||
|
|
||||||
|
Commands in `app.toml` are declarations for Railiance tooling. Stage 1, Stage
|
||||||
|
2, and Stage 3 commands now have local CLI support; workload scripts may still
|
||||||
|
wrap them for service-specific checks.
|
||||||
|
|
||||||
|
Expected mapping:
|
||||||
|
|
||||||
|
- Stage 1 commands are consumed by `bin/railiance run <overlay-dir>`.
|
||||||
|
- Stage 2 commands are consumed by `bin/railiance deploy --stage 2 <overlay-dir>`
|
||||||
|
and `bin/railiance observe --stage 2 <overlay-dir>`.
|
||||||
|
- Stage 3 commands are consumed by `bin/railiance promote <overlay-dir>` and
|
||||||
|
`bin/railiance rollback <overlay-dir>`.
|
||||||
|
|
||||||
|
Tooling must emit machine-readable results with workload identity, candidate
|
||||||
|
revision, checks run, pass/fail status, non-secret evidence, rollback target,
|
||||||
|
and approval state.
|
||||||
|
|
||||||
|
## Minimal Example
|
||||||
|
|
||||||
|
See `examples/railiance/app.toml`. It declares a critical internal service with:
|
||||||
|
|
||||||
|
- immutable image digest requirement;
|
||||||
|
- Stage 1 local validation;
|
||||||
|
- Stage 2 isolated canary;
|
||||||
|
- Stage 3 release replacement;
|
||||||
|
- OpenBao-routed secret references without values;
|
||||||
|
- HTTP, Helm, Kubernetes, and manual approval checks.
|
||||||
|
|
||||||
|
## Adoption Rules
|
||||||
|
|
||||||
|
A workload can enter Stage 1 when `app.toml` passes schema validation and all
|
||||||
|
Stage 1 required checks are declared.
|
||||||
|
|
||||||
|
A workload can enter Stage 2 only when:
|
||||||
|
|
||||||
|
- Stage 1 passed for the same candidate artifact;
|
||||||
|
- Stage 2 namespace, release, canary mode, health checks, dependencies, and
|
||||||
|
rollback target are declared;
|
||||||
|
- secret references use approved routes and contain no values;
|
||||||
|
- production-critical workloads have explicit approval.
|
||||||
|
|
||||||
|
A workload can enter Stage 3 only when:
|
||||||
|
|
||||||
|
- Stage 2 acceptance gates passed for the same candidate artifact;
|
||||||
|
- `previous_stable` and rollback verification are recorded;
|
||||||
|
- backup/restore posture is current for stateful workloads;
|
||||||
|
- production-critical workloads have explicit human approval.
|
||||||
55
docs/canary-helm-template.md
Normal file
55
docs/canary-helm-template.md
Normal file
|
|
@ -0,0 +1,55 @@
|
||||||
|
# Canary Helm Template
|
||||||
|
|
||||||
|
Generated Railiance overlays include a stage-aware Helm chart for Stage 2
|
||||||
|
canaries and Stage 3 stable promotion.
|
||||||
|
|
||||||
|
The chart keeps stable and canary release identities explicit:
|
||||||
|
|
||||||
|
- `railiance.stableRelease` names the current stable release;
|
||||||
|
- `railiance.canaryRelease` names the Stage 2 candidate release;
|
||||||
|
- `railiance.stage` selects the rendered identity, labels, and selectors;
|
||||||
|
- `railiance.previousStable` records rollback context before promotion.
|
||||||
|
|
||||||
|
## Traffic Shape
|
||||||
|
|
||||||
|
The default Stage 2 values use an isolated canary ingress:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
railiance:
|
||||||
|
stage: canary
|
||||||
|
traffic:
|
||||||
|
mode: isolated
|
||||||
|
|
||||||
|
ingress:
|
||||||
|
enabled: true
|
||||||
|
```
|
||||||
|
|
||||||
|
This creates canary Deployment, Service, and Ingress resources without changing
|
||||||
|
the stable release. For environments that use Traefik weighted routing, set:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
railiance:
|
||||||
|
traffic:
|
||||||
|
mode: weighted
|
||||||
|
provider: traefik
|
||||||
|
stableWeight: 95
|
||||||
|
canaryWeight: 5
|
||||||
|
```
|
||||||
|
|
||||||
|
The chart then renders a `TraefikService` and `IngressRoute` that split traffic
|
||||||
|
between the stable and canary services. Other ingress controllers can use the
|
||||||
|
same stable/canary values layout with controller-specific annotations or a later
|
||||||
|
provider template.
|
||||||
|
|
||||||
|
## Observability And Safety
|
||||||
|
|
||||||
|
Generated workloads include:
|
||||||
|
|
||||||
|
- Prometheus-compatible scrape annotations on pods and services;
|
||||||
|
- readiness and liveness HTTP probes;
|
||||||
|
- conservative resource requests/limits for single-node clusters;
|
||||||
|
- separate `values/stage2-canary.yaml` and `values/stage3-production.yaml` so
|
||||||
|
canary exposure and stable promotion can be reviewed independently.
|
||||||
|
|
||||||
|
Run `tests/stage2-template.sh` in the overlay repo before any Stage 2 attempt.
|
||||||
|
It verifies the scaffold and runs `helm template` when Helm is available.
|
||||||
340
docs/deployment-lifecycle.md
Normal file
340
docs/deployment-lifecycle.md
Normal file
|
|
@ -0,0 +1,340 @@
|
||||||
|
# Railiance Deployment Lifecycle
|
||||||
|
|
||||||
|
This document defines the Railiance three-stage promotion lifecycle for
|
||||||
|
workloads that run on the Railiance Kubernetes substrate.
|
||||||
|
|
||||||
|
The lifecycle exists so production workloads move through repeatable gates
|
||||||
|
instead of one-off operator memory. It is intentionally conservative: every
|
||||||
|
stage must leave evidence, every promotion must have a rollback path, and
|
||||||
|
critical workloads require explicit human approval before production traffic is
|
||||||
|
changed.
|
||||||
|
|
||||||
|
Ownership note: this specification now lives in `rail-kubernetes` as the
|
||||||
|
wave-1 home of the default Kubernetes execution rail. `railiance-cluster` may
|
||||||
|
retain compatibility shims during migration, but long-term ownership of the
|
||||||
|
generic lifecycle contract belongs here rather than in the S2 substrate repo.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This specification now lives beside the Kubernetes rail contract rather than
|
||||||
|
beside the cluster substrate. The compatibility tooling may still live
|
||||||
|
elsewhere during migration, but long-term ownership is not S2.
|
||||||
|
|
||||||
|
Repo boundaries:
|
||||||
|
|
||||||
|
- `rail-kubernetes` owns lifecycle semantics, generic promotion gates,
|
||||||
|
canary validation, production routing expectations, rollback expectations,
|
||||||
|
and the generic workload contract.
|
||||||
|
- `railiance-cluster` owns only the substrate prerequisites published to that
|
||||||
|
rail: cluster connectivity, namespace behavior, ingress availability, and
|
||||||
|
operator/runtime expectations.
|
||||||
|
- `railiance-apps` owns workload-specific Helm values, application release
|
||||||
|
definitions, and production workload configuration.
|
||||||
|
- `railiance-platform` owns shared platform services such as databases,
|
||||||
|
caches, object storage, and backup targets.
|
||||||
|
- `railiance-enablement` owns developer-facing templates, CI workflows, and
|
||||||
|
local ergonomics.
|
||||||
|
- `railiance-infra` owns host provisioning, OS hardening, SSH, firewall, and
|
||||||
|
node bootstrap below Kubernetes.
|
||||||
|
|
||||||
|
## Lifecycle Overview
|
||||||
|
|
||||||
|
Railiance promotes a workload through three stages:
|
||||||
|
|
||||||
|
1. Stage 1: local validation.
|
||||||
|
2. Stage 2: production canary.
|
||||||
|
3. Stage 3: production promotion.
|
||||||
|
|
||||||
|
The stages are sequential. A workload may return to an earlier stage at any
|
||||||
|
time, but it must not skip a stage when moving toward production unless an
|
||||||
|
operator records an emergency exception in State Hub.
|
||||||
|
|
||||||
|
Each stage emits a machine-readable result with:
|
||||||
|
|
||||||
|
- workload identity;
|
||||||
|
- source revision or image digest;
|
||||||
|
- target environment;
|
||||||
|
- checks run;
|
||||||
|
- pass/fail status;
|
||||||
|
- non-secret evidence references;
|
||||||
|
- rollback target, when applicable;
|
||||||
|
- approving human or explicit "not required" decision.
|
||||||
|
|
||||||
|
## Workload Declaration
|
||||||
|
|
||||||
|
Each participating workload should declare its promotion contract in a
|
||||||
|
repository-local `railiance/app.toml`. The contract is defined in
|
||||||
|
`docs/app-toml-contract.md`, with a machine-readable schema at
|
||||||
|
`schemas/railiance-app.schema.json`. This lifecycle expects every workload
|
||||||
|
declaration to provide at least:
|
||||||
|
|
||||||
|
- stable workload name and owning repo;
|
||||||
|
- source revision, image tag, or image digest policy;
|
||||||
|
- stage-specific namespaces or release names;
|
||||||
|
- health checks and observability endpoints;
|
||||||
|
- ingress or routing targets;
|
||||||
|
- platform dependencies;
|
||||||
|
- rollback command or previous-stable reference;
|
||||||
|
- secret references by name or path, never plaintext secret values.
|
||||||
|
|
||||||
|
If a workload cannot provide a machine-readable declaration yet, it may still
|
||||||
|
use this lifecycle through a written operator runbook, but that is a temporary
|
||||||
|
compatibility path. The runbook must identify the missing declaration fields.
|
||||||
|
|
||||||
|
## Stage 1: Local Validation
|
||||||
|
|
||||||
|
Stage 1 proves that the workload can be built, configured, and checked outside
|
||||||
|
production traffic.
|
||||||
|
|
||||||
|
Typical Stage 1 targets:
|
||||||
|
|
||||||
|
- local container runtime;
|
||||||
|
- local Kubernetes such as k3d, kind, or a disposable namespace;
|
||||||
|
- dry-run Helm rendering;
|
||||||
|
- unit, integration, migration, and smoke checks that do not require production
|
||||||
|
credentials.
|
||||||
|
|
||||||
|
Required Stage 1 checks:
|
||||||
|
|
||||||
|
- source revision is cleanly identified;
|
||||||
|
- build or artifact selection is deterministic;
|
||||||
|
- Helm templates or manifests render without invalid Kubernetes objects;
|
||||||
|
- local health checks pass;
|
||||||
|
- required secrets are referenced by name only and are not printed;
|
||||||
|
- database migrations, if any, are classified as reversible, forward-only, or
|
||||||
|
requiring human approval;
|
||||||
|
- a Stage 2 candidate artifact is named by immutable digest or equivalent
|
||||||
|
immutable revision.
|
||||||
|
|
||||||
|
Stage 1 fails closed when:
|
||||||
|
|
||||||
|
- local checks are skipped without an approved reason;
|
||||||
|
- generated manifests contain plaintext secrets;
|
||||||
|
- the artifact cannot be traced to source;
|
||||||
|
- the workload cannot state how it will be observed in Stage 2.
|
||||||
|
|
||||||
|
Stage 1 completion does not authorize production traffic. It only makes a
|
||||||
|
workload eligible for Stage 2 review.
|
||||||
|
|
||||||
|
## Stage 2: Production Canary
|
||||||
|
|
||||||
|
Stage 2 deploys the candidate to production infrastructure with limited or
|
||||||
|
isolated exposure. The goal is to observe the candidate against real platform
|
||||||
|
dependencies while keeping blast radius small.
|
||||||
|
|
||||||
|
Acceptable canary forms:
|
||||||
|
|
||||||
|
- weighted ingress split between stable and canary;
|
||||||
|
- header-based or path-based routing for operator traffic only;
|
||||||
|
- shadow deployment receiving replicated non-mutating traffic;
|
||||||
|
- isolated production namespace with manually triggered probes.
|
||||||
|
|
||||||
|
The selected canary form must be declared before deployment. If weighted
|
||||||
|
routing is unavailable, the fallback must preserve the same safety property:
|
||||||
|
the candidate can be observed without silently replacing stable production.
|
||||||
|
|
||||||
|
Required Stage 2 prechecks:
|
||||||
|
|
||||||
|
- Stage 1 result passed for the same candidate artifact;
|
||||||
|
- cluster connectivity and namespace readiness are verified;
|
||||||
|
- target image digest or immutable tag exists in the registry;
|
||||||
|
- Helm server-side dry-run succeeds;
|
||||||
|
- ingress, certificate, and DNS prerequisites are present where applicable;
|
||||||
|
- platform dependencies are healthy or explicitly degraded with operator
|
||||||
|
approval;
|
||||||
|
- rollback target is known before the canary is applied;
|
||||||
|
- monitoring and log queries are available for the canary release.
|
||||||
|
|
||||||
|
Required Stage 2 evidence:
|
||||||
|
|
||||||
|
- rendered release identity;
|
||||||
|
- applied namespace and release name;
|
||||||
|
- pod readiness and restart status;
|
||||||
|
- ingress or routing state;
|
||||||
|
- key health endpoint result;
|
||||||
|
- relevant metrics window or explicit "metrics unavailable" note;
|
||||||
|
- State Hub progress note with non-secret evidence;
|
||||||
|
- operator approval when the workload is production-critical.
|
||||||
|
|
||||||
|
Canary acceptance gates:
|
||||||
|
|
||||||
|
- canary pods remain ready for the configured observation window;
|
||||||
|
- no crash loops, repeated restarts, or pending pods remain unexplained;
|
||||||
|
- health checks pass from inside and outside the cluster when both are
|
||||||
|
applicable;
|
||||||
|
- error rate, latency, and saturation do not regress beyond the workload's
|
||||||
|
declared threshold;
|
||||||
|
- no unexpected schema, storage, or queue side effects are observed;
|
||||||
|
- logs show no secret leakage and no repeated authorization failures;
|
||||||
|
- rollback has been tested previously or is a single documented command with a
|
||||||
|
known previous-stable target.
|
||||||
|
|
||||||
|
Default observation windows:
|
||||||
|
|
||||||
|
- non-critical internal service: 15 minutes;
|
||||||
|
- user-facing or shared platform service: 30 minutes;
|
||||||
|
- production-critical infrastructure such as Forgejo, identity, registry, or
|
||||||
|
State Hub: operator-defined window, minimum 60 minutes unless explicitly
|
||||||
|
waived.
|
||||||
|
|
||||||
|
Stage 2 fails closed when:
|
||||||
|
|
||||||
|
- the canary cannot be distinguished from stable production;
|
||||||
|
- production routing changes more traffic than intended;
|
||||||
|
- any required evidence is missing and no operator waiver is recorded;
|
||||||
|
- rollback target is unknown;
|
||||||
|
- the candidate needs a secret, credential, or platform dependency that was not
|
||||||
|
declared before the canary.
|
||||||
|
|
||||||
|
## Stage 3: Production Promotion
|
||||||
|
|
||||||
|
Stage 3 promotes the accepted candidate to the stable production path.
|
||||||
|
|
||||||
|
Promotion may mean:
|
||||||
|
|
||||||
|
- shifting weighted traffic to the canary release;
|
||||||
|
- replacing the stable Helm release with the accepted candidate;
|
||||||
|
- changing an ingress selector or service target;
|
||||||
|
- activating an operator-approved rollout workflow.
|
||||||
|
|
||||||
|
Required Stage 3 prechecks:
|
||||||
|
|
||||||
|
- Stage 2 acceptance gates passed for the same candidate artifact;
|
||||||
|
- the previous stable version is recorded;
|
||||||
|
- backup and restore posture is current for stateful workloads;
|
||||||
|
- migrations are approved and sequenced;
|
||||||
|
- production-critical workloads have explicit human approval;
|
||||||
|
- a rollback command and rollback verification check are available.
|
||||||
|
|
||||||
|
Required Stage 3 evidence:
|
||||||
|
|
||||||
|
- promotion command or workflow id;
|
||||||
|
- previous stable version;
|
||||||
|
- new stable version;
|
||||||
|
- production routing state after promotion;
|
||||||
|
- smoke result after promotion;
|
||||||
|
- rollback target retained;
|
||||||
|
- State Hub progress note with non-secret evidence.
|
||||||
|
|
||||||
|
Stage 3 is complete only after the post-promotion smoke passes and the
|
||||||
|
workload's stable routing points at the promoted candidate.
|
||||||
|
|
||||||
|
## Rollback Expectations
|
||||||
|
|
||||||
|
Rollback is part of every promotion, not an afterthought.
|
||||||
|
|
||||||
|
Every Stage 2 and Stage 3 action must identify one of:
|
||||||
|
|
||||||
|
- previous stable Helm release revision;
|
||||||
|
- previous image digest and values file;
|
||||||
|
- previous ingress/routing configuration;
|
||||||
|
- documented manual recovery path when automation is not yet safe.
|
||||||
|
|
||||||
|
Rollback must be immediate when:
|
||||||
|
|
||||||
|
- production availability is degraded;
|
||||||
|
- canary traffic escapes the declared blast radius;
|
||||||
|
- the workload emits repeated authorization or secret-handling errors;
|
||||||
|
- data integrity is at risk;
|
||||||
|
- an operator revokes approval during the observation window.
|
||||||
|
|
||||||
|
Rollback may be deferred only when the rollback itself is more dangerous than
|
||||||
|
the incident state. That decision requires a State Hub note and human approval.
|
||||||
|
|
||||||
|
After rollback, record:
|
||||||
|
|
||||||
|
- triggering symptom;
|
||||||
|
- rollback action;
|
||||||
|
- final stable version;
|
||||||
|
- remaining cleanup;
|
||||||
|
- whether the failed candidate is blocked, abandoned, or returned to Stage 1.
|
||||||
|
|
||||||
|
## Human Approval Gates
|
||||||
|
|
||||||
|
Human approval is required before production traffic changes for
|
||||||
|
production-critical workloads.
|
||||||
|
|
||||||
|
Production-critical workloads include:
|
||||||
|
|
||||||
|
- source forge and package registry workloads such as Forgejo or Gitea;
|
||||||
|
- identity, MFA, SSO, or authorization systems;
|
||||||
|
- State Hub, Inter-Hub, and operator coordination services;
|
||||||
|
- databases, object stores, and backup systems;
|
||||||
|
- ingress, certificate, or cluster-wide policy controllers;
|
||||||
|
- any workload whose failure blocks multiple repos or domains.
|
||||||
|
|
||||||
|
Approval must be recorded as a non-secret State Hub note or task comment. The
|
||||||
|
approval record should name:
|
||||||
|
|
||||||
|
- approving operator;
|
||||||
|
- candidate artifact;
|
||||||
|
- stage being approved;
|
||||||
|
- observation window;
|
||||||
|
- rollback target;
|
||||||
|
- any waived gates and why.
|
||||||
|
|
||||||
|
Emergency approval can be retrospective only when delaying the action would
|
||||||
|
increase production risk. Retrospective approval must be recorded immediately
|
||||||
|
after stabilization.
|
||||||
|
|
||||||
|
## Evidence And Secret Handling
|
||||||
|
|
||||||
|
Lifecycle evidence must be useful without being sensitive.
|
||||||
|
|
||||||
|
Allowed evidence:
|
||||||
|
|
||||||
|
- commit ids;
|
||||||
|
- image tags and digests;
|
||||||
|
- workflow ids;
|
||||||
|
- Kubernetes object names;
|
||||||
|
- pod status summaries;
|
||||||
|
- HTTP status codes;
|
||||||
|
- timestamps;
|
||||||
|
- State Hub progress ids;
|
||||||
|
- pass/fail summaries.
|
||||||
|
|
||||||
|
Forbidden evidence:
|
||||||
|
|
||||||
|
- plaintext secrets;
|
||||||
|
- bearer tokens;
|
||||||
|
- static API keys;
|
||||||
|
- kubeconfigs;
|
||||||
|
- private key material;
|
||||||
|
- full environment dumps;
|
||||||
|
- logs that contain credentials or user private data.
|
||||||
|
|
||||||
|
When a check needs secret-backed access, record only the access path and result,
|
||||||
|
for example: "OpenBao path configured, token exchange returned 200".
|
||||||
|
|
||||||
|
## Forgejo Readiness Interpretation
|
||||||
|
|
||||||
|
This lifecycle is clear enough for Forgejo when a future Forgejo workplan can
|
||||||
|
answer these questions before production cutover:
|
||||||
|
|
||||||
|
- What source revision and image digest are being promoted?
|
||||||
|
- What local checks prove the candidate is viable?
|
||||||
|
- How is the production canary isolated or traffic-limited?
|
||||||
|
- Which health, registry, SSH, web, Actions, and email recovery checks define
|
||||||
|
acceptance?
|
||||||
|
- Who approves the Stage 3 traffic switch?
|
||||||
|
- What is the previous stable target?
|
||||||
|
- How is repository data protected before and after promotion?
|
||||||
|
- How will rollback be verified without losing package or repository state?
|
||||||
|
|
||||||
|
If any answer is missing, Forgejo remains in Stage 1 or Stage 2 preparation and
|
||||||
|
must not cut over to Stage 3.
|
||||||
|
|
||||||
|
## Minimum Command Contract
|
||||||
|
|
||||||
|
The Railiance CLI makes these lifecycle operations repeatable:
|
||||||
|
|
||||||
|
```text
|
||||||
|
bin/railiance run <overlay-dir> # Stage 1 local validation
|
||||||
|
bin/railiance deploy --stage 2 <overlay-dir> --plan # Stage 2 canary plan
|
||||||
|
bin/railiance observe --stage 2 <overlay-dir> --plan # Stage 2 evidence targets
|
||||||
|
bin/railiance promote <overlay-dir> --plan # Stage 3 production promotion
|
||||||
|
bin/railiance rollback <overlay-dir> --plan # rollback to previous stable
|
||||||
|
```
|
||||||
|
|
||||||
|
The exact command names may change as implementation lands, but the behavior
|
||||||
|
must preserve the stage gates and evidence requirements in this document.
|
||||||
179
docs/overlay-repo-pattern.md
Normal file
179
docs/overlay-repo-pattern.md
Normal file
|
|
@ -0,0 +1,179 @@
|
||||||
|
# Railiance Overlay Repo Pattern
|
||||||
|
|
||||||
|
A Railiance overlay repo wraps a third-party upstream application without
|
||||||
|
forking Railiance deployment logic into the upstream source repository.
|
||||||
|
|
||||||
|
The overlay repo owns the Railiance deployment contract, promotion evidence,
|
||||||
|
Helm/Kubernetes overlay, values, probes, and runbooks. The upstream repository
|
||||||
|
remains the source of application code and release artifacts.
|
||||||
|
|
||||||
|
This pattern is a compatibility-era scaffolding model inside the
|
||||||
|
`rail-kubernetes` migration. Long-term first-class package ownership should
|
||||||
|
move toward `rapp-*` repos. Until the `rapp-*` path is fully materialized, the
|
||||||
|
overlay structure remains a useful generic pattern for Kubernetes-managed
|
||||||
|
workloads.
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
- Keep upstream code and Railiance deployment mechanics separate.
|
||||||
|
- Make Stage 1, Stage 2, Stage 3, and rollback behavior reproducible from Git.
|
||||||
|
- Declare all platform dependencies, health checks, and secret references in
|
||||||
|
`railiance/app.toml` without plaintext secret values.
|
||||||
|
- Allow third-party applications to adopt the staged promotion lifecycle without
|
||||||
|
requiring changes to their upstream repositories.
|
||||||
|
|
||||||
|
## Repo Layout
|
||||||
|
|
||||||
|
A generated compatibility overlay repo should look like this:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<app>-railiance-overlay/
|
||||||
|
README.md
|
||||||
|
railiance/
|
||||||
|
app.toml
|
||||||
|
upstream.toml
|
||||||
|
charts/
|
||||||
|
<app>/
|
||||||
|
Chart.yaml
|
||||||
|
values.yaml
|
||||||
|
templates/
|
||||||
|
deployment.yaml
|
||||||
|
service.yaml
|
||||||
|
values/
|
||||||
|
stage1.yaml
|
||||||
|
stage2-canary.yaml
|
||||||
|
stage3-production.yaml
|
||||||
|
patches/
|
||||||
|
upstream/.gitkeep
|
||||||
|
tests/
|
||||||
|
stage1.sh
|
||||||
|
runbooks/
|
||||||
|
rollback.md
|
||||||
|
docs/
|
||||||
|
promotion.md
|
||||||
|
```
|
||||||
|
|
||||||
|
When a workload graduates into a durable package repo, the preferred target is
|
||||||
|
to keep the same contract surfaces while moving the repo identity toward
|
||||||
|
`rapp-<workload>`.
|
||||||
|
|
||||||
|
## Ownership Boundary
|
||||||
|
|
||||||
|
The overlay repo owns:
|
||||||
|
|
||||||
|
- `railiance/app.toml` staged promotion declaration;
|
||||||
|
- Railiance-specific Helm chart, values, probes, and runbooks;
|
||||||
|
- canary and promotion evidence expectations;
|
||||||
|
- secret references by approved route and target object name;
|
||||||
|
- compatibility notes for a specific upstream revision or release line.
|
||||||
|
|
||||||
|
The upstream repo owns:
|
||||||
|
|
||||||
|
- application source code;
|
||||||
|
- upstream build and release artifacts;
|
||||||
|
- upstream tests and release notes;
|
||||||
|
- upstream vulnerability and license notices.
|
||||||
|
|
||||||
|
The overlay must not vendor upstream source by default. If a patch is required,
|
||||||
|
store the patch under `patches/upstream/` and record why the patch exists, when
|
||||||
|
it can be retired, and which upstream issue or release should replace it.
|
||||||
|
|
||||||
|
## Required Files
|
||||||
|
|
||||||
|
### `railiance/app.toml`
|
||||||
|
|
||||||
|
This file follows `docs/app-toml-contract.md` and
|
||||||
|
`schemas/railiance-app.schema.json` from `rail-kubernetes`. It is the primary
|
||||||
|
machine-readable contract for promotion tooling.
|
||||||
|
|
||||||
|
### `railiance/upstream.toml`
|
||||||
|
|
||||||
|
This file records non-secret upstream identity:
|
||||||
|
|
||||||
|
```toml
|
||||||
|
[upstream]
|
||||||
|
url = "https://example.com/vendor/app.git"
|
||||||
|
revision = "v1.2.3"
|
||||||
|
tracking = "tag"
|
||||||
|
license = "see-upstream"
|
||||||
|
notes = "Railiance overlay only; upstream code is not vendored here."
|
||||||
|
```
|
||||||
|
|
||||||
|
`revision` should be immutable where possible: tag, commit SHA, release id, or
|
||||||
|
image digest. Mutable branches are acceptable only before a workload becomes a
|
||||||
|
Stage 2 candidate.
|
||||||
|
|
||||||
|
### `charts/<app>/`
|
||||||
|
|
||||||
|
The chart is the Railiance deployment wrapper. It may start as a thin Helm
|
||||||
|
chart around an upstream image and grow only as required by the promotion gates.
|
||||||
|
Generated charts include stable/canary release identities, Prometheus-compatible
|
||||||
|
annotations, HTTP probes, resource limits, isolated canary ingress, and optional
|
||||||
|
Traefik weighted routing. Production-specific choices stay in `values/` files.
|
||||||
|
|
||||||
|
### `values/`
|
||||||
|
|
||||||
|
Stage values separate local validation, canary, and production settings:
|
||||||
|
|
||||||
|
- `stage1.yaml`: local or dry-run defaults;
|
||||||
|
- `stage2-canary.yaml`: limited exposure canary defaults;
|
||||||
|
- `stage3-production.yaml`: stable production defaults.
|
||||||
|
|
||||||
|
Secret values do not belong in these files. Use Kubernetes Secret,
|
||||||
|
ExternalSecret, OpenBao, KeyCape, or another approved route and record only the
|
||||||
|
reference name.
|
||||||
|
|
||||||
|
### `tests/stage1.sh` And `tests/stage2-template.sh`
|
||||||
|
|
||||||
|
Stage 1 should be runnable without production credentials. The generated script
|
||||||
|
performs syntax and Helm rendering checks when the relevant tools are available.
|
||||||
|
|
||||||
|
Stage 2 template validation verifies the canary scaffold, stable/canary values,
|
||||||
|
Prometheus annotations, rollback labels, and Helm rendering when Helm is
|
||||||
|
available. Workload-specific tests can extend either script.
|
||||||
|
|
||||||
|
### `runbooks/rollback.md`
|
||||||
|
|
||||||
|
Rollback instructions must exist before Stage 2. Early overlays may include a
|
||||||
|
placeholder, but it must name the intended rollback target and verification
|
||||||
|
check.
|
||||||
|
|
||||||
|
## Creation Tool
|
||||||
|
|
||||||
|
During the migration window, the compatibility scaffold still comes from:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
railiance-cluster/tools/create_railiance_overlay_repo.sh \
|
||||||
|
--app-id forgejo \
|
||||||
|
--name "Forgejo" \
|
||||||
|
--owner railiance-forge \
|
||||||
|
--criticality critical \
|
||||||
|
--upstream-url https://codeberg.org/forgejo/forgejo.git \
|
||||||
|
--upstream-revision v12.0.0 \
|
||||||
|
--out-dir /tmp/forgejo-railiance-overlay
|
||||||
|
```
|
||||||
|
|
||||||
|
The tool writes only local files. It does not call Gitea, clone upstream code,
|
||||||
|
fetch secrets, or push Git remotes.
|
||||||
|
|
||||||
|
Long-term, the generic scaffolding behavior should move under the rail and/or
|
||||||
|
become `rapp-*`-aware rather than preserving the overlay name forever.
|
||||||
|
|
||||||
|
## Promotion Use
|
||||||
|
|
||||||
|
1. Generate or update the overlay repo.
|
||||||
|
2. Fill in accurate image, namespace, health, dependency, and rollback fields.
|
||||||
|
3. Validate `railiance/app.toml` against the schema.
|
||||||
|
4. Run `tests/stage1.sh` and `tests/stage2-template.sh`.
|
||||||
|
5. Use later Stage 2 and Stage 3 commands to deploy, observe, promote, and
|
||||||
|
rollback.
|
||||||
|
|
||||||
|
## Safety Rules
|
||||||
|
|
||||||
|
- No plaintext secrets in `railiance/app.toml`, values files, tests, runbooks,
|
||||||
|
or generated evidence.
|
||||||
|
- Do not hide deployment logic in upstream source patches.
|
||||||
|
- Do not promote mutable upstream branches to Stage 2/3 without an explicit
|
||||||
|
operator exception.
|
||||||
|
- Production-critical overlays require human approval before canary exposure and
|
||||||
|
production promotion.
|
||||||
71
docs/promote-rollback-onboarding.md
Normal file
71
docs/promote-rollback-onboarding.md
Normal file
|
|
@ -0,0 +1,71 @@
|
||||||
|
# Promote, Rollback, And Onboarding
|
||||||
|
|
||||||
|
This guide shows the representative Railiance lifecycle for an overlay repo.
|
||||||
|
Commands default to plan mode so the path is repeatable before cluster access or
|
||||||
|
operator approval exists.
|
||||||
|
|
||||||
|
## Stage 1
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance run /path/to/overlay --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
Stage 1 validates `railiance/app.toml`, local commands, and local checks. Save
|
||||||
|
the JSON result as non-secret evidence before Stage 2.
|
||||||
|
|
||||||
|
## Stage 2
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance deploy --stage 2 /path/to/overlay --plan --pretty
|
||||||
|
bin/railiance observe --stage 2 /path/to/overlay --plan --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
When Helm, kubectl, cluster access, and approval evidence are ready:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance deploy --stage 2 /path/to/overlay --apply --approval-id <state-hub-id>
|
||||||
|
bin/railiance observe --stage 2 /path/to/overlay --live --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
For critical workloads, Stage 2 apply must not run until the operator has
|
||||||
|
approved canary exposure and rollback context is known.
|
||||||
|
|
||||||
|
## Stage 3
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance promote /path/to/overlay --plan --pretty
|
||||||
|
bin/railiance rollback /path/to/overlay --plan --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
Promotion plan mode emits a `railiance.stage3-promote-result.v1` JSON result
|
||||||
|
with stable release identity, chart and values paths, previous-stable target,
|
||||||
|
expected evidence, and approval requirements.
|
||||||
|
|
||||||
|
Rollback plan mode emits a `railiance.stage3-rollback-result.v1` JSON result
|
||||||
|
with rollback strategy, release identity, verification text, and apply-time
|
||||||
|
requirements.
|
||||||
|
|
||||||
|
When approval evidence and Helm access are ready:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance promote /path/to/overlay --apply --approval-id <state-hub-id>
|
||||||
|
bin/railiance rollback /path/to/overlay --apply --approval-id <state-hub-id> --revision <helm-revision>
|
||||||
|
```
|
||||||
|
|
||||||
|
Stage 3 apply fails closed if the chart or values are missing, previous stable
|
||||||
|
is not recorded, Helm is unavailable, or approval evidence is missing. Rollback
|
||||||
|
apply fails closed if the rollback strategy is missing, Helm is unavailable,
|
||||||
|
approval evidence is missing, or a Helm revision is required but absent.
|
||||||
|
|
||||||
|
## Human Approval Points
|
||||||
|
|
||||||
|
Critical infrastructure workloads require explicit operator approval before:
|
||||||
|
|
||||||
|
- Stage 2 canary exposure;
|
||||||
|
- Stage 3 stable promotion;
|
||||||
|
- rollback apply, unless an incident runbook defines a narrower break-glass
|
||||||
|
process and records the evidence id.
|
||||||
|
|
||||||
|
Progress notes should include only non-secret result summaries: schema version,
|
||||||
|
status, release, namespace, approval id, check counts, and command byte counts.
|
||||||
|
Do not paste command logs, kubeconfigs, tokens, or private service output.
|
||||||
52
docs/railiance-run-command.md
Normal file
52
docs/railiance-run-command.md
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
# Railiance Run Command
|
||||||
|
|
||||||
|
`bin/railiance run` executes Stage 1 local validation for a repository that
|
||||||
|
contains `railiance/app.toml`.
|
||||||
|
|
||||||
|
The command is intentionally local and conservative:
|
||||||
|
|
||||||
|
- reads `railiance/app.toml` using the `railiance.app.v1` contract;
|
||||||
|
- runs `[stages.stage1].commands` from the app directory;
|
||||||
|
- evaluates Stage 1 check ids listed in `[stages.stage1].checks` when they can
|
||||||
|
be checked locally;
|
||||||
|
- emits a machine-readable `railiance.run-result.v1` JSON result;
|
||||||
|
- records command references, exit codes, durations, and output byte counts,
|
||||||
|
but not shell text or command stdout/stderr content;
|
||||||
|
- strips credentials, query strings, and fragments from URLs before reporting HTTP
|
||||||
|
check results.
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance run /path/to/app-or-overlay --pretty
|
||||||
|
bin/railiance run . --json-out .railiance/stage1-result.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The process exits `0` only when all Stage 1 commands and required checks pass.
|
||||||
|
Optional checks may be skipped without failing the run. For example, an optional
|
||||||
|
local health endpoint can be declared before a local server command exists.
|
||||||
|
|
||||||
|
## Supported Local Checks
|
||||||
|
|
||||||
|
- `command`: runs the check `run` command in the app directory.
|
||||||
|
- `http`: calls the declared URL and compares the HTTP status.
|
||||||
|
- `helm`: runs `helm template` when Helm is installed. Required Helm checks fail
|
||||||
|
if Helm is unavailable; optional Helm checks are skipped.
|
||||||
|
|
||||||
|
Other check types are reported as skipped or failed depending on whether the
|
||||||
|
check is required. Stage 2 and Stage 3 checks are never executed by
|
||||||
|
`railiance run`.
|
||||||
|
|
||||||
|
## Result Shape
|
||||||
|
|
||||||
|
The JSON result includes:
|
||||||
|
|
||||||
|
- app identity and source revision;
|
||||||
|
- contract path and app directory;
|
||||||
|
- command/check status summaries using contract references instead of raw shell
|
||||||
|
commands;
|
||||||
|
- expected evidence labels from Stage 1;
|
||||||
|
- timing and exit status metadata.
|
||||||
|
|
||||||
|
The result is suitable for later promotion gates and State Hub progress notes,
|
||||||
|
without embedding secrets or verbose logs.
|
||||||
45
docs/source-import-plan.md
Normal file
45
docs/source-import-plan.md
Normal file
|
|
@ -0,0 +1,45 @@
|
||||||
|
# Source Import Plan
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Record which reviewed source materials are expected to seed this repo.
|
||||||
|
|
||||||
|
## Imported In The Current Wave
|
||||||
|
|
||||||
|
Imported source documents from `railiance-cluster`:
|
||||||
|
|
||||||
|
- `docs/deployment-lifecycle.md`
|
||||||
|
- `docs/app-toml-contract.md`
|
||||||
|
- `docs/overlay-repo-pattern.md`
|
||||||
|
- `docs/canary-helm-template.md`
|
||||||
|
- `docs/stage2-deploy-observe.md`
|
||||||
|
- `docs/promote-rollback-onboarding.md`
|
||||||
|
- `docs/railiance-run-command.md`
|
||||||
|
|
||||||
|
Imported source assets from `railiance-cluster`:
|
||||||
|
|
||||||
|
- `schemas/railiance-app.schema.json`
|
||||||
|
- `examples/railiance/app.toml`
|
||||||
|
|
||||||
|
## Pending First-Wave Imports
|
||||||
|
|
||||||
|
Expected source assets from `railiance-cluster`:
|
||||||
|
|
||||||
|
- `tools/create_railiance_overlay_repo.sh`
|
||||||
|
- `tools/cmd/railiance-run`
|
||||||
|
- `tools/cmd/railiance-stage2`
|
||||||
|
- `tools/cmd/railiance-stage3`
|
||||||
|
|
||||||
|
## Deferred Migration Debt
|
||||||
|
|
||||||
|
Do not import these into `rail-kubernetes` as part of wave 1:
|
||||||
|
|
||||||
|
- workload-specific activity-core helper commands
|
||||||
|
- workload-specific smoke or reconciliation helpers
|
||||||
|
- cluster substrate operations
|
||||||
|
- shared bootstrap helpers that still need a separate placement review
|
||||||
|
|
||||||
|
## Rule
|
||||||
|
|
||||||
|
Imports into this repo should preserve the boundary already reviewed in
|
||||||
|
`railiance-master` and `railiance-cluster`.
|
||||||
49
docs/stage2-deploy-observe.md
Normal file
49
docs/stage2-deploy-observe.md
Normal file
|
|
@ -0,0 +1,49 @@
|
||||||
|
# Stage 2 Deploy And Observe
|
||||||
|
|
||||||
|
`bin/railiance deploy --stage 2` and `bin/railiance observe --stage 2` provide
|
||||||
|
the repeatable command path for production canaries declared in
|
||||||
|
`railiance/app.toml`.
|
||||||
|
|
||||||
|
Both commands default to non-mutating plan mode.
|
||||||
|
|
||||||
|
## Deploy
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance deploy --stage 2 /path/to/overlay --pretty
|
||||||
|
bin/railiance deploy --stage 2 /path/to/overlay --server-dry-run --pretty
|
||||||
|
bin/railiance deploy --stage 2 /path/to/overlay --apply --approval-id <state-hub-id>
|
||||||
|
```
|
||||||
|
|
||||||
|
Plan mode validates the local Stage 2 chart and values paths and emits a
|
||||||
|
`railiance.stage2-deploy-result.v1` JSON plan. It does not contact the cluster.
|
||||||
|
|
||||||
|
`--server-dry-run` runs `helm upgrade --install --dry-run=server` when Helm and
|
||||||
|
cluster access are available. `--apply` runs the Helm canary apply path with
|
||||||
|
`--atomic --wait`. If Stage 2 declares `requires_approval = true`, apply mode
|
||||||
|
fails closed unless `--approval-id` is provided.
|
||||||
|
|
||||||
|
The result records release identity, namespace, chart path, values path,
|
||||||
|
expected checks/evidence, precheck status, and command byte counts. It does not
|
||||||
|
embed Helm or kubectl logs.
|
||||||
|
|
||||||
|
## Observe
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/railiance observe --stage 2 /path/to/overlay --pretty
|
||||||
|
bin/railiance observe --stage 2 /path/to/overlay --live --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
Plan mode emits the rollout, pod selector, ingress selector, health URL, and
|
||||||
|
metrics targets that live observation will query.
|
||||||
|
|
||||||
|
Live mode uses `kubectl` to check rollout status, deployment JSON, canary pods,
|
||||||
|
ingress/routing resources, and pod metrics when metrics-server is available.
|
||||||
|
Metrics unavailability is reported separately so a canary can fail for rollout
|
||||||
|
or readiness problems without hiding missing observability.
|
||||||
|
|
||||||
|
## Safety
|
||||||
|
|
||||||
|
Stage 2 remains blocked when required local paths are missing, Helm is missing
|
||||||
|
for dry-run/apply, `kubectl` is missing for live observe, or approval evidence
|
||||||
|
is missing for an apply that requires approval. Use the emitted JSON as
|
||||||
|
non-secret evidence in State Hub progress notes.
|
||||||
59
docs/wave-1-contract.md
Normal file
59
docs/wave-1-contract.md
Normal file
|
|
@ -0,0 +1,59 @@
|
||||||
|
# Wave-1 Kubernetes Rail Contract
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Summarize the initial contract `rail-kubernetes` owns in wave 1.
|
||||||
|
|
||||||
|
This repo starts from already-reviewed source documents rather than inventing a
|
||||||
|
new runtime model:
|
||||||
|
|
||||||
|
- `railiance-master/docs/rail-kubernetes-boundary.md`
|
||||||
|
- `railiance-cluster/docs/rail-kubernetes-extraction-map.md`
|
||||||
|
- `railiance-cluster/docs/rail-kubernetes-substrate-profile.md`
|
||||||
|
|
||||||
|
## What `rail-kubernetes` Owns
|
||||||
|
|
||||||
|
Wave 1 assigns these concerns to the rail:
|
||||||
|
|
||||||
|
- the generic workload-on-Kubernetes execution contract
|
||||||
|
- the generic `railiance/app.toml` workload declaration surface
|
||||||
|
- generic stage 1 / stage 2 / stage 3 lifecycle semantics
|
||||||
|
- generic canary, observe, promote, and rollback behavior
|
||||||
|
- generic wrapper and overlay guidance for workloads that still use the current
|
||||||
|
overlay pattern during migration
|
||||||
|
|
||||||
|
## What Stays In `railiance-cluster`
|
||||||
|
|
||||||
|
Wave 1 keeps these concerns in the S2 ownership repo:
|
||||||
|
|
||||||
|
- k3s bootstrap and host-to-cluster installation
|
||||||
|
- cluster backup, restore, smoke, and preflight operations
|
||||||
|
- operator ownership for ingress, TLS, and other cluster-level components
|
||||||
|
- cluster access delivery and kubeconfig handling
|
||||||
|
- substrate-level safety rails
|
||||||
|
|
||||||
|
## Published Substrate Assumptions
|
||||||
|
|
||||||
|
Wave 1 `rail-kubernetes` may assume:
|
||||||
|
|
||||||
|
- a Kubernetes API backed by k3s
|
||||||
|
- Helm-based apply paths
|
||||||
|
- operator-driven kubeconfig access
|
||||||
|
- isolated canary exposure as the safe default
|
||||||
|
- namespace creation during workload rollout
|
||||||
|
|
||||||
|
Wave 1 `rail-kubernetes` must not assume:
|
||||||
|
|
||||||
|
- guaranteed weighted routing
|
||||||
|
- a generic workload RBAC model
|
||||||
|
- a generic persistent-storage contract
|
||||||
|
- a secret-delivery contract at the S2 layer
|
||||||
|
|
||||||
|
## Migration Constraint
|
||||||
|
|
||||||
|
Until the extraction is complete, `railiance-cluster` remains the compatibility
|
||||||
|
surface for current operators.
|
||||||
|
|
||||||
|
`rail-kubernetes` should become the durable home of the contract first. Only
|
||||||
|
then should commands, docs, schemas, and helpers move here in a migration-safe
|
||||||
|
sequence.
|
||||||
176
examples/railiance/app.toml
Normal file
176
examples/railiance/app.toml
Normal file
|
|
@ -0,0 +1,176 @@
|
||||||
|
schema_version = "railiance.app.v1"
|
||||||
|
|
||||||
|
[app]
|
||||||
|
id = "example-service"
|
||||||
|
name = "Example Service"
|
||||||
|
repo = "railiance-apps/example-service"
|
||||||
|
owner = "platform"
|
||||||
|
criticality = "critical"
|
||||||
|
description = "Reference declaration for the Railiance staged promotion lifecycle."
|
||||||
|
|
||||||
|
[source]
|
||||||
|
revision = "git:main"
|
||||||
|
artifact = "image"
|
||||||
|
digest_policy = "required"
|
||||||
|
|
||||||
|
[rollback]
|
||||||
|
strategy = "helm-revision"
|
||||||
|
command = "bin/railiance rollback example-service"
|
||||||
|
verification = "GET /health returns 200 on the restored stable release."
|
||||||
|
|
||||||
|
[[platform.dependencies]]
|
||||||
|
name = "state-hub"
|
||||||
|
kind = "state-hub"
|
||||||
|
required = true
|
||||||
|
stage = "stage2"
|
||||||
|
evidence = "State Hub /healthz returns ok from the cluster path."
|
||||||
|
|
||||||
|
[[platform.dependencies]]
|
||||||
|
name = "postgres"
|
||||||
|
kind = "postgres"
|
||||||
|
required = true
|
||||||
|
stage = "stage2"
|
||||||
|
evidence = "Target database reports Ready and backup posture is current."
|
||||||
|
|
||||||
|
[[secrets.references]]
|
||||||
|
name = "runtime-api-key"
|
||||||
|
route = "openbao-api-key"
|
||||||
|
target = "ExternalSecret/example-service-runtime"
|
||||||
|
stage = "stage2"
|
||||||
|
required = true
|
||||||
|
|
||||||
|
[[observability.health_endpoints]]
|
||||||
|
name = "local-health"
|
||||||
|
url = "http://127.0.0.1:8080/health"
|
||||||
|
stage = "stage1"
|
||||||
|
expected_status = 200
|
||||||
|
|
||||||
|
[[observability.health_endpoints]]
|
||||||
|
name = "cluster-health"
|
||||||
|
url = "http://example-service.example-service.svc.cluster.local:8080/health"
|
||||||
|
stage = "stage2"
|
||||||
|
expected_status = 200
|
||||||
|
|
||||||
|
[[observability.metrics]]
|
||||||
|
name = "request-errors"
|
||||||
|
reference = 'promql:rate(http_requests_total{status=~"5.."}[5m])'
|
||||||
|
stage = "stage2"
|
||||||
|
|
||||||
|
[[observability.logs]]
|
||||||
|
name = "secret-leak-scan"
|
||||||
|
reference = "kubectl logs -n example-service deploy/example-service-canary"
|
||||||
|
stage = "stage2"
|
||||||
|
|
||||||
|
[stages.stage1]
|
||||||
|
enabled = true
|
||||||
|
namespace = "local"
|
||||||
|
release = "example-service-local"
|
||||||
|
commands = ["make test", "helm template charts/example-service"]
|
||||||
|
checks = ["unit-tests", "helm-template", "local-health"]
|
||||||
|
evidence = ["pytest output", "helm template success", "local health 200"]
|
||||||
|
requires_approval = false
|
||||||
|
|
||||||
|
[stages.stage2]
|
||||||
|
enabled = true
|
||||||
|
namespace = "example-service"
|
||||||
|
release = "example-service-canary"
|
||||||
|
commands = ["bin/railiance deploy --stage 2 example-service", "bin/railiance observe example-service"]
|
||||||
|
checks = ["server-dry-run", "canary-ready", "cluster-health", "operator-approval"]
|
||||||
|
evidence = ["release name", "pod readiness", "health 200", "State Hub progress id"]
|
||||||
|
requires_approval = true
|
||||||
|
canary_mode = "isolated"
|
||||||
|
observation_minutes = 60
|
||||||
|
|
||||||
|
[stages.stage3]
|
||||||
|
enabled = true
|
||||||
|
namespace = "example-service"
|
||||||
|
release = "example-service"
|
||||||
|
commands = ["bin/railiance promote example-service", "bin/railiance observe example-service"]
|
||||||
|
checks = ["stage2-accepted", "rollback-target", "cluster-health", "operator-approval"]
|
||||||
|
evidence = ["promotion command id", "new stable digest", "post-promotion smoke"]
|
||||||
|
requires_approval = true
|
||||||
|
promotion_mode = "release-replace"
|
||||||
|
previous_stable = "helm:example-service:previous"
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "unit-tests"
|
||||||
|
type = "command"
|
||||||
|
stage = "stage1"
|
||||||
|
description = "Run repository unit tests."
|
||||||
|
required = true
|
||||||
|
run = "make test"
|
||||||
|
timeout_seconds = 600
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "helm-template"
|
||||||
|
type = "helm"
|
||||||
|
stage = "stage1"
|
||||||
|
description = "Render Helm templates locally."
|
||||||
|
required = true
|
||||||
|
chart = "charts/example-service"
|
||||||
|
values = "values/local.yaml"
|
||||||
|
mode = "template"
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "local-health"
|
||||||
|
type = "http"
|
||||||
|
stage = "stage1"
|
||||||
|
description = "Confirm local service health."
|
||||||
|
required = true
|
||||||
|
url = "http://127.0.0.1:8080/health"
|
||||||
|
expected_status = 200
|
||||||
|
timeout_seconds = 10
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "server-dry-run"
|
||||||
|
type = "helm"
|
||||||
|
stage = "stage2"
|
||||||
|
description = "Render and submit a server-side dry run before canary."
|
||||||
|
required = true
|
||||||
|
chart = "charts/example-service"
|
||||||
|
values = "values/canary.yaml"
|
||||||
|
mode = "server-dry-run"
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "canary-ready"
|
||||||
|
type = "kubernetes"
|
||||||
|
stage = "stage2"
|
||||||
|
description = "Canary deployment reaches Available."
|
||||||
|
required = true
|
||||||
|
namespace = "example-service"
|
||||||
|
resource = "deploy/example-service-canary"
|
||||||
|
condition = "Available"
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "cluster-health"
|
||||||
|
type = "http"
|
||||||
|
stage = "stage2"
|
||||||
|
description = "Cluster health endpoint returns 200."
|
||||||
|
required = true
|
||||||
|
url = "http://example-service.example-service.svc.cluster.local:8080/health"
|
||||||
|
expected_status = 200
|
||||||
|
timeout_seconds = 10
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "operator-approval"
|
||||||
|
type = "manual"
|
||||||
|
stage = "stage2"
|
||||||
|
description = "Human approval is recorded before production-critical traffic changes."
|
||||||
|
required = true
|
||||||
|
evidence_required = "State Hub approval note id, candidate digest, rollback target."
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "stage2-accepted"
|
||||||
|
type = "manual"
|
||||||
|
stage = "stage3"
|
||||||
|
description = "Stage 2 gates passed for the same candidate artifact."
|
||||||
|
required = true
|
||||||
|
evidence_required = "State Hub Stage 2 acceptance progress id."
|
||||||
|
|
||||||
|
[[checks]]
|
||||||
|
id = "rollback-target"
|
||||||
|
type = "manual"
|
||||||
|
stage = "stage3"
|
||||||
|
description = "Previous stable release is recorded before promotion."
|
||||||
|
required = true
|
||||||
|
evidence_required = "Previous Helm revision or image digest."
|
||||||
596
schemas/railiance-app.schema.json
Normal file
596
schemas/railiance-app.schema.json
Normal file
|
|
@ -0,0 +1,596 @@
|
||||||
|
{
|
||||||
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||||
|
"$id": "https://railiance.local/schemas/railiance-app.schema.json",
|
||||||
|
"title": "Railiance app.toml contract",
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"schema_version",
|
||||||
|
"app",
|
||||||
|
"source",
|
||||||
|
"platform",
|
||||||
|
"secrets",
|
||||||
|
"observability",
|
||||||
|
"rollback",
|
||||||
|
"stages",
|
||||||
|
"checks"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"schema_version": {
|
||||||
|
"const": "railiance.app.v1"
|
||||||
|
},
|
||||||
|
"app": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"id",
|
||||||
|
"name",
|
||||||
|
"repo",
|
||||||
|
"owner",
|
||||||
|
"criticality",
|
||||||
|
"description"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"id": {
|
||||||
|
"type": "string",
|
||||||
|
"pattern": "^[a-z0-9][a-z0-9-]*$"
|
||||||
|
},
|
||||||
|
"name": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"repo": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"owner": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"criticality": {
|
||||||
|
"enum": [
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high",
|
||||||
|
"critical"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"description": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"source": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"revision",
|
||||||
|
"artifact",
|
||||||
|
"digest_policy"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"revision": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"artifact": {
|
||||||
|
"enum": [
|
||||||
|
"image",
|
||||||
|
"helm-chart",
|
||||||
|
"bundle",
|
||||||
|
"manifest",
|
||||||
|
"other"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"digest_policy": {
|
||||||
|
"enum": [
|
||||||
|
"required",
|
||||||
|
"preferred",
|
||||||
|
"not-applicable"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"platform": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"dependencies"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"dependencies": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/dependency"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"secrets": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"references"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"references": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/secretReference"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"observability": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"health_endpoints"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"health_endpoints": {
|
||||||
|
"type": "array",
|
||||||
|
"minItems": 1,
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/healthEndpoint"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"metrics": {
|
||||||
|
"type": "array",
|
||||||
|
"default": [],
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/observationReference"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"logs": {
|
||||||
|
"type": "array",
|
||||||
|
"default": [],
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/observationReference"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"rollback": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"strategy",
|
||||||
|
"command",
|
||||||
|
"verification"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"strategy": {
|
||||||
|
"enum": [
|
||||||
|
"helm-revision",
|
||||||
|
"image-digest",
|
||||||
|
"traffic-shift",
|
||||||
|
"manual-runbook",
|
||||||
|
"none"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"command": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"verification": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"stages": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"stage1",
|
||||||
|
"stage2",
|
||||||
|
"stage3"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"stage1": {
|
||||||
|
"$ref": "#/$defs/stage1"
|
||||||
|
},
|
||||||
|
"stage2": {
|
||||||
|
"$ref": "#/$defs/stage2"
|
||||||
|
},
|
||||||
|
"stage3": {
|
||||||
|
"$ref": "#/$defs/stage3"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"checks": {
|
||||||
|
"type": "array",
|
||||||
|
"minItems": 1,
|
||||||
|
"items": {
|
||||||
|
"$ref": "#/$defs/check"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"$defs": {
|
||||||
|
"stageName": {
|
||||||
|
"enum": [
|
||||||
|
"stage1",
|
||||||
|
"stage2",
|
||||||
|
"stage3"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"dependency": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"name",
|
||||||
|
"kind",
|
||||||
|
"required",
|
||||||
|
"stage",
|
||||||
|
"evidence"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"name": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"kind": {
|
||||||
|
"enum": [
|
||||||
|
"postgres",
|
||||||
|
"redis",
|
||||||
|
"object-store",
|
||||||
|
"identity",
|
||||||
|
"state-hub",
|
||||||
|
"inter-hub",
|
||||||
|
"network",
|
||||||
|
"other"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"required": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"stage": {
|
||||||
|
"$ref": "#/$defs/stageName"
|
||||||
|
},
|
||||||
|
"evidence": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"secretReference": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"name",
|
||||||
|
"route",
|
||||||
|
"target",
|
||||||
|
"stage",
|
||||||
|
"required"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"name": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"route": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"target": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"stage": {
|
||||||
|
"$ref": "#/$defs/stageName"
|
||||||
|
},
|
||||||
|
"required": {
|
||||||
|
"type": "boolean"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"not": {
|
||||||
|
"anyOf": [
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"value"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"token"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"password"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"secret"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"private_key"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"required": [
|
||||||
|
"kubeconfig"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"healthEndpoint": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"name",
|
||||||
|
"url",
|
||||||
|
"stage",
|
||||||
|
"expected_status"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"name": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"url": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"stage": {
|
||||||
|
"$ref": "#/$defs/stageName"
|
||||||
|
},
|
||||||
|
"expected_status": {
|
||||||
|
"type": "integer",
|
||||||
|
"minimum": 100,
|
||||||
|
"maximum": 599
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"observationReference": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"name",
|
||||||
|
"reference",
|
||||||
|
"stage"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"name": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"reference": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"stage": {
|
||||||
|
"$ref": "#/$defs/stageName"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"check": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": true,
|
||||||
|
"required": [
|
||||||
|
"id",
|
||||||
|
"type",
|
||||||
|
"stage",
|
||||||
|
"description",
|
||||||
|
"required"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"id": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"type": {
|
||||||
|
"enum": [
|
||||||
|
"command",
|
||||||
|
"http",
|
||||||
|
"kubernetes",
|
||||||
|
"helm",
|
||||||
|
"metric",
|
||||||
|
"log",
|
||||||
|
"manual"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"stage": {
|
||||||
|
"$ref": "#/$defs/stageName"
|
||||||
|
},
|
||||||
|
"description": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"required": {
|
||||||
|
"type": "boolean"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"stage1": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"enabled",
|
||||||
|
"namespace",
|
||||||
|
"release",
|
||||||
|
"commands",
|
||||||
|
"checks",
|
||||||
|
"evidence",
|
||||||
|
"requires_approval"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"enabled": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"namespace": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"release": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"commands": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"checks": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"evidence": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"requires_approval": {
|
||||||
|
"type": "boolean"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"stage2": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"enabled",
|
||||||
|
"namespace",
|
||||||
|
"release",
|
||||||
|
"commands",
|
||||||
|
"checks",
|
||||||
|
"evidence",
|
||||||
|
"requires_approval",
|
||||||
|
"canary_mode",
|
||||||
|
"observation_minutes"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"enabled": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"namespace": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"release": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"commands": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"checks": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"evidence": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"requires_approval": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"canary_mode": {
|
||||||
|
"enum": [
|
||||||
|
"weighted",
|
||||||
|
"header",
|
||||||
|
"path",
|
||||||
|
"shadow",
|
||||||
|
"isolated"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"observation_minutes": {
|
||||||
|
"type": "integer",
|
||||||
|
"minimum": 1
|
||||||
|
},
|
||||||
|
"traffic_percent": {
|
||||||
|
"type": "integer",
|
||||||
|
"minimum": 0,
|
||||||
|
"maximum": 100
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"stage3": {
|
||||||
|
"type": "object",
|
||||||
|
"additionalProperties": false,
|
||||||
|
"required": [
|
||||||
|
"enabled",
|
||||||
|
"namespace",
|
||||||
|
"release",
|
||||||
|
"commands",
|
||||||
|
"checks",
|
||||||
|
"evidence",
|
||||||
|
"requires_approval",
|
||||||
|
"promotion_mode",
|
||||||
|
"previous_stable"
|
||||||
|
],
|
||||||
|
"properties": {
|
||||||
|
"enabled": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"namespace": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"release": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
},
|
||||||
|
"commands": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"checks": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"evidence": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"requires_approval": {
|
||||||
|
"type": "boolean"
|
||||||
|
},
|
||||||
|
"promotion_mode": {
|
||||||
|
"enum": [
|
||||||
|
"traffic-shift",
|
||||||
|
"release-replace",
|
||||||
|
"selector-switch",
|
||||||
|
"workflow"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"previous_stable": {
|
||||||
|
"type": "string",
|
||||||
|
"minLength": 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
129
workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md
Normal file
129
workplans/RAIL-K8S-WP-0001-bootstrap-and-wave1-import.md
Normal file
|
|
@ -0,0 +1,129 @@
|
||||||
|
---
|
||||||
|
id: RAIL-K8S-WP-0001
|
||||||
|
type: workplan
|
||||||
|
title: "Bootstrap rail-kubernetes and import the wave-1 contract"
|
||||||
|
domain: financials
|
||||||
|
repo: rail-kubernetes
|
||||||
|
status: active
|
||||||
|
owner: codex
|
||||||
|
topic_slug: railiance
|
||||||
|
created: "2026-07-25"
|
||||||
|
updated: "2026-07-25"
|
||||||
|
state_hub_workstream_id: "1f2254f5-a873-46fc-a77c-bc725eeaaeb8"
|
||||||
|
---
|
||||||
|
|
||||||
|
# RAIL-K8S-WP-0001 - Bootstrap rail-kubernetes and import the wave-1 contract
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Turn `rail-kubernetes` into the concrete home of the default Railiance
|
||||||
|
execution rail.
|
||||||
|
|
||||||
|
## Boundaries
|
||||||
|
|
||||||
|
This workplan may bootstrap the repo, publish the initial declaration, and
|
||||||
|
import the generic wave-1 contract.
|
||||||
|
|
||||||
|
It must not:
|
||||||
|
|
||||||
|
- absorb S2 substrate ownership from `railiance-cluster`
|
||||||
|
- normalize workload-specific helper commands into the rail
|
||||||
|
- pretend the migration is complete before compatibility shims exist
|
||||||
|
|
||||||
|
## T01 - Establish the repo baseline and rail declaration
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: RAIL-K8S-WP-0001-T01
|
||||||
|
status: done
|
||||||
|
priority: high
|
||||||
|
state_hub_task_id: "9bab36db-e094-4e41-af76-b6bf5e73be78"
|
||||||
|
```
|
||||||
|
|
||||||
|
Create the baseline repo files, the initial rail declaration, and the first
|
||||||
|
wave-1 contract summary.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- baseline repo files exist
|
||||||
|
- `declarations/rail.yaml` exists
|
||||||
|
- the repo states its boundary with `railiance-cluster`
|
||||||
|
|
||||||
|
2026-07-25: Created the initial repo baseline, declaration, and wave-1
|
||||||
|
contract summary.
|
||||||
|
|
||||||
|
## T02 - Import the reviewed generic contract material
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: RAIL-K8S-WP-0001-T02
|
||||||
|
status: progress
|
||||||
|
priority: high
|
||||||
|
state_hub_task_id: "148276a2-2e40-4d56-89dd-2930c5e4309c"
|
||||||
|
```
|
||||||
|
|
||||||
|
Bring the reviewed generic lifecycle docs, schema, examples, and helper
|
||||||
|
surfaces into this repo in a migration-safe sequence.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- the import sequence is anchored in the reviewed extraction map
|
||||||
|
- imported material stays generic to the Kubernetes rail
|
||||||
|
- the repo becomes a real home for the default rail rather than only a shell
|
||||||
|
|
||||||
|
2026-07-25: Established the repo-local contract summary and explicit source
|
||||||
|
import plan. Full import of docs, schema, examples, and helper surfaces is
|
||||||
|
still pending.
|
||||||
|
|
||||||
|
2026-07-25: Imported the first core generic contract artifacts from
|
||||||
|
`railiance-cluster`: `docs/deployment-lifecycle.md`,
|
||||||
|
`docs/app-toml-contract.md`, `schemas/railiance-app.schema.json`, and
|
||||||
|
`examples/railiance/app.toml`. Helper command surfaces and the remaining
|
||||||
|
supporting docs are still pending.
|
||||||
|
|
||||||
|
2026-07-25: Imported the remaining generic supporting docs from
|
||||||
|
`railiance-cluster`: `docs/overlay-repo-pattern.md`,
|
||||||
|
`docs/canary-helm-template.md`, `docs/stage2-deploy-observe.md`,
|
||||||
|
`docs/promote-rollback-onboarding.md`, and
|
||||||
|
`docs/railiance-run-command.md`. Helper scripts and command implementations
|
||||||
|
remain pending.
|
||||||
|
|
||||||
|
## T03 - Prepare the compatibility handoff from `railiance-cluster`
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: RAIL-K8S-WP-0001-T03
|
||||||
|
status: wait
|
||||||
|
priority: medium
|
||||||
|
state_hub_task_id: "18140d90-d10e-47fe-b7da-d454550466b1"
|
||||||
|
```
|
||||||
|
|
||||||
|
Define the thin compatibility path that keeps current operators functional
|
||||||
|
while the extraction proceeds.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- the compatibility path is explicit
|
||||||
|
- current operators do not need a flag day
|
||||||
|
- the retained S2 surface stays narrow and temporary
|
||||||
|
|
||||||
|
## T04 - Register `rail-kubernetes` in State Hub and Fabric
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: RAIL-K8S-WP-0001-T04
|
||||||
|
status: progress
|
||||||
|
priority: medium
|
||||||
|
state_hub_task_id: "d78b5712-f305-4967-9034-9b9471633755"
|
||||||
|
```
|
||||||
|
|
||||||
|
Register the repo from classification and add it to the Fabric repository
|
||||||
|
registry with rail-family metadata.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- State Hub recognizes the repo and its workplan
|
||||||
|
- Fabric recognizes the repo as `repo_family: rail`
|
||||||
|
- ownership and primary contract metadata are visible
|
||||||
|
|
||||||
|
2026-07-25: Registered `rail-kubernetes` in State Hub and generated
|
||||||
|
`WORK-RECORDS.md` plus `.custodian-brief.md`. Added `rail-kubernetes` to
|
||||||
|
`railiance-fabric/registry/railiance-repos.yaml` as `repo_family: rail` with
|
||||||
|
`ownership_repo: railiance-cluster`. Live Fabric registry sync and Forgejo
|
||||||
|
remote provisioning remain pending.
|
||||||
1
workplans/archived/.gitkeep
Normal file
1
workplans/archived/.gitkeep
Normal file
|
|
@ -0,0 +1 @@
|
||||||
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue