resource-control/docs/evidence-basis.md
tegwick 13c2b82281 feat: restate the backup case in ITC-CAP terms; add evidence basis; publish consumption-mode signal
Three things.

1. CANON RESTATEMENT (info-tech-canon's ask after accepting our demand)
data/capability/platform-audit-storage.json restates the backup case against
ITC-CAP 0.2.0: requirement with profile, targets and the failure-domain
constraint that decided the procurement; two provisions (data.object and
data.backup); all four data.backup evidence hooks satisfied and measured; and
consumption in native units — GB, hours, tokens — with unknown never zero.

tools/capability.py reads their capabilities.yaml directly rather than copying
it, so drift in either repo fails here. The requirement asks D5, the provision
is D4, and the review reports below_requirement rather than inflating maturity.

2. EVIDENCE BASIS (tools/basis.py, docs/evidence-basis.md)
Every value declares how it was obtained on an ordered scale: invoiced,
measured, quoted, derived, projected, estimated, assumed, unknown. A derived
value resolves to the weakest basis among its inputs, so precise arithmetic
cannot launder weak assumptions.

First application is a finding about our own biggest decision: the Scaleway vs
Hetzner comparison, EUR 29.14/month stated to the cent, grades "indicative" —
1 of 4 load-bearing values evidenced, weakest "assumed". The direction is
robust; the magnitude is a model output. The cheapest fix is recording real
operator hours, not better arithmetic.

3. CONSUMPTION-MODE SIGNAL (railiance-platform RAILIANCE-WP-0017)
settlement.py gains a consumption-mode command projecting statements into the
signal they consume; make consumption-mode PERIOD=YYYY-MM publishes
data/consumption-mode/current.json. Currently an empty list: no live charges for
2026-09, so no entity is restricted. Publishing the empty list makes that an
assertion rather than an absence, which their contract distinguishes. The
validator fails if the published signal is stale.

185 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 18:42:06 +02:00

103 lines
4.6 KiB
Markdown

# Evidence basis — where we know, and where we guess
A counted object and an assumed hourly rate are both numbers. They are not both
knowledge. Every quantity in this repository declares **how it was obtained**,
so a decision can be graded by the weakest thing it actually rests on.
Vocabulary and propagation: `tools/basis.py`.
## The scale
Ordered strongest to weakest. The order is the point — it is what makes
"weakest input wins" computable.
| Basis | Meaning | Example here |
|---|---|---|
| `invoiced` | A booked financial fact, authoritative from fin-hub | none yet |
| `measured` | Directly observed from the authoritative system | 8 objects, 6 604 031 B in the backup prefix |
| `quoted` | Stated by a provider or counterparty in a citable source | Scaleway €0.01606/GB-month |
| `derived` | Computed from other values by a stated rule | €7.35/month infrastructure |
| `projected` | Interpolated between, or extrapolated beyond, observations | 457.968 GB stored at month 12 |
| `estimated` | Human judgement, neither observed nor computed | apps-pg 6 operator-hours setup |
| `assumed` | A modelling constant we chose | €60/hour operator rate |
| `unknown` | No value exists | railiance01 monthly price |
`invoiced`, `measured`, and `quoted` are **evidenced**: they assert an observed
or contracted fact. Everything below them is inference.
## The propagation rule
> A derived value is only as strong as its weakest input.
Without this rule a precise-looking figure launders weak assumptions. The
monthly labour line for the backup store is €60.00, which reads like a
measurement. It is `derived` from an `assumed` operator-hour allowance and an
`assumed` hourly rate, so it resolves to **`assumed`** — the weakest class of
input we have.
Deriving in the other direction is not penalised: GB computed from measured
bytes resolves to `measured`, because both inputs are.
## Two orthogonal qualifiers
**`proxy_for`** — the value is a genuine measurement of a *different* quantity
than the one being reasoned about. It does not weaken the measurement; it
weakens the inference. The apps-pg allocation driver measures database bytes
accurately, and one of those databases has never recorded a transaction, so as
a driver for cost it is a proxy whose validity is separate from its precision.
**`observed_at`** — a measurement has an age. A single sample from a
four-hour-old resource is `measured` and still thin.
## Unknown is never zero
A value with basis `unknown` carries no quantity and must name the gap and its
owner. A value with any other basis must carry a quantity. Zero means "measured
and none"; unknown means "not measured". The validator enforces both directions.
## Decision grades
`decision_grade(values)` reports the distribution, the weakest basis, the
evidenced ratio, and a grade:
| Grade | Condition |
|---|---|
| `evidenced` | Every load-bearing value is invoiced, measured, or quoted |
| `projected` | The conclusion rests on values projected from observations |
| `indicative` | The conclusion is no stronger than an estimated or assumed value |
| `insufficient` | At least one load-bearing value is unknown |
### The first real application
The provider comparison that selected Scaleway over Hetzner — a €29.14/month
difference stated to the cent — grades **`indicative`**:
```
by_basis: {quoted: 1, projected: 1, assumed: 2}
weakest: assumed
evidenced_ratio: 0.25
```
One of four load-bearing values is evidenced. The two labour figures, which are
what actually inverted the ranking, are both `assumed` on top of `assumed`.
This does not overturn the decision. The direction is robust — Hetzner needs
more operator time under any plausible rate, and the alternative was five times
dearer. But it does say plainly that the *magnitude* is a model output, not a
measurement, and that the cheapest way to strengthen it is to record real
operator hours rather than to refine the arithmetic.
That is the whole purpose of the scale: to make the difference between a firm
conclusion and a confident-looking one visible before it is acted on.
## Where it applies today
- `data/capability/*.json` — every consumption row and evidence item.
- Provider comparisons carried in a capability restatement.
Not yet retrofitted: `data/actuals/`, `data/control-cycle/`, and the option
fields of `data/optimization/`. Those already distinguish known from unknown and
carry `measurement_gaps` and `uncertainty` notes, but they do not yet grade a
present value. That is the next step, and it is the one that would let
`make thresholds` report a breach alongside the strength of the evidence
producing it.