freedom-intelligence/inventory/collection-policy.md
tegwick 8a65297d56 Stream model reserve to Scaleway without local weight staging
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a09cbd-43c1-79f3-809e-1ee97b40b64d
2026-09-14 00:17:15 +02:00

155 lines
5.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Collection policy — open-weight reserve
**Status:** updated for Scaleway object reserve (2026-09-13)
**Related:** `schema.yaml`, `docs/backup-storage-policy.md`,
`docs/decisions/2026-09-13-scaleway-object-reserve.md`, `INTENT.md`
---
## Purpose
Decide **what** enters the open-weight reserve, **who** may approve it, and
**when** a daily-brief candidate becomes a catalog entry with blobs on
Scaleway (streamed through a remote worker with bounded RAM; no local staging).
---
## Goals
* Hold a **capability-first strategic reserve** of the best open weights we may
need later — **including models too large to run on current lab hardware**
* Also hold a **runnable spine** for day-to-day local ops and fine-tunes
* Enforce **license and integrity** before download completes
* Stay within the **€15 / month** soft budget — quality over mirror volume
* Prefer official provenance; not a full Hugging Face scrape
---
## Two envelopes (do not conflate)
| Envelope | Question | Effect on collection |
| -------- | -------- | -------------------- |
| **Run** | Can we infer / FT this *now*? | Tags `hardware_class`, prioritizes R tier for local work |
| **Reserve** | Is this among the best open artifacts worth keeping *just in case*? | **Not gated by current VRAM** |
**Runnability is not an eligibility gate.** An unrunnable SOTA open MoE can be
priority **high** if license and euro budget allow.
---
## Eligibility (must pass all)
1. **Open weights** — obtainable under terms that allow offline retention for lab use
2. **Clear license** — SPDX or linkable license text; `allows_offline_retention: true`
3. **Stable provenance** — official org, tagged release, or commit revision
4. **Lab rationale** — written `reason` (capability SOTA, runnable spine, embed, FT base, …)
5. **Capacity** — estimated monthly storage cost fits under remaining
€15 soft budget (Glacier for S, One Zone IA for R)
Fail any gate → `rejected` or never enter catalog.
---
## Collection tiers
| Tier | Code | Meaning |
| ---- | ---- | ------- |
| **Runnable spine** | **R** | Default local chat/code/embed/FT bases; keep resident |
| **Strategic capability** | **S** | Most capable open models (often large); may be T4+ only to *run* |
| **Watch / optional** | **W** | Secondary; collect only with clear headroom |
Catalog field: use `tags` including `tier-r` / `tier-s` / `tier-w` and
`priority: high|medium|low`.
### Priority rubric
| Priority | Guidance |
| -------- | -------- |
| **high** | Top open capability (S) or essential runnable spine (R); hard to re-obtain; license/access risk |
| **medium** | Strong but not unique; mid-size upgrades; companions (rerankers) |
| **low** | Nice-to-have; W tier |
---
## Approval rule of thumb
| Estimated total size | Approval |
| -------------------- | -------- |
| **< 5 GiB** | Operator or lab agent after license check |
| **540 GiB** | Explicit operator approval |
| **> 40 GiB** | Operator approval + remaining **euro-budget** check |
| **> 200 GiB (typical S giants)** | Operator approval + written note that the object goes to **Glacier** |
| **Any size if month ≥ 70% of €15** | Operator approval required |
| **Unclear license or ToS risk** | Do not collect; `rejected` |
Agents may **nominate** freely; they **collect** only in the < 5 GiB band with
unambiguous licenses unless the operator has approved the catalog entry.
---
## What we prefer to collect
### Runnable spine (R)
* Small/mid instruct and code models that fit the run envelope
* Strong embedding / rerank models for local RAG
* Bases known to fine-tune well under QLoRA on lab GPUs
### Strategic capability (S)
* **Best available open general / reasoning / code weights** at the frontier of open
* Large MoE or dense models even if current infra cannot serve them
* Prefer official compressed releases (FP8, published quant) when full
precision would blow the euro budget; Glacier makes full official trees
of the *top* S identities affordable (V4-Flash ~€0.42/mo, K3 ~€3.69/mo)
* One clear “best open” per capability niche is better than five near-duplicates
### Companions
Tokenizers, LoRA adapters, small eval fixtures when required to use a reserved base.
---
## What we usually skip
* Duplicate quants of the same revision already reserved
* Anonymous merges/repacks without provenance
* Closed weights
* Entire org mirrors
* Giant pretraining corpora (default out of band unless separately justified)
* Anything whose license forbids offline retention
---
## Lifecycle
```text
brief nominates
→ candidate
→ approved
→ collecting (remote HTTP stream → S3 multipart)
→ collected
→ verified
→ superseded|evicted
```
---
## Brief integration
1. Brief **Collection candidates** nominates R/S/W.
2. Catalog YAML under `inventory/catalog/`.
3. Download only after the Scaleway bucket is live (FI-WP-0004-T08) and
approval rules pass. Use [diskless streaming](../docs/streaming-reserve.md); SoT is `s3://`.
4. `collection.brief_refs` / research refs for provenance of the nomination.
---
## Eviction rule of thumb
When over the euro soft budget:
1. **W** tier and easily re-obtainable duplicates
2. Superseded revisions with a stronger verified successor
3. Never silent-delete **S** SOTA or sole **R** spine without operator note
4. Always set `status: evicted` and append `history`