freedom-intelligence/inventory/collection-policy.md
tegwick b5f911140b Establish Freedom Intelligence lab foundation and baseline research.
Add INTENT/SCOPE, daily-brief playbook, activity-core definition (disabled),
workplans FI-WP-0001..0003, baseline field survey with open-weight collection
recommendations, and inventory catalog candidates for the model reserve.
2026-07-24 00:15:27 +02:00

134 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Collection policy — open-weight reserve
**Status:** foundation
**Related:** `schema.yaml`, `docs/backup-storage-policy.md`, `INTENT.md`
---
## Purpose
Decide **what** enters the open-weight reserve, **who** may approve it, and
**when** a daily-brief candidate becomes a catalog entry with blobs on backup
storage.
---
## Goals
* Keep a **small, high-leverage** reserve — not a Hugging Face mirror
* Enforce **license and integrity** before download completes
* Match capacity to the backup storage soft quota
* Prefer models that serve axes **B** and **C**, plus strategic **A** open releases
---
## Eligibility (must pass all)
1. **Open weights** — weights obtainable under terms that allow offline retention for lab use
2. **Clear license** — SPDX or linkable license text; `allows_offline_retention: true`
3. **Stable provenance** — official org, tagged release, or commit revision (not anonymous drive-by reupload as sole source)
4. **Lab rationale** — written `reason` tied to at least one axis AD (usually B/C)
5. **Capacity** — estimated size fits under remaining soft quota (see backup storage policy)
Fail any gate → status `rejected` with reason, or never enter catalog.
---
## Priority rubric
| Priority | Guidance |
| -------- | -------- |
| **high** | Rare or strategically important; license/access risk of disappearance; uniquely strong for B/C at our hardware class; hard to re-obtain |
| **medium** | Clear lab use within 12 quarters; good quality/cost; easy enough to re-download but worth having cold |
| **low** | Nice to have; only collect if quota headroom is large and pull is cheap |
Daily brief **collection candidates** should set a suggested priority; approval may change it.
---
## Approval rule of thumb
| Estimated total size | Approval |
| -------------------- | -------- |
| **< 5 GiB** | Operator or lab agent may collect after license check; catalog entry required before or immediately after |
| **540 GiB** | Explicit operator approval (chat, workplan task, or signed catalog `approved_by`) |
| **> 40 GiB** | Operator approval **plus** check against soft quota and whether a smaller quant/variant suffices |
| **Any size if quota ≥ 70% used** | Operator approval required regardless of size |
| **Unclear license or ToS risk** | Do not collect; status `rejected` |
“Operator” means the human lab owner (or a documented delegate). Agents may
**nominate** (`status: candidate`) freely from briefs; they may **collect** only
within the < 5 GiB band when licenses are unambiguous — otherwise stop at
`candidate` / `approved`.
---
## Lifecycle
```text
brief nominates
→ candidate (catalog YAML, no blobs required)
→ approved (license + size + quota OK)
→ collecting (download in staging/)
→ collected (blobs complete, checksums recorded, storage_path set)
→ verified (optional re-hash / smoke load)
→ superseded|evicted (replaced or removed; metadata kept)
```
Rejected candidates stay in catalog only if useful as a decision record; otherwise omit.
---
## What we prefer to collect
* Small/mid instruct and code models that fit the hardware envelope
* Strong embedding / rerank models for local RAG
* Base models known to fine-tune well under QLoRA/LoRA on lab GPUs
* Official quant releases when they are the supported distribution
* Adapters and tokenizers that unlock a reserved base (as companions)
## What we usually skip
* Duplicate quants of the same revision already reserved
* Huge models with no near-term local run/train path and no access-risk story
* Merges/repacks without provenance
* Datasets larger than model weights unless separately justified (default: out of band)
* Anything requiring acceptance flows we cannot satisfy offline
---
## Companions
Tokenizers, LoRA adapters, and small eval fixtures may be collected when:
* they are required to use a reserved base, or
* they are small (< 1 GiB) and high leverage
Link via `companions` in the catalog schema.
---
## Brief integration
1. Brief section **Collection candidates** nominates items.
2. Operator/agent opens `inventory/catalog/{id}.yaml` with `status: candidate`.
3. Approval and download follow this policy and `docs/backup-storage-policy.md`.
4. Brief `brief_refs` on the entry point back to the nominating day(s).
---
## Eviction rule of thumb
When over quota or cleaning:
1. `low` priority, easily re-obtainable from still-live official URLs
2. Superseded revisions with a newer `verified` replacement
3. Never silent-delete: set `status: evicted`, clear or note `storage_path`, append `history`
---
## Non-goals
* Automatic bulk mirrors of entire orgs
* Collecting on every brief mention without priority
* Bypassing license gates for “research only” convenience