Make the sensing loop durable and pin the reserve to Scaleway

The weekday brief only counts when the file is on origin/main. Hub
events without git objects are failures. The open-weight SoT moves
from the VAULT 850 GiB quota to a dedicated Scaleway bucket (One Zone
IA for R, Glacier for S) so V4-Flash and K3 are economically in-scope.

Catalogs the real DeepSeek-V4-Flash-0731 (MIT ~167 GiB MoE, not 12B)
and nominates Qwen3.8-27B. Bucket create remains FI-WP-0004-T08.

Assistant: grok
Assistant-Session: 01a09c6a-1cfc-75b1-a78d-c13eaf22241d
This commit is contained in:
tegwick 2026-09-13 22:45:43 +02:00
parent 8832652ad3
commit a78187c33e
18 changed files with 1125 additions and 226 deletions

View file

@ -1,8 +1,8 @@
# Model inventory
In-repo **catalog of open-weight revisions** reserved (or nominated) for the lab.
Weight blobs live on the **VAULT HD** at `D:\vault\coulomb\freedom-intelligence\`
see `docs/backup-storage-policy.md`.
Weight blobs live on **Scaleway Object Storage** (`s3://railiance-fi-open-weight-reserve/…`).
VAULT HD is staging only — see `docs/backup-storage-policy.md`.
## Layout

View file

@ -1,8 +1,10 @@
# Open-weight reserve status
**As of:** 2026-08-03
**Store:** `/mnt/d/vault/coulomb/freedom-intelligence/` (`D:\vault\coulomb\freedom-intelligence\`)
**Soft quota:** 850 GiB · **Hard stop:** 920 GiB
**As of:** 2026-09-13
**SoT store:** Scaleway Object Storage `nl-ams` bucket
`railiance-fi-open-weight-reserve` (**planned** — FI-WP-0004-T08)
**Staging:** `/mnt/d/vault/coulomb/freedom-intelligence/` (VAULT HD)
**Capacity gate:** **€15 / month** soft (not 850 GiB)
---
@ -10,64 +12,70 @@
| Class | Count | Status mix |
| ----- | ----: | ---------- |
| Catalog entries | 13 | + Kimi K3 S1 |
| **verified** on disk | 4 | R embeds + Qwen3-8B + R1-Distill-14B |
| **approved** (pull deferred) | S giants + Llama edge + **Kimi K3** | capacity / HF gate |
| **candidate** (W) | 3 | headroom-gated |
| Catalog entries | 15 | + V4-Flash S, + Qwen3.8-27B W/R |
| **verified** on VAULT | 4 | R embeds + Qwen3-8B + R1-Distill-14B |
| **approved**, blobs missing | S giants + Llama-3.2-3B + **V4-Flash** | HF gate / bucket not live |
| **candidate** | 4 | W + Qwen3.8-27B |
### Disk use (model tree)
### Disk / object use
| Path | Role | Approx |
| ---- | ---- | ------ |
| `models/nomic-ai__nomic-embed-text-v1.5/main` | R embed | ~0.5 GiB verified |
| `models/BAAI__bge-m3/main` | R embed | ~2.1 GiB verified |
| `models/Qwen__Qwen3-8B/main` | R instruct | ~15.3 GiB verified |
| `models/deepseek-ai__DeepSeek-R1-Distill-Qwen-14B/main` | R reason | ~27.5 GiB verified |
| Location | Role | Approx |
| -------- | ---- | ------ |
| VAULT `models/nomic-ai__nomic-embed-text-v1.5/main` | R embed (staging cache) | ~0.5 GiB verified |
| VAULT `models/BAAI__bge-m3/main` | R embed (staging cache) | ~2.1 GiB verified |
| VAULT `models/Qwen__Qwen3-8B/main` | R instruct (staging cache) | ~15.3 GiB verified |
| VAULT `models/deepseek-ai__DeepSeek-R1-Distill-Qwen-14B/main` | R reason (staging cache) | ~27.5 GiB verified |
| Scaleway bucket | SoT | **empty** (not created) |
**~45 GiB** used → **~800+ GiB** soft-quota headroom (insufficient for full Kimi K3 ~1454 GiB).
VAULT free (2026-08-03): ~1.7 TB free of 1.9 TB — full K3 would still dominate the volume and **violate soft quota**.
VAULT ~45 GiB used remains a **cache**, not the reserve.
---
## Decision matrix (2026-08-03)
## Decision matrix (2026-09-13)
| Id | Tier | Decision | Disk |
| -- | ---- | -------- | ---- |
| Qwen3-8B | R | **verified** | yes |
| Id | Tier | Decision | Blobs |
| -- | ---- | -------- | ----- |
| Qwen3-8B | R | **verified** on VAULT cache | local only |
| Llama-3.2-3B-Instruct | R | **approved** | deferred — HF gate |
| BGE-M3 | R | **verified** | yes |
| R1-Distill-Qwen-14B | R | **verified** | yes |
| nomic-embed-text-v1.5 | R | **verified** | yes |
| **Kimi-K3** | **S1** | **approved**, **collection deferred** | no — ~1454 GiB > soft quota |
| DeepSeek-V3 | S1→secondary | **approved**, deferred | no — K3 is primary open SOTA identity |
| BGE-M3 | R | **verified** on VAULT cache | local only |
| R1-Distill-Qwen-14B | R | **verified** on VAULT cache | local only |
| nomic-embed-text-v1.5 | R | **verified** on VAULT cache | local only |
| **DeepSeek-V4-Flash-0731** | **S1** | **approved****first Scaleway pull** | no — ~167 GiB Glacier |
| Kimi-K3 | S1 secondary | **approved**, deferred until after V4-Flash | no — ~1454 GiB Glacier ≈ €3.69/mo |
| DeepSeek-V3 | S superseded identity | **approved**, deferred | no — V4-Flash is the collectable S1 |
| DeepSeek-R1 full | S2 | **approved**, deferred | no |
| Qwen3-72B | S4 | **approved**, deferred | no |
| Llama-3.3-70B-Instruct | S4 | **approved**, deferred | no |
| Qwen3.8-27B | W/R | **candidate** | no |
| Qwen3-14B | W | candidate | no |
| R1-Distill-32B | W | candidate | no |
| bge-reranker-v2-m3 | W | candidate | no |
**Correction:** there is no `DeepSeek-V4-Flash-0731-12B`. Automated briefs
that used that id were wrong. Do not collect under that name.
---
## Hardware coverage gaps
| Need | Status |
| ---- | ------ |
| Default local instruct (8B) | **Covered** (Qwen3-8B) |
| Default local instruct (8B) | **Covered** (Qwen3-8B cache) |
| Stronger local instruct (27B) | **Candidate** Qwen3.8-27B |
| Edge micro-agent | Llama-3.2-3B blocked on HF auth |
| Multilingual RAG embed | **Covered** (BGE-M3 + nomic) |
| Local reason | **Covered** (R1-Distill-14B) |
| Frontier-open general MoE offline | **Cataloged** Kimi K3 — pull blocked on capacity |
| Full open reasoner MoE | Deferred S2 |
| Frontier-open MIT MoE offline | **Cataloged** V4-Flash — pull blocked on bucket |
| Full open giant (K3) | Cataloged; Glacier-affordable; after V4-Flash |
---
## Next pulls (operator)
1. **Capacity decision for Kimi K3** (raise soft quota / dedicate disk / official compressed only).
2. Set `HF_TOKEN` and pull Llama-3.2-3B-Instruct.
3. Optional S4: one of Qwen3-72B vs Llama-3.3-70B under remaining quota.
4. DeepSeek V4-Flash-0731 as mid open value candidate (catalog later if desired).
1. **FI-WP-0004-T08** create `railiance-fi-open-weight-reserve` (no 30-day expiry).
2. **FI-WP-0004-T09** collect `deepseek-ai/DeepSeek-V4-Flash-0731``strategic/` Glacier.
3. Set `HF_TOKEN` and pull Llama-3.2-3B-Instruct into `models/` One Zone IA.
4. Decide Qwen3.8-27B vs keeping Qwen3-8B as the R instruct default.
5. Optional: Kimi K3 on Glacier if the month is still under €15.
Tool: `scripts/collect_model.py` (weights-only, sequential, writes `MANIFEST.json`).
Tool: `scripts/collect_model.py` (weights-only, sequential, MANIFEST, optional `--s3-bucket`).

View file

@ -0,0 +1,74 @@
id: Qwen__Qwen3.8-27B__candidate
status: candidate
name: Qwen3.8-27B
org: Qwen
source:
kind: huggingface
url: https://huggingface.co/Qwen/Qwen3.8-27B
revision: main
model_card_url: https://huggingface.co/Qwen/Qwen3.8-27B
project_url: https://qwen.ai/
license:
spdx: Apache-2.0
url: https://huggingface.co/Qwen/Qwen3.8-27B
allows_offline_retention: true
allows_local_ops: true
allows_fine_tune: true
notes: "Apache 2.0 on the 27B dense multimodal card (2026-08-14). Confirm at pin. Not the Qwen3.8-Max bespoke licence."
profile:
summary: "Dense 27B multimodal Apache-2.0 — candidate R-spine upgrade vs Qwen3-8B."
original_source: https://huggingface.co/Qwen/Qwen3.8-27B
use_cases:
- "Local instruct / vision on T2T3 (24 GB class at Q4)"
- "Coding and office agents that need more than 8B"
- "Homelab QLoRA base"
sweet_spots:
- "Apache-2.0, ungated, actually runnable unlike V4-Flash / K3"
- "Native vision; 262k context (YaRN toward 1M)"
not_ideal_for:
- "Replacing the 8B always-on default until measured"
- "Standing in for Qwen3.8-2.4T-A95B (different licence, multi-TB)"
capability_notes: >
Released 2026-08-14, the day the automated brief last landed on
origin — so the sensing loop never recorded it. Dense 27.8B,
multimodal, Apache 2.0. Distinct from Qwen3.8-2.4T-A95B (custom
Max licence, ~4.89 TB). Nominate as W/R; do not pull until euro
budget and R-spine review after V4-Flash lands.
swot:
strengths:
- "Apache-2.0 dense multimodal at a size we can actually serve later"
- "Likely better local default than Qwen3-8B if VRAM allows"
weaknesses:
- "Larger than current verified R instruct (8B)"
- "Hardware envelope hosts still TBD"
opportunities:
- "R-spine upgrade path that does not need Glacier"
- "Vision in the local working set"
threats:
- "Qwen 4.x could supersede before we pull"
size:
total_bytes: 0
total_human: "~56 GiB class (reports); measure at pull — Q4 much smaller"
hardware_class:
min_vram_gb_q4: 20
min_vram_gb_fp16: 56
notes: "T2 stretch / T3. Not T1."
axes: [B, C]
priority: medium
collection:
approved_by: ""
approved_at: null
downloaded_at: null
downloaded_by: ""
storage_path: ""
brief_refs:
- briefs/2026/09/2026-09-13.md
reason: "Missed by the Aug 14 brief cutoff; Apache-2.0 dense 27B is the obvious R-upgrade candidate."
tags: [tier-w, tier-r, instruct, vision, apache]
companions: []
notes: "Do not confuse with Qwen3.8-2.4T-A95B (S-class, custom licence, multi-TB)."
history:
- at: "2026-09-13"
event: nominated
by: grok
detail: "Catch-up brief after cadence hole. Candidate only."

View file

@ -0,0 +1,94 @@
id: deepseek-ai__DeepSeek-V4-Flash-0731__strategic
status: approved
name: DeepSeek-V4-Flash-0731
org: deepseek-ai
source:
kind: huggingface
url: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
revision: main
model_card_url: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
project_url: https://www.deepseek.com/
paper_url: https://arxiv.org/abs/2606.19348
license:
spdx: MIT
url: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
allows_offline_retention: true
allows_local_ops: true
allows_fine_tune: true
notes: "Official card: repository and weights MIT. Confirm at pin."
profile:
summary: "S-tier open MIT MoE — 284B base / ~13B active (304B with DSpark); not a 12B dense model."
original_source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
use_cases:
- "Primary open-weight agentic / coding identity after Jul 2026"
- "Offline optionality if DeepSeek API access changes"
- "Future multi-GPU serve (vLLM DSpark / SGLang); not current R-spine"
- "Capability bar for daily briefs against closed frontier"
sweet_spots:
- "Best MIT-licensed open near-frontier that actually fits cheap Glacier (~167 GiB HF tree)"
- "Fits the €15/month Scaleway budget with room for K3"
- "Displaces DeepSeek-V3 as the collectable S1 identity"
not_ideal_for:
- "Current single-consumer-GPU daily ops (use R spine / Qwen3.8-27B)"
- "Treating 'Flash' as a 12B dense model (that reading was an executor error)"
capability_notes: >
Official V4-Flash release (2026-07-31). Same structure as
DeepSeek-V4-Flash-DSpark (speculative decoding module attached).
HF card reports 304B params (284B base + draft). ~13B active.
1M context. Terminal Bench 2.1 82.7. HF tree listed ~167 GB.
Automated briefs 2026-08-04…14 repeatedly called this a 12B dense
consumer-GPU model and once claimed it was already collected. It
was not. Size class is S, not R.
swot:
strengths:
- "MIT, official org, actually downloadable"
- "Near-frontier agentic scores at ~13B active"
- "Footprint (~167 GiB) is cheap on Glacier and was always inside the old 850 GiB quota"
weaknesses:
- "Serve needs multi-GPU / high unified memory, not T1T2"
- "Executor hallucinated it as 12B for two weeks — catalog must stay the authority"
opportunities:
- "First S pull onto Scaleway (FI-WP-0004-T09)"
- "Compressed / GGUF community packs exist if we want a later R experiment (not a substitute for the official tree)"
threats:
- "Successor V4.x could land while we still have no object on disk"
- "Quota-era deferral already cost a month of optionality"
size:
total_bytes: 179000000000
total_human: "~167 GiB official HF tree (card file listing 2026-09-13); 304B params / ~13B active"
hardware_class:
min_vram_gb_q4: 0
min_vram_gb_fp16: 0
notes: "Beyond current lab run envelope (T3T4). Community 2-bit reports exist for 128 GB unified memory; not an R-spine claim."
axes: [A, B, D]
priority: high
collection:
approved_by: "grok"
approved_at: "2026-09-13"
downloaded_at: null
downloaded_by: ""
storage_path: ""
brief_refs:
- briefs/2026/08/2026-08-03.md
- briefs/2026/09/2026-09-13.md
- docs/decisions/2026-09-13-scaleway-object-reserve.md
reason: "S1 collectable — strongest MIT open MoE that fits the Scaleway euro budget; first pull after bucket create."
tags: [tier-s, strategic, moe, mit, beyond-run-envelope]
companions: []
notes: >
Do not confuse with a 12B dense checkpoint. There is no
deepseek-ai/DeepSeek-V4-Flash-0731-12B. Pull the official card repo.
Storage class at upload: strategic/ → Glacier.
history:
- at: "2026-08-03"
event: nominated
by: grok
detail: "Catch-up brief named V4-Flash-0731 as mid-tier open value leader."
- at: "2026-08-04"
event: misidentified
by: rein-aharness
detail: "Automated briefs treated it as 12B dense; never cataloged; 08-07 claimed on-disk collection that did not happen."
- at: "2026-09-13"
event: approved
by: grok
detail: "Corrected to 304B-class MoE ~167 GiB MIT. Approved for Scaleway Glacier once FI-WP-0004-T08 bucket exists."

View file

@ -1,15 +1,16 @@
# Collection policy — open-weight reserve
**Status:** updated for 1 TB NAS + strategic capability reserve (2026-07-24)
**Status:** updated for Scaleway object reserve (2026-09-13)
**Related:** `schema.yaml`, `docs/backup-storage-policy.md`,
`docs/decisions/2026-07-24-nas-strategic-reserve.md`, `INTENT.md`
`docs/decisions/2026-09-13-scaleway-object-reserve.md`, `INTENT.md`
---
## Purpose
Decide **what** enters the open-weight reserve, **who** may approve it, and
**when** a daily-brief candidate becomes a catalog entry with blobs on the NAS.
**when** a daily-brief candidate becomes a catalog entry with blobs on
Scaleway (staged locally first).
---
@ -19,7 +20,7 @@ Decide **what** enters the open-weight reserve, **who** may approve it, and
need later — **including models too large to run on current lab hardware**
* Also hold a **runnable spine** for day-to-day local ops and fine-tunes
* Enforce **license and integrity** before download completes
* Stay within the **1 TB NAS soft quota (850 GiB)** — quality over mirror volume
* Stay within the **€15 / month** soft budget — quality over mirror volume
* Prefer official provenance; not a full Hugging Face scrape
---
@ -32,7 +33,7 @@ Decide **what** enters the open-weight reserve, **who** may approve it, and
| **Reserve** | Is this among the best open artifacts worth keeping *just in case*? | **Not gated by current VRAM** |
**Runnability is not an eligibility gate.** An unrunnable SOTA open MoE can be
priority **high** if license and quota allow.
priority **high** if license and euro budget allow.
---
@ -42,7 +43,8 @@ priority **high** if license and quota allow.
2. **Clear license** — SPDX or linkable license text; `allows_offline_retention: true`
3. **Stable provenance** — official org, tagged release, or commit revision
4. **Lab rationale** — written `reason` (capability SOTA, runnable spine, embed, FT base, …)
5. **Capacity** — estimated size fits under remaining soft quota (850 GiB)
5. **Capacity** — estimated monthly storage cost fits under remaining
€15 soft budget (Glacier for S, One Zone IA for R)
Fail any gate → `rejected` or never enter catalog.
@ -75,9 +77,9 @@ Catalog field: use `tags` including `tier-r` / `tier-s` / `tier-w` and
| -------------------- | -------- |
| **< 5 GiB** | Operator or lab agent after license check |
| **540 GiB** | Explicit operator approval |
| **> 40 GiB** | Operator approval + remaining soft-quota check |
| **> 200 GiB (typical S giants)** | Operator approval + written note on which other S models may need to wait |
| **Any size if quota ≥ 70% used** | Operator approval required |
| **> 40 GiB** | Operator approval + remaining **euro-budget** check |
| **> 200 GiB (typical S giants)** | Operator approval + written note that the object goes to **Glacier** |
| **Any size if month ≥ 70% of €15** | Operator approval required |
| **Unclear license or ToS risk** | Do not collect; `rejected` |
Agents may **nominate** freely; they **collect** only in the < 5 GiB band with
@ -97,8 +99,9 @@ unambiguous licenses unless the operator has approved the catalog entry.
* **Best available open general / reasoning / code weights** at the frontier of open
* Large MoE or dense models even if current infra cannot serve them
* Prefer official compressed releases (FP8, published quant) when full precision
would exhaust the 1 TB NAS
* Prefer official compressed releases (FP8, published quant) when full
precision would blow the euro budget; Glacier makes full official trees
of the *top* S identities affordable (V4-Flash ~€0.42/mo, K3 ~€3.69/mo)
* One clear “best open” per capability niche is better than five near-duplicates
### Companions
@ -124,7 +127,7 @@ Tokenizers, LoRA adapters, small eval fixtures when required to use a reserved b
brief nominates
→ candidate
→ approved
→ collecting (NAS staging/)
→ collecting (local staging/ then S3 upload)
→ collected
→ verified
→ superseded|evicted
@ -136,14 +139,15 @@ brief nominates
1. Brief **Collection candidates** nominates R/S/W.
2. Catalog YAML under `inventory/catalog/`.
3. Download only after storage path is pinned (`docs/backup-storage-policy.md`) and approval rules pass.
3. Download only after the Scaleway bucket is live (FI-WP-0004-T08) and
approval rules pass. Stage on VAULT; SoT is `s3://`.
4. `collection.brief_refs` / research refs for provenance of the nomination.
---
## Eviction rule of thumb
When over soft quota:
When over the euro soft budget:
1. **W** tier and easily re-obtainable duplicates
2. Superseded revisions with a stronger verified successor