freedom-intelligence/research/2026-07-24-baseline-field-survey.md
tegwick a83ef0a79a Register reserve under agents hub and pin VAULT storage.
Classify freedom-intelligence for State Hub (agents domain), sync FI-WP-0001..0003 with hub IDs, enrich catalog entries with profile/SWOT, and pin the open-weight reserve to D:\vault\coulomb\freedom-intelligence\ with strategic S-tier policy.
2026-07-28 00:25:21 +02:00

286 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Baseline field survey — Freedom Intelligence starting point
**Date:** 2026-07-24
**Type:** deep baseline (not a daily delta brief)
**Purpose:** Establish longitudinal memory so the **first daily research brief**
can be a true delta, not a full survey. Also recommend the **initial open-weight
collection** for the lab reserve.
**Axes:** A frontier · B edge/local · C training · D harness/fleet
**Sources:** public web surveys, HF/model docs, pricing roundups, harness project
sites (midlate 2026 public reporting). Prices and leaderboard claims move weekly —
treat numbers as **order-of-magnitude** and re-check on first daily.
---
## 1. Executive snapshot
| Theme | State of the field (mid-2026) | Lab implication |
| ----- | ----------------------------- | --------------- |
| Frontier | Multi-vendor race: OpenAI GPT-5.x, Anthropic Claude 4.x/Opus, Google Gemini 3.x, xAI Grok 4.x; coding agents dominate product narrative | Keep API access via `llm-connect`; do not try to host frontier closed weights |
| Price collapse | DeepSeek-class APIs often **10100×** cheaper than flagship closed models for many tasks | Route cheap/easy work to low-cost APIs; reserve local for privacy, offline, specialization |
| Open weights | Strong mid-size dense + MoE story (Qwen3, DeepSeek distill/V3-class, Gemma/Llama small, Mistral small) | Homelab can be *useful* without 8×H100 |
| Local stack | llama.cpp / vLLM / MLX / Ollama mature; Q4Q5 quants standard | Axis B is operationally ready |
| Training | QLoRA/Unsloth/TRL still the homelab path; full pretrain not lab-realistic | Axis C = adapters + domain SFT, not from-scratch |
| Harness | OpenHands, Aider, SWE-agent, OpenCode, Cline; closed CLIs (Claude Code, Codex) lead quality | Axis D: watch + integrate with sand-boxer / agent-harness, dont fork everything |
**Central proposition for the reserve:** collect a **compact spine** of (1) small
general instruct, (2) mid-size code/reason distill, (3) multilingual embed,
(4) optional larger “lab ceiling” base — not a mirror of Hugging Face.
---
## 2. Axis A — Frontier capabilities, benchmarks, price
### 2.1 Capability landscape
Public 2026 reporting groups the closed frontier into:
- **Flagship general / agentic** — OpenAI GPT-5.x family, Anthropic Claude Opus/Sonnet 4.x+, Google Gemini 3.x Pro
- **Fast / cheap tiers** — GPT mini/nano, Claude Haiku, Gemini Flash
- **Reasoning-heavy** — “thinking” / extended variants; cost dominated by output tokens
- **Coding productization** — agent harnesses matter as much as base model (SWE-bench Verified numbers often quoted 70%+ with strong closed models + good harness; novel-issue sets much lower ~1820%)
Benchmark caveats for the lab:
- Arena and public leaderboards are **gamed and lag** product quality
- SWE-bench Verified ≠ “our monorepos + NetKingdom tools”
- Prefer **internal harness scores** once we have them; use public numbers only as directional
### 2.2 Price-point trends (API, ~per 1M tokens)
Synthesized from 2026 pricing roundups (verify on provider pages before budget decisions):
| Band | Examples (reported ranges) | Role |
| ---- | -------------------------- | ---- |
| **Ultra-cheap** | DeepSeek V3.x / V4 Flash-class (~$0.140.28 in / ~$0.280.42 out; cache hits lower) | Bulk classify, draft, simple tools |
| **Budget closed** | Gemini Flash, GPT mini/nano, Haiku-class | Default automation when quality bar is modest |
| **Mid flagship** | Claude Sonnet-class, GPT mid-tier (~$25 in / ~$1015 out) | Daily agent coding, serious writing |
| **Top flagship** | Opus / GPT pro tiers (~$5+ in / $2530+ out; pro tiers much higher) | Hard reasoning, high-stakes review |
**Trend:** continuous compression of mid-tier quality cost; flagship **output**
tokens remain expensive. Caching and batch APIs swing real cost 50%+.
**Lab stance:**
- Frontier APIs stay first-class for hard work (via llm-connect / OpenRouter-class)
- Daily *research* about price moves is high value for routing policy
- Local models win on **privacy, offline, predictable unit cost, specialization** — not always raw SOTA
### 2.3 Open near-frontier (API or weights)
DeepSeek open-weight + cheap API continues to be the structural disruptor:
MIT-leaning releases for R1/V3-class artifacts, distill series for single-GPU
reasoning. Track **V3/V4 Flash vs Pro** naming carefully — product lines rename often.
---
## 3. Axis B — Edge, local, open compute
### 3.1 What “homelab useful” means now
| Class | Params (ballpark) | Typical local use |
| ----- | ----------------- | ----------------- |
| Edge / CPU-OK | ≤3B Q4 | Chat toys, classification, always-on small agents |
| Consumer GPU | 714B Q4/Q5 | Daily coding assist, RAG chat, tool loops |
| Single enthusiast GPU | 32B Q4 | Stronger reason/code; slower |
| Lab multi-GPU | 70B Q4 or sparse MoE | Optional ceiling; not required for v1 reserve |
Runtimes to track: **llama.cpp**, **vLLM**, **MLX** (Apple), **Ollama** (UX),
TensorRT-LLM (NVIDIA optimize). Speculative decoding + good quant matter more
than chasing every new 70B.
### 3.2 Open model families to watch (weights)
| Family | Why it matters | License caution |
| ------ | -------------- | --------------- |
| **Qwen3** (Alibaba) | Dense + MoE line; strong multilingual & tool use; embeddings/VL variants | Apache-class for many; check model card |
| **DeepSeek** (V3/R1 + distills) | Reasoning + cost; MIT for key R1/distill artifacts | Confirm per-revision card |
| **Llama 3.x / 4 text** | Ecosystem, tooling, small 1B3B instruct | Llama community license (not pure MIT) |
| **Gemma 3** | Google small/efficient; EmbeddingGemma | Gemma terms |
| **Mistral Small / Ministral** | EU-friendly commercial open options | Apache / research licenses vary |
| **Phi / Smol** class | Edge CPU experiments | Check commercial terms |
### 3.3 Embeddings & retrieval (local RAG spine)
For NetKingdom docs, code, and ops notes:
| Model | Role | Notes |
| ----- | ---- | ----- |
| **BGE-M3** (BAAI) | Default multilingual dense (+ multi-granularity) | MIT; ~0.5B; self-host staple |
| **Nomic Embed** text v1/v2 | Strong English/multilingual text | Apache |
| **Nomic Embed Code** / code-specialized | Code retrieval | Larger; optional P1 |
| **EmbeddingGemma-300M** | Tiny multilingual | Edge RAG |
| **Qwen3 embedding / VL-small** | Multimodal retrieval experiments | Watch cards |
Rerankers (BGE reranker, mini cross-encoders) are high leverage and small — good
companions once embed base is reserved.
---
## 4. Axis C — Homelab training & domain specialization
### 4.1 Realistic lab methods
| Method | Homelab fit | Use |
| ------ | ----------- | --- |
| QLoRA / LoRA (Unsloth, PEFT, TRL) | Excellent | Domain assistants (security, ops, railiance) |
| Preference (DPO/ORPO/KTO) | Good on 714B | Style and safety for agents |
| Continued pretrain | Possible on small bases | Domain language (logs, policies) |
| Full SFT of 70B+ | Poor without multi-GPU budget | Prefer distill/adapter on smaller |
| From-scratch pretrain | Out of scope | — |
### 4.2 Bases worth specializing
1. **78B instruct** (Qwen3-8B or Llama-3.1-8B-Instruct class) — default FT target
2. **14B** — quality step if VRAM allows
3. **DeepSeek-R1-Distill 14B/32B** — reason-heavy agents after FT discipline
4. **Code-specialized mid** — if code assistants dominate workload
Data: prefer **synthetic + curated NetKingdom docs** over scraping; keep PII out
of training sets (separate policy later).
### 4.3 Tooling to track (not collect as weights)
Unsloth, Axolotl, TRL, llama-factory, mergekit, preference datasets (UltraFeedback-class),
eval harnesses (lm-eval, custom domain suites).
---
## 5. Axis D — Agent harnesses, orchestration, sandboxing, fleet
### 5.1 Open harness landscape (2026)
| Project | Shape | Why watch |
| ------- | ----- | --------- |
| **OpenHands** | Full agent platform + Canvas; MIT; multi-backend | Leading open coding-agent platform; multi-agent |
| **Aider** | Git-native CLI pair programmer | Atomic commits; simple; great for ops scripts |
| **SWE-agent / mini-SWE-agent** | Research harness, clean ACI | Benchmark + security CTF modes |
| **OpenCode** | Open Claude Code alternative | High interest community |
| **Cline** | IDE agent | Editor-centric workflows |
| **Closed CLIs** | Claude Code, Codex, Gemini CLI | Quality bar for comparison |
SWE-bench: harness + model co-determine scores; Mini-SWE-Agent claims strong
results in very little code — architecture lessons for Coulomb harnesses.
### 5.2 NetKingdom adjacency
| Coulomb piece | Relation to D |
| ------------- | ------------- |
| **sand-boxer** | Where agents run safely |
| **activity-core** | When work is due (this labs daily brief) |
| **agent-harness / kaizen** | Who executes with profiles |
| **state-hub** | What changed; task spawn completion |
| **llm-connect** | Which model backend |
Freedom Intelligence **reports** on harness tech; it does not replace sand-boxer.
### 5.3 Sandboxing & fleet themes to track
- Default-deny network + TTL sandboxes
- Credential routing (OpenBao / ops-warden) separate from model weights
- Multi-agent kanban / workplan-tied loops (Ralph-workplan class)
- Cost and rate-limit fleet management for API models
---
## 6. Collection recommendations (open-weight reserve)
### 6.1 Decision framework (recap)
- Prefer **B/C** leverage + clear license + stable org provenance
- Size gates per `inventory/collection-policy.md`
- Backup path must be pinned (FI-WP-0001-T04) before large pulls
- Quants: prefer **official or well-known GGUF/AWQ** only after base id is cataloged; v1 recommends **safetensors official** first, quant as companion later
### 6.2 P0 — seed now (approve after storage pin)
| # | Model (representative HF id) | Est. size class | Axes | Why |
| - | ---------------------------- | --------------- | ---- | --- |
| 1 | `Qwen/Qwen3-8B` (or current 8B-Instruct sibling on card) | ~16 GB fp16 / ~5 GB Q4 | B, C | Best default general + FT base; multilingual; tool-friendly |
| 2 | `meta-llama/Llama-3.2-3B-Instruct` | ~6 GB fp16 / ~2 GB Q4 | B | Tiny always-on / edge; huge ecosystem |
| 3 | `BAAI/bge-m3` | ~2 GB | B | Multilingual RAG staple; MIT |
| 4 | `deepseek-ai/DeepSeek-R1-Distill-Qwen-14B` (or 8B if VRAM tight) | ~28 GB fp16 / ~9 GB Q4 | B, C | Local reasoning without full R1; MIT distill story |
| 5 | `nomic-ai/nomic-embed-text-v1.5` (or current v2 text) | <1 GB | B | Lightweight embed alternative / A/B with BGE |
**Rough P0 cold storage (fp16 worst case):** ~5060 GB.
**With Q4 preference for LLMs + fp for embeds:** often **~2030 GB**. Prefer quants only when catalog notes the exact quant provenance.
### 6.3 P1 — next wave (quota permitting)
| # | Model | Why |
| - | ----- | --- |
| 6 | `Qwen/Qwen3-14B` (instruct) | Quality step for single-GPU chat/code |
| 7 | `google/gemma-3-4b-it` or EmbeddingGemma-300M | Efficient Google line; embed edge |
| 8 | `mistralai/Mistral-Small-3.x` instruct (current card) | EU-adjacent open option |
| 9 | `deepseek-ai/DeepSeek-R1-Distill-Qwen-32B` | Stronger local reason if ≥24 GB VRAM class |
| 10 | Code embed: `nomic-ai/nomic-embed-code` or Jina code embed | Code RAG for repos |
| 11 | Small reranker: `BAAI/bge-reranker-v2-m3` | Cheap quality win for RAG |
### 6.4 Strategic large (updated 2026-07-24)
**Superseded by operator decision:** 1 TB local NAS + capability-first reserve.
Full DeepSeek-V3/R1-class and 70B dense opens **are in scope** even if unrunnable
on current GPUs. See:
- `docs/decisions/2026-07-24-nas-strategic-reserve.md`
- `research/2026-07-24-nas-strategic-collection-plan.md`
| Item | Guidance |
| ---- | -------- |
| Full **DeepSeek-V3 / R1** class | **Strategic (S)** — collect under 850 GiB soft quota; prefer official compressed if needed |
| **Llama 70B / Qwen 72B** class | **Strategic (S)** dense — good second fill after one primary MoE |
| Every new quant repack on HF | Skip; track base official only |
### 6.5 Explicit non-recommendations (for now)
- Closed API-only weights
- Anonymous “uncensored” reuploads without provenance
- Duplicate quants of the same revision
- Giant pretraining corpora as “models”
### 6.6 Catalog seed
YAML candidates for P0P1 are under `inventory/catalog/` with
`status: candidate` and `brief_refs` pointing at this survey. Approval and
download are **FI-WP-0003**.
---
## 7. Implications for the first daily brief
The first `briefs/YYYY/MM/YYYY-MM-DD.md` after this survey should:
1. **Not** restate this whole document
2. Cover **only deltas** since 2026-07-24 (new releases, price changes, harness versions)
3. Revisit P0/P1 if a **new open model** clearly displaces a candidate
4. Note any **license or access shocks** (takedowns, geo blocks)
5. Optionally mark FI-WP-0001-T04 / FI-WP-0003 progress under Lab implications
Automation: activity-core def `fi-daily-research-brief` (disabled until FI-WP-0002
resolver + dry-run). Manual playbook: `docs/daily-brief-playbook.md`.
---
## 8. Open questions for later briefs
1. Exact **hardware envelope** measured VRAM on lab hosts?
2. Prefer **GGUF Q5_K_M** vs official safetensors for P0?
3. Is **German-first** multilingual quality a hard requirement (affects embed + instruct choice)?
4. Will local models serve **air-gapped NetKingdom** scenarios or only cost/privacy convenience?
5. Single **code model** vs general instruct for domain FT?
---
## 9. Source anchors (non-exhaustive)
- DeepSeek R1 release / MIT notes — DeepSeek API docs & HF model cards
- Pricing roundups 2026 — IntuitionLabs, MorphLLM, provider pricing pages
- Embeddings 2026 comparisons — BentoML, Milvus/Zilliz, BGE-M3 cards
- Harness landscape — OpenHands, Aider, SWE-agent comparisons (LocalAI Master, vendor blogs)
- Open-weight commercial lists — community curations (verify each card)
Re-verify any number before spend or multi-hundred-GB download.
---
*Baseline established 2026-07-24 for Freedom Intelligence. Daily briefs begin as deltas from here.*