Classify freedom-intelligence for State Hub (agents domain), sync FI-WP-0001..0003 with hub IDs, enrich catalog entries with profile/SWOT, and pin the open-weight reserve to D:\vault\coulomb\freedom-intelligence\ with strategic S-tier policy.
14 KiB
Baseline field survey — Freedom Intelligence starting point
Date: 2026-07-24
Type: deep baseline (not a daily delta brief)
Purpose: Establish longitudinal memory so the first daily research brief
can be a true delta, not a full survey. Also recommend the initial open-weight
collection for the lab reserve.
Axes: A frontier · B edge/local · C training · D harness/fleet
Sources: public web surveys, HF/model docs, pricing roundups, harness project
sites (mid–late 2026 public reporting). Prices and leaderboard claims move weekly —
treat numbers as order-of-magnitude and re-check on first daily.
1. Executive snapshot
| Theme | State of the field (mid-2026) | Lab implication |
|---|---|---|
| Frontier | Multi-vendor race: OpenAI GPT-5.x, Anthropic Claude 4.x/Opus, Google Gemini 3.x, xAI Grok 4.x; coding agents dominate product narrative | Keep API access via llm-connect; do not try to host frontier closed weights |
| Price collapse | DeepSeek-class APIs often 10–100× cheaper than flagship closed models for many tasks | Route cheap/easy work to low-cost APIs; reserve local for privacy, offline, specialization |
| Open weights | Strong mid-size dense + MoE story (Qwen3, DeepSeek distill/V3-class, Gemma/Llama small, Mistral small) | Homelab can be useful without 8×H100 |
| Local stack | llama.cpp / vLLM / MLX / Ollama mature; Q4–Q5 quants standard | Axis B is operationally ready |
| Training | QLoRA/Unsloth/TRL still the homelab path; full pretrain not lab-realistic | Axis C = adapters + domain SFT, not from-scratch |
| Harness | OpenHands, Aider, SWE-agent, OpenCode, Cline; closed CLIs (Claude Code, Codex) lead quality | Axis D: watch + integrate with sand-boxer / agent-harness, don’t fork everything |
Central proposition for the reserve: collect a compact spine of (1) small general instruct, (2) mid-size code/reason distill, (3) multilingual embed, (4) optional larger “lab ceiling” base — not a mirror of Hugging Face.
2. Axis A — Frontier capabilities, benchmarks, price
2.1 Capability landscape
Public 2026 reporting groups the closed frontier into:
- Flagship general / agentic — OpenAI GPT-5.x family, Anthropic Claude Opus/Sonnet 4.x+, Google Gemini 3.x Pro
- Fast / cheap tiers — GPT mini/nano, Claude Haiku, Gemini Flash
- Reasoning-heavy — “thinking” / extended variants; cost dominated by output tokens
- Coding productization — agent harnesses matter as much as base model (SWE-bench Verified numbers often quoted 70%+ with strong closed models + good harness; novel-issue sets much lower ~18–20%)
Benchmark caveats for the lab:
- Arena and public leaderboards are gamed and lag product quality
- SWE-bench Verified ≠ “our monorepos + NetKingdom tools”
- Prefer internal harness scores once we have them; use public numbers only as directional
2.2 Price-point trends (API, ~per 1M tokens)
Synthesized from 2026 pricing roundups (verify on provider pages before budget decisions):
| Band | Examples (reported ranges) | Role |
|---|---|---|
| Ultra-cheap | DeepSeek V3.x / V4 Flash-class (~$0.14–0.28 in / ~$0.28–0.42 out; cache hits lower) | Bulk classify, draft, simple tools |
| Budget closed | Gemini Flash, GPT mini/nano, Haiku-class | Default automation when quality bar is modest |
| Mid flagship | Claude Sonnet-class, GPT mid-tier (~$2–5 in / ~$10–15 out) | Daily agent coding, serious writing |
| Top flagship | Opus / GPT pro tiers (~$5+ in / $25–30+ out; pro tiers much higher) | Hard reasoning, high-stakes review |
Trend: continuous compression of mid-tier quality cost; flagship output tokens remain expensive. Caching and batch APIs swing real cost 50%+.
Lab stance:
- Frontier APIs stay first-class for hard work (via llm-connect / OpenRouter-class)
- Daily research about price moves is high value for routing policy
- Local models win on privacy, offline, predictable unit cost, specialization — not always raw SOTA
2.3 Open near-frontier (API or weights)
DeepSeek open-weight + cheap API continues to be the structural disruptor: MIT-leaning releases for R1/V3-class artifacts, distill series for single-GPU reasoning. Track V3/V4 Flash vs Pro naming carefully — product lines rename often.
3. Axis B — Edge, local, open compute
3.1 What “homelab useful” means now
| Class | Params (ballpark) | Typical local use |
|---|---|---|
| Edge / CPU-OK | ≤3B Q4 | Chat toys, classification, always-on small agents |
| Consumer GPU | 7–14B Q4/Q5 | Daily coding assist, RAG chat, tool loops |
| Single enthusiast GPU | 32B Q4 | Stronger reason/code; slower |
| Lab multi-GPU | 70B Q4 or sparse MoE | Optional ceiling; not required for v1 reserve |
Runtimes to track: llama.cpp, vLLM, MLX (Apple), Ollama (UX), TensorRT-LLM (NVIDIA optimize). Speculative decoding + good quant matter more than chasing every new 70B.
3.2 Open model families to watch (weights)
| Family | Why it matters | License caution |
|---|---|---|
| Qwen3 (Alibaba) | Dense + MoE line; strong multilingual & tool use; embeddings/VL variants | Apache-class for many; check model card |
| DeepSeek (V3/R1 + distills) | Reasoning + cost; MIT for key R1/distill artifacts | Confirm per-revision card |
| Llama 3.x / 4 text | Ecosystem, tooling, small 1B–3B instruct | Llama community license (not pure MIT) |
| Gemma 3 | Google small/efficient; EmbeddingGemma | Gemma terms |
| Mistral Small / Ministral | EU-friendly commercial open options | Apache / research licenses vary |
| Phi / Smol class | Edge CPU experiments | Check commercial terms |
3.3 Embeddings & retrieval (local RAG spine)
For NetKingdom docs, code, and ops notes:
| Model | Role | Notes |
|---|---|---|
| BGE-M3 (BAAI) | Default multilingual dense (+ multi-granularity) | MIT; ~0.5B; self-host staple |
| Nomic Embed text v1/v2 | Strong English/multilingual text | Apache |
| Nomic Embed Code / code-specialized | Code retrieval | Larger; optional P1 |
| EmbeddingGemma-300M | Tiny multilingual | Edge RAG |
| Qwen3 embedding / VL-small | Multimodal retrieval experiments | Watch cards |
Rerankers (BGE reranker, mini cross-encoders) are high leverage and small — good companions once embed base is reserved.
4. Axis C — Homelab training & domain specialization
4.1 Realistic lab methods
| Method | Homelab fit | Use |
|---|---|---|
| QLoRA / LoRA (Unsloth, PEFT, TRL) | Excellent | Domain assistants (security, ops, railiance) |
| Preference (DPO/ORPO/KTO) | Good on 7–14B | Style and safety for agents |
| Continued pretrain | Possible on small bases | Domain language (logs, policies) |
| Full SFT of 70B+ | Poor without multi-GPU budget | Prefer distill/adapter on smaller |
| From-scratch pretrain | Out of scope | — |
4.2 Bases worth specializing
- 7–8B instruct (Qwen3-8B or Llama-3.1-8B-Instruct class) — default FT target
- 14B — quality step if VRAM allows
- DeepSeek-R1-Distill 14B/32B — reason-heavy agents after FT discipline
- Code-specialized mid — if code assistants dominate workload
Data: prefer synthetic + curated NetKingdom docs over scraping; keep PII out of training sets (separate policy later).
4.3 Tooling to track (not collect as weights)
Unsloth, Axolotl, TRL, llama-factory, mergekit, preference datasets (UltraFeedback-class), eval harnesses (lm-eval, custom domain suites).
5. Axis D — Agent harnesses, orchestration, sandboxing, fleet
5.1 Open harness landscape (2026)
| Project | Shape | Why watch |
|---|---|---|
| OpenHands | Full agent platform + Canvas; MIT; multi-backend | Leading open coding-agent platform; multi-agent |
| Aider | Git-native CLI pair programmer | Atomic commits; simple; great for ops scripts |
| SWE-agent / mini-SWE-agent | Research harness, clean ACI | Benchmark + security CTF modes |
| OpenCode | Open Claude Code alternative | High interest community |
| Cline | IDE agent | Editor-centric workflows |
| Closed CLIs | Claude Code, Codex, Gemini CLI | Quality bar for comparison |
SWE-bench: harness + model co-determine scores; Mini-SWE-Agent claims strong results in very little code — architecture lessons for Coulomb harnesses.
5.2 NetKingdom adjacency
| Coulomb piece | Relation to D |
|---|---|
| sand-boxer | Where agents run safely |
| activity-core | When work is due (this lab’s daily brief) |
| agent-harness / kaizen | Who executes with profiles |
| state-hub | What changed; task spawn completion |
| llm-connect | Which model backend |
Freedom Intelligence reports on harness tech; it does not replace sand-boxer.
5.3 Sandboxing & fleet themes to track
- Default-deny network + TTL sandboxes
- Credential routing (OpenBao / ops-warden) separate from model weights
- Multi-agent kanban / workplan-tied loops (Ralph-workplan class)
- Cost and rate-limit fleet management for API models
6. Collection recommendations (open-weight reserve)
6.1 Decision framework (recap)
- Prefer B/C leverage + clear license + stable org provenance
- Size gates per
inventory/collection-policy.md - Backup path must be pinned (FI-WP-0001-T04) before large pulls
- Quants: prefer official or well-known GGUF/AWQ only after base id is cataloged; v1 recommends safetensors official first, quant as companion later
6.2 P0 — seed now (approve after storage pin)
| # | Model (representative HF id) | Est. size class | Axes | Why |
|---|---|---|---|---|
| 1 | Qwen/Qwen3-8B (or current 8B-Instruct sibling on card) |
~16 GB fp16 / ~5 GB Q4 | B, C | Best default general + FT base; multilingual; tool-friendly |
| 2 | meta-llama/Llama-3.2-3B-Instruct |
~6 GB fp16 / ~2 GB Q4 | B | Tiny always-on / edge; huge ecosystem |
| 3 | BAAI/bge-m3 |
~2 GB | B | Multilingual RAG staple; MIT |
| 4 | deepseek-ai/DeepSeek-R1-Distill-Qwen-14B (or 8B if VRAM tight) |
~28 GB fp16 / ~9 GB Q4 | B, C | Local reasoning without full R1; MIT distill story |
| 5 | nomic-ai/nomic-embed-text-v1.5 (or current v2 text) |
<1 GB | B | Lightweight embed alternative / A/B with BGE |
Rough P0 cold storage (fp16 worst case): ~50–60 GB.
With Q4 preference for LLMs + fp for embeds: often ~20–30 GB. Prefer quants only when catalog notes the exact quant provenance.
6.3 P1 — next wave (quota permitting)
| # | Model | Why |
|---|---|---|
| 6 | Qwen/Qwen3-14B (instruct) |
Quality step for single-GPU chat/code |
| 7 | google/gemma-3-4b-it or EmbeddingGemma-300M |
Efficient Google line; embed edge |
| 8 | mistralai/Mistral-Small-3.x instruct (current card) |
EU-adjacent open option |
| 9 | deepseek-ai/DeepSeek-R1-Distill-Qwen-32B |
Stronger local reason if ≥24 GB VRAM class |
| 10 | Code embed: nomic-ai/nomic-embed-code or Jina code embed |
Code RAG for repos |
| 11 | Small reranker: BAAI/bge-reranker-v2-m3 |
Cheap quality win for RAG |
6.4 Strategic large (updated 2026-07-24)
Superseded by operator decision: 1 TB local NAS + capability-first reserve. Full DeepSeek-V3/R1-class and 70B dense opens are in scope even if unrunnable on current GPUs. See:
docs/decisions/2026-07-24-nas-strategic-reserve.mdresearch/2026-07-24-nas-strategic-collection-plan.md
| Item | Guidance |
|---|---|
| Full DeepSeek-V3 / R1 class | Strategic (S) — collect under 850 GiB soft quota; prefer official compressed if needed |
| Llama 70B / Qwen 72B class | Strategic (S) dense — good second fill after one primary MoE |
| Every new quant repack on HF | Skip; track base official only |
6.5 Explicit non-recommendations (for now)
- Closed API-only weights
- Anonymous “uncensored” reuploads without provenance
- Duplicate quants of the same revision
- Giant pretraining corpora as “models”
6.6 Catalog seed
YAML candidates for P0–P1 are under inventory/catalog/ with
status: candidate and brief_refs pointing at this survey. Approval and
download are FI-WP-0003.
7. Implications for the first daily brief
The first briefs/YYYY/MM/YYYY-MM-DD.md after this survey should:
- Not restate this whole document
- Cover only deltas since 2026-07-24 (new releases, price changes, harness versions)
- Revisit P0/P1 if a new open model clearly displaces a candidate
- Note any license or access shocks (takedowns, geo blocks)
- Optionally mark FI-WP-0001-T04 / FI-WP-0003 progress under Lab implications
Automation: activity-core def fi-daily-research-brief (disabled until FI-WP-0002
resolver + dry-run). Manual playbook: docs/daily-brief-playbook.md.
8. Open questions for later briefs
- Exact hardware envelope measured VRAM on lab hosts?
- Prefer GGUF Q5_K_M vs official safetensors for P0?
- Is German-first multilingual quality a hard requirement (affects embed + instruct choice)?
- Will local models serve air-gapped NetKingdom scenarios or only cost/privacy convenience?
- Single code model vs general instruct for domain FT?
9. Source anchors (non-exhaustive)
- DeepSeek R1 release / MIT notes — DeepSeek API docs & HF model cards
- Pricing roundups 2026 — IntuitionLabs, MorphLLM, provider pricing pages
- Embeddings 2026 comparisons — BentoML, Milvus/Zilliz, BGE-M3 cards
- Harness landscape — OpenHands, Aider, SWE-agent comparisons (LocalAI Master, vendor blogs)
- Open-weight commercial lists — community curations (verify each card)
Re-verify any number before spend or multi-hundred-GB download.
Baseline established 2026-07-24 for Freedom Intelligence. Daily briefs begin as deltas from here.