Enable fi-daily-research-brief, prove fi_brief_status idempotence with the 2026-07-28 brief, approve R/S catalog entries, collect and verify embeds plus Qwen3-8B and R1-Distill-14B on VAULT, and document residuals (HF-gated Llama, deferred S giants, railiance ConfigMap apply).
91 lines
3.3 KiB
YAML
91 lines
3.3 KiB
YAML
id: Qwen__Qwen3-72B__strategic
|
||
status: approved
|
||
name: Qwen3-72B
|
||
org: Qwen
|
||
source:
|
||
kind: huggingface
|
||
url: https://huggingface.co/Qwen/Qwen3-72B
|
||
revision: main
|
||
model_card_url: https://huggingface.co/Qwen/Qwen3-72B
|
||
project_url: https://qwenlm.github.io/
|
||
license:
|
||
spdx: Apache-2.0
|
||
url: https://huggingface.co/Qwen/Qwen3-72B
|
||
allows_offline_retention: true
|
||
allows_local_ops: true
|
||
allows_fine_tune: true
|
||
notes: "Confirm card; use Instruct/Coder sibling if that is the capability peak for the line."
|
||
profile:
|
||
summary: "Strategic large Qwen dense — multilingual/tool open 70B-class reserve."
|
||
original_source: https://huggingface.co/Qwen/Qwen3-72B
|
||
use_cases:
|
||
- "Future multi-GPU dense instruct without MoE serving complexity"
|
||
- "Strong multilingual (DE/EN) offline assistant at 70B scale"
|
||
- "Domain FT base when 8B/14B capacity is too small"
|
||
- "Dense alternative to MoE giants for portable lab stacks"
|
||
sweet_spots:
|
||
- "S4 dense open with Qwen tooling continuity from the R 8B default"
|
||
- "Easier multi-GPU dense serve story than 600B MoE"
|
||
- "Apache-friendly licensing typical of Qwen3 line"
|
||
not_ideal_for:
|
||
- "Current default single-GPU ops"
|
||
- "Collecting both this and Llama-70B-class first under tight quota"
|
||
- "Embedding / retrieval"
|
||
capability_notes: >
|
||
Large dense strategic pick. Prefer Instruct/Coder sibling if that is the line
|
||
peak at pin time. If soft quota is tight after S1, pick at most one of
|
||
Qwen-72B vs Llama-70B-class first.
|
||
swot:
|
||
strengths:
|
||
- "High multilingual dense quality in open 70B class"
|
||
- "Same family as R-tier Qwen3-8B — transfer of prompts/tools"
|
||
- "Usually simpler ops than giant MoE"
|
||
weaknesses:
|
||
- "~40–150 GiB depending on quant; still heavy"
|
||
- "May lag top MoE on some frontier-open benches"
|
||
opportunities:
|
||
- "First dense strategic fill after R + optional S1 compressed"
|
||
- "FT / LoRA at 70B when facility upgrades"
|
||
threats:
|
||
- "Llama 70B / next Qwen large may be better use of same bytes"
|
||
- "Supersession by Qwen next large release"
|
||
size:
|
||
total_bytes: 0
|
||
total_human: "~145 GB fp16 / ~40 GB Q4 (estimate)"
|
||
hardware_class:
|
||
min_vram_gb_q4: 40
|
||
min_vram_gb_fp16: 145
|
||
notes: "T3+ to run; strategic dense multilingual open."
|
||
axes: [A, B, C]
|
||
priority: high
|
||
collection:
|
||
approved_by: "bernd"
|
||
approved_at: "2026-07-28"
|
||
downloaded_at: null
|
||
downloaded_by: ""
|
||
storage_path: ""
|
||
brief_refs:
|
||
- research/2026-07-24-nas-strategic-collection-plan.md
|
||
- docs/decisions/2026-07-24-nas-strategic-reserve.md
|
||
reason: "S4 strategic — large Qwen dense open for multilingual/tool capability reserve."
|
||
tags: [tier-s, strategic, dense, qwen3]
|
||
companions:
|
||
- Qwen__Qwen3-8B__candidate
|
||
notes: "Pick at most one of Llama-70B-class vs Qwen-72B-class first if quota tight after S1 MoE."
|
||
history:
|
||
- at: "2026-07-24"
|
||
event: nominated
|
||
by: operator-policy
|
||
detail: "Strategic dense open for NAS."
|
||
- at: "2026-07-28"
|
||
event: profile_swot_added
|
||
by: grok
|
||
detail: "schema 0.2 profile + SWOT."
|
||
- at: "2026-07-28"
|
||
event: approved
|
||
by: bernd
|
||
detail: "S4 dense Qwen — deferred after R fill; pick vs Llama-70B"
|
||
- at: "2026-07-28"
|
||
event: collection_deferred
|
||
by: grok
|
||
detail: "S4 dense Qwen — deferred after R fill; pick vs Llama-70B"
|