id: moonshotai__Kimi-K3__strategic status: approved name: Kimi-K3 org: moonshotai source: kind: huggingface url: https://huggingface.co/moonshotai/Kimi-K3 revision: main model_card_url: https://huggingface.co/moonshotai/Kimi-K3 project_url: https://www.kimi.com/ paper_url: https://huggingface.co/papers/2607.24653 license: spdx: LicenseRef-Kimi-K3 url: https://huggingface.co/moonshotai/Kimi-K3 allows_offline_retention: true allows_local_ops: true allows_fine_tune: true notes: > Custom Kimi K3 License (not plain MIT/Apache). Review commercial SaaS revenue thresholds and attribution before external product embedding. Confirm card text at pin time. profile: summary: "S1 open frontier MoE (2.8T) — best open general/agentic class mid-2026." original_source: https://huggingface.co/moonshotai/Kimi-K3 use_cases: - "Long-horizon coding and agent workflows (offline optionality)" - "Strategic reserve if closed APIs are unavailable or restricted" - "Future multi-GPU / training-facility inference when hardware catches up" - "Capability baseline against closed frontier for daily briefs" - "Native vision + 1M-context knowledge work (when runnable)" sweet_spots: - "Top open-weight intelligence class after Jul 2026 weight release" - "Agentic / coding / long-context open MoE identity" - "Displaces prior S1 DeepSeek-V3 as primary open general target" not_ideal_for: - "Current single-GPU or small multi-GPU lab serve (far beyond run envelope)" - "Fitting under 850 GiB soft quota at full precision (~1.45 TiB tree)" - "Casual always-on chat (use R spine)" capability_notes: > Moonshot Kimi K3: ~2.8T total MoE parameters, sparse activation (~16/896 experts), Kimi Delta Attention + Attention Residuals, native multimodal, ~1M context. Official HF tree measured ~1454 GiB (96 shards). Collection deferred until capacity exception or official compressed distribution. swot: strengths: - "Leading open-weight general/agentic capability at release" - "Open weights enable true offline retention optionality" - "Strong coding / long-horizon / vision narrative" weaknesses: - "Disk footprint (~1.45 TiB) exceeds current soft quota and most lab GPUs" - "Custom license with commercial thresholds" - "Serve cost multi-node enterprise class" opportunities: - "Hold identity in catalog; pull compressed later if published" - "Anchor daily briefs / evals against open SOTA" - "Displace V3 as primary S1 when space allows" threats: - "Quota: full pull blocks all other S niches on VAULT" - "Rapid supersession by K3.x or peer open giants" - "Third-party quant quality variance if forced to unofficial packs" size: total_bytes: 1560998984390 total_human: "~1454 GiB official HF tree (measured 2026-08-03); compressed TBD" hardware_class: min_vram_gb_q4: 0 min_vram_gb_fp16: 0 notes: "Beyond current lab run envelope (multi-node). Strategic reserve only." axes: [A, B, C, D] priority: high collection: approved_by: "bernd" approved_at: "2026-08-03" downloaded_at: null downloaded_by: "" storage_path: "" brief_refs: - briefs/2026/08/2026-08-03.md - research/2026-07-24-nas-strategic-collection-plan.md - docs/decisions/2026-07-24-nas-strategic-reserve.md reason: > S1 strategic — best open-weight general/agentic model class after Jul 2026 release; keep capability identity even though full tree exceeds soft quota. tags: [tier-s, strategic, beyond-run-envelope, moe, general, vision, kimi, s1] companions: [] notes: > Do not bulk-download full tree under 850 GiB soft quota. Prefer official compressed if/when available. Revisit capacity exception with operator. history: - at: "2026-08-03" event: nominated by: grok detail: "Week catch-up brief — Kimi K3 open weights Jul 27 as new open SOTA." - at: "2026-08-03" event: approved by: bernd detail: "Operator direction: catalog and plan; collection deferred on capacity." - at: "2026-08-03" event: collection_deferred by: grok detail: "Official tree ~1454 GiB > soft quota 850 GiB; await capacity decision or compressed release." - at: "2026-08-03" event: profile_swot_added by: grok detail: "schema 0.2 profile + SWOT."