Open Weights July 22, 2026 ~22 min read Kimi K3 July 27

Kimi K3 Open Weights:
Release Date, Benchmarks & How to Run It

Five days to weight drop · Modified MIT · 2.8T MoE · 1M context · 1.4TB @ 4-bit · API $3/$15

Kimi K3 vs Claude Fable 5 open weights release benchmark comparison chart concept

Direct answer: Moonshot AI will release the full weights of Kimi K3—a 2.8-trillion-parameter MoE model with a 1,048,576-token context window—on July 27, 2026, under a Modified MIT license on Hugging Face, alongside a technical report. The API and Kimi product stack went live July 16. This guide covers timeline confusion, specs, benchmarks vs Claude Fable 5 and GPT-5.6 Sol, real API cost with caching, self-hosting hardware reality, industry impact, and a release-day checklist—without pretending you can run it on a laptop.

01

TL;DR: What Changes on July 27

  • July 16: Kimi K3 shipped via API, Kimi App, Kimi Work, Kimi Code; Artificial Analysis published an independent score (57.1).
  • July 17: WAIC opened in Shanghai; state media framed K3 as a national milestone.
  • July 27: Full open weights + technical report (architecture, training, evals).

Bottom line: K3 is not a “Fable 5 killer.” It ranks third on the AA Intelligence Index (57.1 vs 59.9 / 58.9), but delivers frontier-adjacent capability at roughly one-third the cost, plus two things closed models cannot offer: downloadable weights and a 1M-token context window.

Pain points before you commit

  1. 01

    Two launch dates: API live ≠ weights downloadable yet.

  2. 02

    “Run locally” clickbait: ~1.4 TB at 4-bit, 64+ accelerators—not Ollama-on-a-laptop territory.

  3. 03

    Closed-model premium: Fable 5 per-task cost is ~65.8% higher (AA measured K3 at $0.94).

  4. 04

    Long-context KV cost: Many models advertise long windows but cannot afford them in production.

  5. 05

    Regulatory off-switch: Closed weights can be pulled; published open weights cannot be “un-released.”

02

Kimi K3 Release Date: API vs Open Weights

DateEvent
2026-07-16API + Kimi stack live; Artificial Analysis benchmark
2026-07-17WAIC; national-milestone framing
2026-07-27Hugging Face weight download + technical report

Search intent peaks on kimi k3 release date, kimi k3 download, and kimi k3 hugging face around July 27. Update dateModified and swap “scheduled” copy to “now available” within hours of the drop.

03

What Is Kimi K3? Key Specs

SpecValue
Parameters2.8 trillion (largest open-weight release to date)
ArchitectureSparse MoE: Stable LatentMoE, 16 of 896 experts active per token
AttentionKimi Delta Attention (KDA) + AttnRes + Gated MLA
Context1,048,576 tokens
ModalitiesNative vision (text + image; video in product); text out
Weight formatMXFP4 weights + MXFP8 activations (native low-precision training)
Scaling vs K2~2.5× efficiency at equal compute
Long-context decodeUp to 6.3× faster at 1M tokens (KDA)
04

Kimi K3 Benchmarks vs Claude Fable 5 and GPT-5.6

ModelAA Intelligence Index
Claude Fable 559.9
GPT-5.6 Sol58.9
Kimi K357.1 (#3 of 189)

Where Kimi K3 wins

  • Frontend Code Arena #1 (blind dev preference over Fable and Sol)
  • Automation Bench, SpreadsheetBench 2 #1
  • BrowseComp 91.2 #1 (1M no-compression strategy ~90.4+)
  • SWE Marathon 42.0—long coding sessions (GPT-5.5 / GLM-5.2 collapse to teens)
  • Terminal Bench 2.1 88.3—near parity with Sol 88.8
  • Program Bench 77.8—edges Sol 77.6
  • FrontierSWE 81.2—well above Sol 71.3 (below Fable 86.6)

Where it still falls short

  • GDPval v2 Elo 1668–1687 vs Fable 1760 / Sol 1748
  • DeepSWE 67.5 vs Sol 73.0
  • Hallucination rate up vs K2.6 (Moonshot acknowledged); Reddit reports vs top closed models
  • Polish and session variance still trail Fable 5 / Sol

Moonshot unusually admitted trailing Fable 5 and Sol overall, citing harness sensitivity and over-eager behavior on ambiguous intent. Benchmarks ran through different harnesses (eval date: July 16, 2026)—treat as directional until independent reruns land.

05

Kimi K3 API Pricing (Real Cost With Caching)

ItemPrice per 1M tokens
Input (cache miss)$3.00
Input (cache hit, automatic)$0.30
Output$15.00
  • Pricing aligns with Western quality tiers—highest among Chinese vendors, still below Claude Opus 4.8 per-task cost.
  • AA measured $0.94 per task: 9.4% below GPT-5.6 Sol, 65.8% below Claude Fable 5.
  • Coding workloads report 90%+ cache hit rates—effective input cost far below list price.
06

Is Kimi K3 Open Source? Open Weights vs Open Source

Use open weights in titles and snippets—not “open source” alone. K3 releases weights + a technical report under Modified MIT; training code and data are not fully open. That distinction matters on r/LocalLLaMA and HN, and it is itself high-intent FAQ content for is kimi k3 open source.

On July 27 expect: Moonshot Hugging Face org upload, LICENSE text, vLLM KDA/prefix-cache support aligned with partners, and later Ollama/GGUF builds following the K2 pattern.

07

Can You Run Kimi K3 Locally? Hardware Requirements

Honest answer: not on consumer hardware.

  • ~1.4 TB at 4-bit; Moonshot recommends a 64+ accelerator super-node with expert + tensor parallelism.
  • Reference: K2.7 Code (1T) needed ~577 GB VRAM at INT4; K3 is 2.8× that scale.
  • Self-host only if you need data residency, fine-tuning, or API-scale spend where owned iron wins—otherwise use API ($3/$15) or cloud inference.

Release-day checklist

  1. 01

    HF repo completeness (shards, index, config)

  2. 02

    LICENSE verbatim (Modified MIT terms)

  3. 03

    Quantized artifacts (MXFP4 / GGUF)

  4. 04

    vLLM / SGLang launch commands and minimum hardware notes

  5. 05

    Independent long-context and BrowseComp reruns

08

Why the July 27 Drop Matters

  1. 01

    Frontier gap nearly closed: AA spread shrunk to a few points—harder to justify closed-only premium on raw capability.

  2. 02

    Geopolitics: WAIC timing; US briefly delisted Anthropic Fable/Mythos (restored July 1)—regulators can block closed API access, not already-published weights.

  3. 03

    Chinese wave: GLM-5.2, DeepSeek V4 Pro (1.6T), MiniMax—K3 is the peak of the open-weight push.

  4. 04

    Moonshot comeback: from post–DeepSeek R1 slump to K2 → K2.5 → K3 open-weight cadence.

Decision matrix

ScenarioPathWhy
Try now / coding agentsAPI or Kimi CodeLive since July 16; caching cuts cost
Post-7/27 fine-tune / residencySelf-host (64+ GPUs)Needs full weights + compliance
Cost-sensitiveDeepSeek V4 Pro API$3.48/M output
Long-horizon reasoningClaude Fable 5GDPval / overall index lead
Wait for Ollama/GGUFTrack K2 follow-onsConsumer quantizations lag the HF drop

Sources: kimi.com/en/blog/kimi-k3, platform.kimi.ai, Artificial Analysis (July 16), VentureBeat, Northflank, r/LocalLLaMA, Hacker News.

09

FAQ

July 27, 2026 on Hugging Face with a technical report. API shipped July 16.

Open weights under Modified MIT—not fully open source (no complete training stack/data release).

$3 / $0.30 cached / $15 per million tokens; ~$0.94 per task in AA testing.

No—~1.4 TB at 4-bit, 64+ accelerators recommended. Use the API.

Overall index: no (57.1 vs 59.9). Several coding/UI benches: yes. Cost per task: much lower.

Modified MIT—confirm exact terms in the Hugging Face LICENSE on July 27.

Closing

Open weights on July 27 do not magically shrink a 2.8T model onto your desk GPU. Most builders will consume K3 through API, Kimi Code, or OpenRouter; self-hosting remains an enterprise-scale bet. If your workflow also needs a real macOS GUI—Kimi Code agents, OpenClaw routing, Xcode signing dialogs SSH cannot click—buying hardware is expensive and idle nodes depreciate fast. VNCMac remote Mac rental gives you an on-demand VNC desktop to validate long-context coding and agent flows, then stop paying when the sprint ends. See Mac rental pricing or our July 17 Kimi K3 review for deeper architecture detail.