Summary
Moonshot AI’s Kimi K3 (released 2026-07-16) is the largest open model ever announced: a 2.8T-parameter Mixture-of-Experts (~A50B active) with a 1M-token context and a new “Kimi Delta Attention.” Unlike GLM 5.2 (clearly behind frontier, wins on cost), K3 trades blows with the top proprietary models on several benchmarks — but at real Sonnet-level pricing, and the weights + agentic behavior are unproven until the promised 2026-07-27 open-weight drop.
Key Claims
- Benchmarks trade blows with frontier (largely self-reported / early third-party, numbers vary by source):
- SWE Marathon (long sustained sessions) 42.0 — leads Opus 4.8 (40), GPT-5.6 Sol (39), Fable 5 (35)
- Program Bench 77.8 — tops field, edges GPT-5.6 Sol (77.6)
- Terminal Bench 2.1 88.3 — just behind GPT-5.6 Sol (88.8), ahead of Opus 4.8 / Fable 5 (84.6)
- FrontierSWE ~81 — beats GPT-5.6 Sol (~71), trails Fable 5 (~86)
- DeepSWE 67.5 — behind GPT-5.6 Sol (73), Fable 5 (70)
- Arena blind: preferred over every US model for front-end; overall Elo ~1547 (behind only Fable 5)
- Pricing 15 per M in/out — “Opus-class at Sonnet pricing.” ~3x jump from Kimi K2.6 (4) and ~2x GLM 5.2’s rate. NOT the rock-bottom-cost play GLM is.
- Weights not out yet — API/app/Playground only until 2026-07-27; everything now is trust-the-table.
My Reactions
Two caveats sink most of the hype for my use case per LLM Benchmark Reference:
- Simon Willison’s core critique: the marketing table “doesn’t touch the thing that matters most — agentic tool calling.” Strong static benchmarks != strong agent.
- Single reasoning level + heavy token consumption (~13k reasoning tokens for a simple SVG) erodes the price story on real workloads.
- Nearly all the headline benches (SWE Marathon, Program Bench, Terminal Bench, DeepSWE, FrontierSWE) are self-reported here — exactly what the reference note says not to update beliefs on until an independent run reproduces them.
Connections
- GLM 5.2 - Near-Frontier Open-Weight Coding Model — the other mid-2026 China open-weight release; GLM = value play, K3 = frontier-capability-at-real-price play
- LLM Benchmark Reference — most K3 headline benches are newer than that note; see the 2026 coding successors I added there
- LLM Comparison Sources — Arena Elo is the source class Karpathy/Willison distrust; wait for OpenRouter revealed preference + independent SWE-bench runs
- Dimensions of LLM Quality
Additional coverage: MarkTechPost, Latent Space AINews, Axios.