Summary

Moonshot AI’s Kimi K3 (released 2026-07-16) is the largest open model ever announced: a 2.8T-parameter Mixture-of-Experts (~A50B active) with a 1M-token context and a new “Kimi Delta Attention.” Unlike GLM 5.2 (clearly behind frontier, wins on cost), K3 trades blows with the top proprietary models on several benchmarks — but at real Sonnet-level pricing, and the weights + agentic behavior are unproven until the promised 2026-07-27 open-weight drop.

Key Claims

  • Benchmarks trade blows with frontier (largely self-reported / early third-party, numbers vary by source):
    • SWE Marathon (long sustained sessions) 42.0leads Opus 4.8 (40), GPT-5.6 Sol (39), Fable 5 (35)
    • Program Bench 77.8 — tops field, edges GPT-5.6 Sol (77.6)
    • Terminal Bench 2.1 88.3 — just behind GPT-5.6 Sol (88.8), ahead of Opus 4.8 / Fable 5 (84.6)
    • FrontierSWE ~81 — beats GPT-5.6 Sol (~71), trails Fable 5 (~86)
    • DeepSWE 67.5 — behind GPT-5.6 Sol (73), Fable 5 (70)
    • Arena blind: preferred over every US model for front-end; overall Elo ~1547 (behind only Fable 5)
  • Pricing 15 per M in/out — “Opus-class at Sonnet pricing.” ~3x jump from Kimi K2.6 (4) and ~2x GLM 5.2’s rate. NOT the rock-bottom-cost play GLM is.
  • Weights not out yet — API/app/Playground only until 2026-07-27; everything now is trust-the-table.

My Reactions

Two caveats sink most of the hype for my use case per LLM Benchmark Reference:

  • Simon Willison’s core critique: the marketing table “doesn’t touch the thing that matters most — agentic tool calling.” Strong static benchmarks != strong agent.
  • Single reasoning level + heavy token consumption (~13k reasoning tokens for a simple SVG) erodes the price story on real workloads.
  • Nearly all the headline benches (SWE Marathon, Program Bench, Terminal Bench, DeepSWE, FrontierSWE) are self-reported here — exactly what the reference note says not to update beliefs on until an independent run reproduces them.

Connections

Additional coverage: MarkTechPost, Latent Space AINews, Axios.