On July 16, a Beijing startup put a number on the table that made Wall Street flinch: 2.8 trillion parameters. Moonshot AI’s new Kimi K3 is the largest open-weight model ever released, and chip stocks slid into bear market territory within a day of the announcement. The last time a Chinese lab caused that kind of reaction, it was called DeepSeek.
So is K3 the real thing, or another panic that fades in a month? The answer matters for any business paying for AI right now, because this release changes the math.
What Moonshot Actually Announced
Kimi K3 is a massive language model with a 1-million-token context window, native visual understanding, and an always-on reasoning mode Moonshot calls “thinking mode.” Reports on the exact size vary slightly, with Bloomberg citing 2.7 trillion parameters and Tom’s Hardware and Hugging Face citing 2.8 trillion. Either way, it is the biggest model anyone has ever offered as open weights.
That open-weight part is the headline. Moonshot says the full weights land on July 27. Once they do, any company can download the model, run it on its own hardware, fine-tune it for its own data, and never pay Moonshot a licensing fee.
Founder and CEO Yang Zhilin has been building toward this since Kimi K2 caught developers’ attention last year. K3 is his claim to a seat at the frontier table.
The Benchmark Claims, and the Fine Print
Moonshot’s own charts show K3 beating Claude Opus 4.8 and GPT-5.5 on coding and agent benchmarks. The company reports 76.8% on SWE-bench Verified, 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, and first-place finishes on SWE Marathon and Program Bench. Tom’s Hardware notes it even beat Claude Fable 5, Anthropic’s current flagship, on the Frontend Code Arena benchmark.
Now the fine print. Moonshot itself admits K3 still trails OpenAI’s GPT-5.6 and Anthropic’s Fable 5 on overall capability and user experience. Several of the coding scores come from Moonshot’s own test harness rather than standardized third-party evaluations, so independent labs have not yet confirmed them. Benchmark wins also do not guarantee the judgment and stability you feel in a long working session.
The honest summary: K3 sits just behind the two American flagships and ahead of nearly everything else. For an open model, that has never happened before.

The Price Is the Real Story
K3’s API costs $3 per million input tokens and $15 per million output tokens, with cached input at $0.30. Compare that to roughly $50 per million output tokens for Claude Fable 5. That is a 70% discount for a model that wins some head-to-head coding benchmarks against models in that class.
One catch: K3 runs full reasoning on every request, and every thinking token bills as output. Short, simple tasks can cost more than you expect. The model earns its price on hard work like large-repo coding and long agent runs, where the 1M context and cache discount pay off.
It is not the cheapest Chinese model either. z.ai’s GLM-5.2 charges $4.40 per million output tokens and DeepSeek V4 charges $0.87. We broke down how DeepSeek V4 compares to US models on price and performance earlier this year, and K3 slots in as the premium option in that lineup: Chinese pricing, near-frontier capability.
What the Experts Are Saying
Reaction splits along a familiar line. Patrick Moorhead of Moor Insights called the market response “an over-reaction shockingly similar to the DeepSeek panic” and reminded everyone that “we are far away from super-intelligence.”
Builders are more impressed. Simon Koser, chief product officer at AI startup Tzafon, called K3 legitimately impressive on coding and said cost has become a deciding factor for AI labs choosing models. Analysts tracking open models say the gap between open weights and the closed Western frontier is now measured in months, not years.
The real test comes July 27. If the community can reproduce Moonshot’s numbers once the weights are public, the claims hold. If not, expect the skeptics to say so loudly.
What This Means for Small Businesses and E-commerce
You probably will not download a 2.8-trillion-parameter model. You do not have the hardware, and you do not need to. K3 still matters to you in three ways.
Your AI bills should fall
Every time a credible cheap alternative appears, the big labs respond on price. That happened after DeepSeek, and the pressure just returned. If you are running chatbots, product description generators, or support automation, this is a good month to renegotiate or re-shop. If you are on Claude today, start by cutting your Claude token usage before switching anything.
Your tools will quietly switch models
The apps you already pay for, from email writers to inventory forecasters, choose models based on cost per task. Many will test K3 the week the weights drop. You may get better output at the same subscription price without lifting a finger. It is also a reason to recheck what you pay for automation, because your vendors’ costs just dropped.
Agents keep getting better at commerce
K3’s strongest scores are on agent benchmarks, the tests that measure whether a model can complete multi-step tasks on its own. Cheaper, stronger agents accelerate the shift we covered when AI agents started buying directly inside ChatGPT. If machines are becoming your customers, the machines just got smarter and cheaper to run.
The Takeaway
Kimi K3 does not dethrone GPT-5.6 or Claude Fable 5, and Moonshot admits as much. What it does is prove the frontier is no longer a walled garden. A model you can download for free now trades benchmark wins with systems that cost three times as much to use.
For businesses, the lesson is simple: never sign a long AI contract in a market moving this fast. The best model in July may be the second-best bargain by August.








