China's fastest-rising AI lab says its new flagship trades blows with Claude and GPT. In a week, it plans to give the whole thing away.
Moonshot AI switched on Kimi K3 on July 16, pushing a 2.8 trillion parameter model to its apps and API, then demonstrated it a day later at the World Artificial Intelligence Conference in Shanghai. The Beijing startup calls it the most capable system it has ever shipped and claims benchmark parity with the American frontier. The full model weights are promised by July 27 under a permissive license.
That last detail is the one rattling the industry. Once those files go public, K3 becomes the largest open-weight AI model ever released, available for any developer or government on earth to download and adapt. Rival labs can dissect it. Enterprises can run it inside their own data centers, beyond any provider's terms of service. And no regulator, in Washington or Beijing, can order it switched off afterward.
Inside the 2.8 Trillion Parameter Machine
The headline number needs an immediate footnote: it describes capacity rather than the compute spent on each word. K3 is a sparse mixture-of-experts model that routes every token through 16 of its 896 expert subnetworks, under two percent of the total pool. The design extends the philosophy of its predecessor, Kimi K2.6, which packed one trillion parameters into roughly 32 billion active ones, and it is the trick that makes a model this size servable at all.
The rest of the spec sheet reads like a wish list. A context window of 1,048,576 tokens with no long-context surcharge. Native visual understanding, so the model can iterate against screenshots and rendered output. Reasoning that runs at maximum effort by default, since Moonshot ships no cheaper non-thinking variant. Underneath sits a reworked attention stack headlined by Kimi Delta Attention, which substitutes inexpensive linear attention through most of the network and reserves full attention for periodic layers, while a companion technique called Attention Residuals lets each layer reach back for useful representations from earlier ones. Moonshot claims roughly 2.5 times the scaling efficiency of the K2 generation, though the technical report defining that figure has yet to appear.
The company positions K3 as a long-horizon agent model, built for marathon software engineering sessions and sustained knowledge work with minimal human supervision.
The Benchmark Scorecard
Moonshot's own numbers place K3 in rare company. The company said the model "substantially outperformed" Anthropic's Claude Opus 4.8 along with both of OpenAI's current flagships, GPT 5.6 Sol and GPT 5.5, while performing competitively against Claude Fable 5, the strongest model in general release.
Independent testing broadly supports the claim while adding texture. Artificial Analysis scored K3 at 57 on its Intelligence Index, a composite spanning reasoning, knowledge, mathematics and coding, where comparable models average 31. That places the system among the top handful worldwide, though still short of Fable 5 overall. On Arena's crowd-judged leaderboards K3 took first place in frontend and interface engineering, beating Fable in blind human-preference matchups. One quirk surfaced along the way: the model is strikingly verbose, generating 130 million tokens during the Intelligence Index run against an average of 63 million, which pushed the cost of simply evaluating it past $2,700.
Then there is the price. API access costs $3 per million input tokens and $15 per million output tokens, flat across the entire context window, with cached input discounted 90 percent to $0.30. That undercuts Anthropic's Opus 4.8 by roughly half and sits far below Fable, whose output runs about $50 per million tokens. The framing flips inside China, where K3 counts as a premium product: z.ai's GLM-5.2 charges $4.40 per million output tokens and DeepSeek V4 just $0.87.
The Company Behind Kimi
Moonshot was founded in Beijing in 2023 by Yang Zhilin, a former Google researcher whose Kimi chatbot has grown into one of China's most popular consumer AI products. Annual recurring revenue passed $200 million in April on subscriptions and API usage. The investor list doubles as a map of Chinese tech power: Alibaba, Tencent, Meituan and HongShan all hold stakes, with IDG Capital and 5Y Capital also on the roster, and total funding stands near $3.8 billion. A $2 billion round in May valued the company above $20 billion, and reports say a fresh raise now under discussion would push that figure past $30 billion.
Its models have already crossed the Pacific. Cursor used Kimi to help build Composer 2, its coding agent. DoorDash's chief technology officer has said the delivery company hands lower-level work to Kimi K2.6, and Thinking Machines used K2.5 to generate early post-training data for Inkling, the open-weight model it released on July 15.
For Silicon Valley, the awkward part is how much of Silicon Valley already runs on Kimi.
A Release Washington Cannot Recall
The launch lands in the middle of the tensest stretch yet in the AI standoff between the United States and China. Export controls still bar Chinese labs from the most advanced accelerators, and although the Commerce Department began approving limited Nvidia H200 shipments this month under a licensing regime, a House Foreign Affairs Committee hearing on July 14 spent hours probing the gaps in that system. Bank of America analysts drew the uncomfortable conclusion from K3's debut: constrained chips have pushed Moonshot into training and architecture efficiencies that are now paying off.
The sharper contrast involves Anthropic. In June, days after the company launched Claude Fable 5, the US government invoked export-control authorities to bar foreign nationals from accessing Fable 5 and its less restricted sibling, Mythos 5, citing a reported technique for bypassing the model's cybersecurity guardrails. Because the order covered Anthropic's own foreign-national employees, the company disabled both models for every customer overnight. Anthropic publicly disputed the rationale, arguing the jailbreak was narrow and that comparable capability already existed in rival systems. Washington relented at the start of July after a review, restoring Fable 5 to general availability and Mythos 5 to a vetted partner program, with tighter safeguards attached.
The episode established that the US now treats frontier models as controllable national-security assets, software that can be ordered offline by directive. K3 is the counterexample. An open-weight release has no off switch, and Moonshot is about to hand the world its largest one.
Markets React Like It Is DeepSeek Again
Investors caught the reference immediately. Larry Dignan of Constellation Research wrote that "Kimi K3 may be the Deepseek Act II," recalling January 2025, when a cheap Chinese model erased nearly $600 billion of Nvidia's market value in a single session.
This time the heaviest damage hit closer to home. Shares of Moonshot's Hong Kong-listed domestic rivals sank on launch day, with Zhipu down about 28 percent and MiniMax off nearly 16 percent, as investors concluded the local race had a runaway leader. Semiconductor names on the Nasdaq slipped too, on the recurring worry that efficient Chinese models weaken the case for America's enormous compute build-out.
The pricing pressure lands on the closed labs as well. If a frontier-class open model can be self-hosted, every dollar of proprietary API spend gets renegotiated.
What July 27 Will Actually Settle
For all the noise, K3 remains a promise in one important respect. As of this weekend the weights are not downloadable; the model runs only through Moonshot's API and apps, plus OpenRouter for developers without a Moonshot account. The company says the files will publish by July 27 under a modified MIT license, and the exact terms deserve scrutiny when they land, along with the long-awaited technical report.
Self-hosting will be no small feat either. A 2.8 trillion parameter mixture-of-experts model demands serious hardware even with only 16 experts firing per token, which means the first wave of outside deployment will come from cloud providers and well-funded labs rather than hobbyists. Distilled and quantized offshoots will follow; they always do.
Moonshot, meanwhile, is clearing the runway. New users have already lost access to K2.5, and older models are marked for retirement at the end of August.
The open question is whether publishing the weights closes the gap with the American frontier or resets the whole race around a Chinese baseline. Seven days from now, anyone will be able to start finding out.
Comments 0
Join the discussion and share your perspective.
Sign in to post a comment and reply to other readers.
No comments yet
Be the first to share your perspective on this article.