Open weights, frontier price
Moonshot is giving away a frontier model and charging frontier prices to run it, which tells you exactly what it thinks it's selling
Moonshot AI, the Beijing lab behind the Kimi assistant (founded 2023, last valued above $20bn in a May round backed by Alibaba and Tencent), shipped Kimi K3 on July 16th and called it the largest open-weight model ever built. It is: 2.8 trillion parameters, fourth on the Artificial Analysis intelligence index at 57.1, a hair above Anthropic's Opus 4.8. Then it did two things that sit oddly next to the word "open." It throttled access within 48 hours, telling users demand had pushed "close to the limits of our current capacity" and that paying customers came first — the notice any hosted business issues, open weights or not. And it priced itself at $3 per million input tokens and $15 per million output — call it $5.40 blended at a standard 80/20 mix — roughly triple the previous Kimi, and only about 40% under the American frontier (Opus 4.8 and GPT-5.5 both sit near $9 blended, on identical $5/$25 rate cards) rather than the 90%-plus discounts that have defined nearly every prior Chinese release.
The weights themselves don't drop until July 27th. So for eleven days the largest open-weight model in the world is a hosted product you can rent and cannot run, sold by a company rationing it and charging near-frontier prices for the privilege. Looks like a contradiction. It's a business model, a well-worn one, and the interesting question for anyone sizing Moonshot isn't how good K3 is. It's which open-source playbook this becomes.
Start with what "free weights" actually cost to run, because that is the number that quietly decides everything else. Kimi K3 is a very large model, and the awkward thing about a very large model is that you have to keep all of it loaded in expensive memory even though only a fraction of it does any work on a given request. Holding it in place takes something like fifty of Nvidia's H100 chips — a cluster that runs on the order of a million dollars a year before anyone has paid for networking, redundancy, or the engineers to keep it alive. Set that fixed bill against K3's own hosted service, which rents the identical model for about $5.40 per million tokens, and the free weights only start to win once you are running something like half a billion tokens a day. Even that flatters the self-hosting case, since it assumes you keep every one of those chips busy around the clock, which essentially no one manages. In the real world the crossover sits higher still.
This article is for Vector members. Start a 7-day free trial to keep reading.
Start your free trial