We're all just reselling tokens
Cursor lost money on every subscription, then sold for $60 billion — and the same spread runs through every layer of the AI economy
The Information reported in the spring that Anysphere, the company behind the Cursor code editor, had posted a gross margin of negative 23% in the quarter ended January, at a moment when its revenue was approaching $2 billion annualized. Five months later, on June 16th, SpaceX agreed to buy it for $60 billion in stock, the largest acquisition of a venture-backed startup on record; the deal closed on August 14th. Between those two facts sits the arithmetic of the whole AI application layer. Cursor sold subscriptions at $20 to $200 a month and spent, by most estimates, 40 to 70 cents of every revenue dollar buying inference from Anthropic and OpenAI, the same companies whose coding agents compete with it. It added a markup of roughly 20% to someone else's tokens and sold them inside an editor. For at least one quarter, the markup did not cover the tokens.
Cursor is the clearest case, but the structure it exposes does not stop at the application layer. Anthropic spent 71 cents of every revenue dollar on compute in the first quarter, according to a Yahoo Finance analysis of its financials, and projected 56 cents in the second. OpenAI told investors that its inference costs quadrupled in 2025, pushing its adjusted gross margin down to 33% from 40% the year before, per The Information; the figure recovered to roughly 39% in the first quarter of 2026. The compute that eats those cents is rented from Google, Amazon, Microsoft, Oracle and CoreWeave — Anthropic alone has committed $200 billion to Google Cloud over five years — and the clouds in turn buy their accelerators from Nvidia, which buys wafers from TSMC. Each layer sells the layer above it the same underlying unit, a token or the capacity to produce one, at a markup, and each layer keeps somewhere between 30 and 50 cents of the dollar. The labs are resellers of hyperscaler compute. The hyperscalers are resellers of Nvidia silicon. Only TSMC owns the thing that cannot be rented.
The chain extends past the last company, too. At Nvidia's GTC conference in March, Jensen Huang, the company's chief executive, said engineers should carry annual token budgets equal to half their salaries, about $250,000 worth. An engineer with a $250,000 token allowance is buying inference from an employer and reselling it as labor, with the employer reselling the output to customers as product. Huang was describing a demand curve; he was also describing the final link in a resale chain that begins at a fab in Hsinchu. Hence the title.
Long-distance charges
The pattern has a precedent, and the precedent has an ending. After the 1984 breakup of AT&T, a class of firms known as "switchless resellers" (they owned no network, only a billing system and a sales force) bought bulk long-distance capacity from AT&T at volume discounts and resold minutes to small businesses at prices below AT&T's own retail tariff. By the early 1990s there were hundreds of them. Their margin was the spread between the wholesale rate and the retail rate, and both rates were set by the carriers. Through the decade AT&T, MCI and Sprint cut retail prices, the spread narrowed, and the resellers sorted themselves into three groups: those that bought switches and became facilities-based carriers, those that sold themselves to carriers, and those that vanished. Excel Communications, the largest of the pure resellers, built a business with more than $1 billion in annual revenue selling long distance through network marketing and sold itself to Teleglobe, a Canadian carrier, in 1998. The mobile version of the same story ran in 2023, when T-Mobile paid up to $1.35 billion for Mint Mobile, a reseller of T-Mobile's own network.
The AI resale spread is currently wide, and it is wide for the same reason the long-distance spread was wide in 1987. The wholesale price of a token is collapsing. Stanford's AI Index found that the cost of inference at GPT-3.5-level performance fell roughly 280-fold between November 2022 and October 2024; OpenAI has said its inference cost per query is down about 95% since GPT-4 launched. Retail prices — a $20 monthly subscription, a per-seat enterprise contract, a $200 Ultra plan — fall much more slowly, because they are set by what a customer will pay rather than by what a GPU costs. SemiAnalysis, a research firm, estimates that Anthropic's gross margin on inference rose from 38% to 70% in a year. That is the reseller's spread widening because wholesale is falling faster than retail. It is also, as every switchless reseller learned, a spread the wholesaler can close whenever it decides to.
The trouble at the application layer is that volume moved before price did. OpenAI's API went from 6 billion tokens a minute in October 2025 to 15 billion by April; Google said at I/O in May that its APIs were processing about 19 billion tokens a minute and that 375 customers had each consumed more than a trillion tokens in the past year; OpenRouter, a routing marketplace, reported 25 trillion tokens a week in May, up from 5 trillion six months earlier. Agentic workloads are the reason. A single user action in a coding agent triggers five to 20 sequential model calls, so per-user consumption rises faster than per-token cost falls, and a flat-rate reseller — Cursor Pro at $20 — finds itself short tokens on every heavy user. Cursor's negative quarter was the flat-rate tariff meeting the agent. So was GitHub Copilot's move to usage-based billing on June 1st. Usage pricing hands the price signal back down the chain to the wholesaler, which is what the telecom resellers were eventually forced to do as well. ICONIQ, a growth investor, surveys about 300 AI-product companies twice a year; its January report put average gross margins for AI-native products at 52% for 2026, up from 41% in 2024 and still some 30 points below the software median. The spread is improving. It is not becoming software.
Which leaves the two exits, and both point at the network. The first is downward: own the facility. Anthropic agreed to buy roughly $21 billion of Google's tensor processing units and is building its own capacity on top of the cloud commitments; Cursor shipped Composer, an in-house model, in November 2025, and by early 2026 was routing about half its autocomplete requests through it, later admitting that Composer 2 was trained on Moonshot's open Kimi base. That is the switchless reseller buying a switch. The SpaceX acquisition is the other version of the same move — a carrier buying the reseller. SpaceX owns xAI's Colossus cluster and, since the August close, the developer relationship Cursor built with a million paying engineers; the docs now list SpaceXAI as a model provider beside OpenAI, Anthropic and Google. T-Mobile bought Mint for its subscribers, not its technology. SpaceX bought Cursor for the same reason.
The second exit is upward: own the customer and the meter. Enterprise customers generate three to five times the revenue per token that consumers do, by Forbes's estimate, and their workloads are more predictable and cheaper to serve. Cursor's enterprise accounts reached gross-margin profitability in the spring while its individual accounts stayed underwater, which is why $2.6 billion of its $4 billion in annualized revenue now comes from enterprise contracts. Anthropic draws about 85% of its revenue from enterprise and developer customers, booked $11.5 billion in the second quarter, and reported positive adjusted operating income — the first frontier lab to do so. OpenAI, whose revenue is roughly 85% consumer and whose users are roughly 95% non-paying, filed its confidential S-1 on June 8th with a gross margin in the high 30s. The two labs are resellers of the same input at the same wholesale price. The difference is who they resell to.
For an allocator, the question about any AI company is therefore not whether it is a "wrapper," a word that has stopped carrying information. The question is where in the chain the company sits and which direction it is moving. Every layer is squeezed from both sides: the labs by hyperscaler rents below and by their own resellers above, the clouds by Nvidia below and the labs' custom silicon above, the apps by lab pricing below and by the labs' own agents above. A reseller whose wholesale cost falls faster than its retail price is a fine business until the wholesaler notices, and the wholesalers have noticed. Anthropic and OpenAI both ship coding agents into Cursor's market. Google priced Gemini 3.5 Flash in May explicitly to pull enterprise workloads off frontier models, with Sundar Pichai, Alphabet's boss, telling customers they could save more than $1 billion a year by moving 80% of their traffic to it. AT&T competed with its resellers, too.
The chain will consolidate the way the last one did, toward whoever owns a facility or whoever owns a customer, and the companies with neither will be absorbed or priced out as the spread closes. What is different this time is the bottom of the chain. There is one TSMC, and it is not for sale. And what is different at the top is Huang's engineer, buying $250,000 of tokens a year and reselling judgment to an employer — the last reseller in the chain, and the only one who cannot buy a switch.