AI Learned to bill like a law firm

Every metered business eventually learns it is charging for effort instead of result, and knowing has never been the same as changing

// Share
AI Learned to bill like a law firm

Consider the rate card, that most unglamorous of documents. Anthropic's Claude Opus 4.8 lists at $25 for a million tokens of output; OpenAI's GPT-5.5 at $30; xAI's Grok 4.5 at $6; and the cheapest of the open-weight models, the ones trained in Hangzhou and Paris and handed to anyone who wants them, at something nearer two dollars for the same million. The spread runs past tenfold, and around it the whole of enterprise software has arranged its buying the way a household on a budget arranges a supermarket run: find the cheapest unit that will do, and take that one. The unit, in this case, is the token, the small fragment of text a model reads and writes, and for two years the price of it has done nothing but fall.

There is a long history of trades that priced their work by a unit and then discovered, slowly and at some expense, that the unit was measuring the wrong thing.

The law firm is the cleanest case, and the most humbling. American lawyers built their business in the last century on the billable hour, a unit with the great merit of being easy to count and the quiet defect of paying the lawyer more the longer he took; a partner who drafted a contract in three hours earned less than a partner who produced the same contract, no better, in nine. Clients have complained about the arithmetic for forty years (a Silicon Valley general counsel told a legal newspaper in 2003 that "the days of blank checks are over") and pressed for flat fees, fixed prices, and the thing the profession still names with faint distaste, "value-based billing." They have mostly lost. Something close to 90% of legal spending in America still passes through the hourly meter, and the alternatives have hovered near a fifth of billings for years without climbing. Everyone involved has known for a generation that the unit measures effort rather than the brief. The unit has not cared.

The token is the billable hour of the intelligence business, and it is arriving at the same reckoning on a far shorter clock. What a company buys from a model is never a token but a finished task, a resolved support ticket or a merged pull request or a cleared invoice, and the tokens spent reaching one vary in a way the rate card cannot see. A modern agent does not answer and stop; it plans, consults a tool, reads the result, revises, and goes again, resending the whole of its accumulated context at every step, so that one task through one agent can burn thirty times the tokens depending only on how cleanly the model reasons its way to the end. Price per token is what the vendor prints. Tokens per task is what the buyer actually pays, and the widening space between the two is where the entire argument has quietly relocated.

Measured against finished tasks, the ranking came apart in the hand. Artificial Analysis, a firm that benchmarks models, ran a suite of real tasks through the current field and found Kimi K2 — the open-weight model from the Chinese lab Moonshot — finishing the average task for about ninety-five cents; a GPT-5.1-class model for a little over a dollar; and an Opus-class model, dearest of them all by the token, for roughly $2.75. The cheap model's per-token price was half the middle model's, and it spent close to twice the tokens getting to the same place, and the two facts, multiplied together, very nearly canceled. The bargain on the sticker was not a bargain on the invoice. It had never really been one.

The most telling evidence is not in the benchmark but in the conduct of the people selling the models, who have quietly stopped selling the meter. When OpenAI released GPT-5.6 and xAI released Grok 4.5 within days of each other in July, both led with frugality rather than the benchmark scores that had governed every launch before: OpenAI reporting that its model beat Opus on a computer-use test while spending, by its own count, 85% fewer output tokens; xAI reporting that Grok finished coding tasks on roughly a quarter of what Opus required. Anthropic, reaching the same destination from the pricing side, now discounts the tokens an agent spends rereading its own history by as much as 90%. When the men who own the meter begin apologizing for how fast it spins, the meter has stopped being the product.

The telegraph companies of the last century, which charged by the word, taught a generation to write in the clipped, articleless shorthand that came to be called cablese; then the price of a word fell, and the shorthand died with it, a small casualty of a change in a pricing table. The token will leave its own residue. What is worth noticing now, before the residue sets, is how much capital has already committed itself on the strength of the departing unit. Roughly $1.8 billion flowed into agent startups across a dozen-odd deals this past July, at valuations some 40% above the quarter before, nearly all of it underwritten on decks whose retention curves and gross margins are computed, at the base, in tokens.

None of this required anyone to behave badly, which is what makes it the ordinary kind of mistake rather than the interesting kind. An industry found a unit it could count, priced the most valuable commodity of the age against it, and mistook the ease of the counting for the truth of the measure. The lawyers show how long such a mistake can run: everyone can know the meter is measuring the wrong thing and keep feeding it for a generation. The token may prove stickier than its critics expect, and the billion-dollar decks built on it safer, for now, than the arithmetic implies. Arithmetic wins in the end. It is patient about when.

// The Daily

Get Vector in your inbox.

A free morning briefing on the AI revolution. Weekdays at 6am CT.