When Legora, one of the larger legal-AI platforms, recently moved its most capable product to usage-based pricing, the reaction inside firms was largely operational. How do we budget for tokens. Who owns the cost. Do we pass it on or absorb it. These are reasonable questions, and they miss the larger one.
The shift to usage-based AI is not really a procurement change. It is the return of an idea the profession has spent the better part of two decades trying to leave behind.
An old logic in new clothes
The billable hour has a single structural flaw that no amount of refinement ever fixed. It prices an input, the lawyer's time, rather than an output, the result the client actually wanted. For most of the profession's history this did not matter, because clients had no way to separate the two. They paid for hours because hours were the only thing on offer.
Then they stopped accepting it. The years since the financial crisis have been one long client revolt against paying for time: fixed fees, capped fees, alternative fee arrangements, the slow rise of value-based pricing. The argument never changed. Time spent is not the same as value delivered, and clients no longer wish to fund the difference.
Usage-based AI pricing is built on the very logic the profession has been trying to escape. You pay for tokens consumed. You pay for the input. The unit has changed from the hour to the token, but the thinking has not moved at all.
There is a fair objection to make here. Usage-based pricing actually repairs the billable hour's worst feature, the incentive to run the meter, because the firm now has every reason to spend less rather than more. That is true, and it matters. What it does not repair is the framing. It still trains the client to look at what was consumed rather than at what was delivered, and that habit was always the part that did the damage.
Usage-based AI is the billable hour reincarnated one layer down. The meter moved from the lawyer to the machine. What carried over was not the mechanics but the mistake beneath them, pricing the input and teaching the client to watch it.
Why can't a law firm mark up the cost of AI?
There is a deeper problem beneath the change of unit.
The billable hour was at least a defensible input. A client could not buy your senior associate's hour on the open market at cost, so the firm could attach a margin to it and the client had little choice but to accept. The input was scarce, and its true cost was opaque.
Compute is the opposite. It is abundant, and its price is roughly public. A token is a cost the firm pays, not a fee the client agreed, and a general counsel can open a pricing page in much the same range the firm sees. The billable hour blurred that distinction on purpose, turning the firm's input cost straight into the client's bill. Passing compute through as a line item repeats the move one layer down, except now the input's price is published, so the markup is visible from the first invoice.
The hour earned a margin because it was scarce and its cost was hidden. Compute is abundant and its cost is in plain sight. You cannot build a margin on an input the client can price for themselves.
Two firms, one announcement
Picture two firms reading the same pricing announcement.
The first is built on recurring, standardised work, priced more or less by the hour. It was already the firm most exposed to clients bringing routine work in-house. Now it faces a second squeeze. Its production costs have just become visible and variable. Its instinct will be to pass the compute through, which only invites the client to compare prices and to ask why the firm sits in the middle at all.
The second firm priced outcomes years ago. It charges for advice, for judgment, for carrying risk the client does not want to hold alone. To this firm, compute is a cost of goods, no more interesting than the cost of the coffee. The price the client sees never mentioned inputs, so a change in the cost of an input changes nothing.
Only one of these firms is troubled by usage-based pricing. The other stopped selling inputs long ago, and barely looks up.
Few firms are purely one or the other. Most are a blend, outcome-priced on the advisory work and hourly on the overflow. The useful question is not which firm you are today. It is which way you let the mix drift as compute gets cheaper.
The independent's quiet advantage
Most of the advice being written about this shift is written for large firms. It speaks of standing up a "TokenOps" function to manage consumption, of training lawyers to choose the right model for each task, of the internal politics of who gets which token budget. All of it is sensible. It is also a large institution building machinery to manage a cost that, until recently, did not exist, machinery it needs only because it intends to treat compute as a billable input in the first place.
The independent and the boutique have none of that machinery, and for once that is the advantage. There is no token politics to litigate and no legacy billing culture to unlearn. The independent can do the thing the commentators reserve for the AI-native newcomers: treat the tool as a small, lightly managed cost and price the outcome on top. The capability the large firm must spend heavily to control is a burden the independent can largely avoid carrying.
There is a sharper recognition here for anyone who left a big firm to build their own. The back-office machine you walked away from, and quietly missed in the first year, is the machine the large firm must now expand again. What once looked like the thing you gave up is turning into the thing you no longer have to defend.
At the layer of routine, standardised work, scale becomes a tax. The cost the large firm builds a department to manage is one the independent can mostly decline to carry.
What actually defends a law firm's fee?
Strip the question to its core, and one thing remains. As the cost of producing the standardisable parts of legal work falls towards the price of compute, what is left that a client cannot buy at cost?
Only the judgment. The decision about what to do, what is enough, what is too risky, and what the firm is prepared to stand behind. That has never been a function of how many hours or how many tokens it took to reach. It is a function of being trusted to decide, and trust does not have a published price.
None of this means a firm can simply declare itself a judgment business. A practice whose clients buy throughput, and chose it for price and capacity, cannot charge for judgment those clients never came for. Repositioning is a real fork, not a free one. But a firm living on an input markup should at least see it clearly for what it is, a borrowed margin on a clock that is running down.
This is why the firms that learned to hold a confident value conversation will find the next few years easier, not harder. The harder it becomes to justify a fee by pointing at effort, the more a firm must justify it by pointing at worth. AI does not retire that conversation. It makes it the only one that matters.
As production collapses towards the cost of compute, the only durable price is the price of being trusted to decide.
What it looks like from the outside
Viewed through an ownership lens, the distinction is sharper still. One firm's revenue is a markup on inputs, hours today and tokens tomorrow. The other's is a charge for judgment. The first margin is borrowed from a pricing model the market is steadily leaving. The second is the firm's own.
The difference will not appear in this year's accounts. It will appear the year compute pricing turns transparent to clients, the input markup quietly disappears, and one of the two firms discovers it was never really being paid for the thing it thought it was selling.
The invitation, again
The move to usage-based AI is being read as a problem to manage. Budget the tokens, police the usage, decide who pays. Read that way, it is a decade of new overhead.
Read properly, it is the same invitation the billable hour issued years ago and most firms declined: stop selling inputs. The firms that accept it this time will let the machine be cheap and charge for the judgment. The firms that decline it will spend the next few years building careful machinery to protect a margin that was never quite theirs to keep.
The billable hour did not die. It moved into the machine. The firms that already learned to price something else will not even notice it arrive.
Frequently asked questions
Why is usage-based legal AI pricing like the billable hour?
Both price an input rather than an output. The billable hour charged for the lawyer's time. Usage-based AI charges for tokens consumed. The unit has changed but the thinking has not, and it still trains the client to watch what was consumed rather than what was delivered, which was always the habit that did the damage.
Can a law firm pass AI compute costs on to clients with a markup?
Not for long. The hour earned a margin because it was scarce and its true cost was hidden. Compute is abundant and its price is roughly public, so a general counsel can open a pricing page in much the same range the firm sees. Passing compute through as a line item makes the markup visible from the first invoice.
Why do independent and boutique firms have an advantage here?
Large firms are building "TokenOps" machinery to manage consumption, model choice and token politics, all because they intend to treat compute as a billable input. Independents have no such machinery and no legacy billing culture to unlearn. They can treat the tool as a small, lightly managed cost and price the outcome on top. At the layer of routine work, scale becomes a tax.
What can a law firm still charge for when production costs fall?
Judgment. The decision about what to do, what is enough, what is too risky, and what the firm is prepared to stand behind. That was never a function of hours or tokens. It is a function of being trusted to decide, and trust has no published price. A firm living on an input markup is holding a borrowed margin on a clock that is running down.
Put a number on your own firm
The Pricing Power Score is a four-question, 60-second test. A score out of 100 and a figure, in pounds, on what holding your price would be worth. No email, no sign-up.
Run the Pricing Power Score