The Unaudited Trillion

No one can independently tell you whether enterprise AI pays off, because every headline ROI number in the debate was produced by someone with a position in the answer. That missing measurement, an independent, recurring, instrument-grade reading of value realized against AI spend, is the most valuable vacant position in the field.

The most cited number in enterprise AI is that 95 percent of corporate generative-AI pilots deliver no measurable return. It comes from a mid-2025 report out of MIT's NANDA initiative, and it has traveled into board decks, earnings calls, short theses, and at least one central-bank financial-stability review. It also rests on 52 interviews and 153 survey responses gathered at four conferences over six months. The second most cited number arrived three weeks ago and points the other way. A fine-tuned open model that Bridgewater built with Thinking Machines scored 84.7 percent on a battery of financial-reasoning tasks, beating the best frontier model while running at roughly a fourteenth of the cost. That figure comes from Bridgewater's own evaluation. No independent party has reproduced it.

Sit with those two numbers, because they define the terrain. The bear case and the bull case for the largest capital reallocation in enterprise software both rest on figures the interested party produced and nobody audited. That is not a temporary gap or an accident of a young field. It is the structural condition of the whole discourse, and I have started calling it the self-report economy: a market for claims in which every load-bearing statistic about whether AI pays off was generated by someone with a position in the answer. Vendors publish numbers. Consultancies publish numbers. Venture firms, model labs, and the enterprises buying the models all publish numbers. None of them is disinterested, and there is no scorekeeper. The most valuable unoccupied position in AI right now is not a model architecture or an agent protocol. It is the referee's chair, and it is empty.

The launch that turned a research argument into a spending fight

Thinking Machines released Inkling on July 15. It is a Mixture-of-Experts model with 975 billion total parameters and 41 billion active, a million-token context window, multimodal across text, image, and audio, shipped under a permissive Apache 2.0 license with the weights on Hugging Face. The company was unusually plain about what the model is not. In its own materials it states that Inkling is not the strongest overall model available, open or closed. That is a strange thing for a lab to say on launch day, and it is the most important sentence in the release.

The pitch is not benchmark supremacy. The pitch is that Inkling is a good base to make your own. It ships alongside Tinker, the company's fine-tuning platform, and a smaller preview model, and the whole package is aimed at an enterprise that wants to post-train a capable open model on its private data and run something specialized rather than rent a general one. Deployment is not trivial. Running the full-precision checkpoint needs on the order of two terabytes of aggregated GPU memory, which is a cluster of eight of Nvidia's newest accelerators or sixteen of the prior generation, though a quantized version cuts that by more than half. The model is available for inference across Together AI, Fireworks, Modal, Databricks, and Baseten, so a firm that wants to try before it builds can.

That positioning is not a marketing choice. It is the direct product expression of a pattern this site named some time ago as Model Convergence Pressure : as raw capability across the frontier converges, the marginal advantage stops coming from the base model and moves to what sits on top, the topology, the rubrics, the memory, the scaffolding, and above all the proprietary data you fine-tune with. Inkling is Model Convergence Pressure sold as a SKU. Mira Murati's company is betting that enterprises will care less about renting the single smartest general model and more about owning one they can shape to their own work. The bet only makes sense in a world where the base models have converged enough that "good enough and mine" beats "best and rented" for a meaningful slice of workloads. Two years ago that world did not exist. It does now, which is why the launch matters more than the model.

For most of the last two years the custom-versus-API question was a research-blog debate about fine-tuning efficacy, argued by people who read arXiv for fun. Inkling, and the Bridgewater result that preceded it by two weeks, turned it into a boardroom capital-allocation question with nine-figure stakes. Should a large enterprise route its agentic workloads to a frontier API from OpenAI, Anthropic, or Google, paying per token for the best available intelligence and accepting the dependency? Or should it license an open-weight base, fine-tune it on its own data, run it in its own environment, and own the result along with the cost and the maintenance burden? The honest answer, which almost nobody selling into this market will give you cleanly, is that it depends on four variables. The people offering you a clean answer are usually the ones with the most to gain from the answer they are offering.

The Bridgewater number deserves a much closer look than it got

The 84.7 percent has been passed around as if it settled something. Read the actual research, published at the end of June by Bridgewater's AI research group with Thinking Machines, and it is more interesting and more limited than the headline.

The task was not "beat the frontier at finance" in general. It was a specific, narrow, high-stakes job: replicating the judgment Bridgewater's own analysts apply when triaging financial information across six defined task types. The team set 80 percent accuracy as the threshold below which a model cannot be trusted in a daily workflow, a bar chosen by the people who would have to live with the output. Naive prompting of a strong general model scored around 50 percent, which is a coin flip and useless. The best hand-engineered expert prompts got into the mid-to-high 70s, better but still under the trust line. The fine-tuned model cleared it, at 84.7 percent against 78.2 percent for the strongest frontier model they tested, and it did so at roughly a fourteenth of the per-task inference cost. On the numbers Bridgewater reports, the fine-tuned model made about 30 percent fewer errors than the best general model while costing dramatically less to run at volume.

That is a real result, and it is exactly the kind of result the ownership camp needs. It is also a first-party evaluation on a benchmark the same team constructed, measuring a task defined by the same firm that stood to benefit from the finding, published in collaboration with the vendor whose model and platform were being validated. Bridgewater is not a Thinking Machines investor, which removes one obvious conflict and leaves several others in place. Nobody outside the collaboration has run the evaluation. The tasks are proprietary, so nobody outside can. One of the sharper independent write-ups of the work made the point that the deliverable here is not really the model at all, it is the auditable process that produced it, and the process is what Bridgewater is actually selling the market on. The number is a claim. A good one, from serious people, pointing at something probably real. Still a claim, and the difference between a claim and a measurement is the entire subject of this piece.

Apply the same skepticism to the company doing the launching, because the pattern runs all the way down. Thinking Machines raised $2 billion in seed funding at a $12 billion valuation in the summer of 2025, before it had shipped a product, in a round led by Andreessen Horowitz with Nvidia among the backers. A reported follow-on that would have valued the company near $50 billion stalled and fell apart around the turn of the year, and several senior researchers left, including at least one to Meta. The company has an undisclosed revenue figure, one fine-tuning API in market, a large Nvidia compute partnership, and a multibillion-dollar cloud deal with Google. A $12 billion valuation resting on one API and a research reputation deserves the same discount as a hedge fund's internal benchmark. I am not saying the company is not worth it. I am saying that its worth, like everything else in this market, is a number produced by parties with a position in it, and that you should notice how many of those numbers you have been quoting without a haircut.

The chief executives arguing this out are all talking their book

Watch who is loudest. Then watch what they sell.

Nikesh Arora, the chief executive of Palo Alto Networks, went on CNBC in early July and argued that token prices have to fall roughly 20 percent within a year and something like 90 percent within two before enterprise AI economics work at depth. His reasoning has real content. He points out that more than half of AI compute today serves money-losing free consumer usage, that enterprises adopt AI easily in narrow domains like coding and then stall when they need memory, context, and guardrails, and that a growing share of token volume is already routing to open-weight models through services like OpenRouter and not coming back. All of that is worth taking seriously. Arora also runs a company that sells enterprise security and benefits directly if AI inference gets cheaper and more commoditized. His analysis and his interests point the same direction. That does not make him wrong. It does mean his number is not a neutral reading off an instrument.

Alex Karp, the chief executive of Palantir, has been blunter. His line, delivered on the same network, is that something has gone completely wrong with the rent-a-token model, that enterprises pay, get little value, and hand over their intellectual property in the process, and that "the jig is up." Palantir stock rose about 9 percent on the appearance. Palantir also sells the alternative he is describing, an ontology layer wrapped around sovereign models a company controls, running on its own choice of open-weight base. One financial outlet put the tension precisely, noting that his numbers are real but so is his motive. Chamath Palihapitiya has made a similar case with a harder figure, claiming his firm found open-source on private infrastructure more than sixteen times cheaper than a frontier API for real workloads, with a specific example putting a month of one frontier model's usage above a hundred thousand dollars against a few thousand for an open alternative. He is an investor in that thesis.

The most developed version of the ownership argument came from Satya Nadella. In a mid-July essay he called the Reverse Information Paradox, he argued that a firm renting frontier intelligence pays for it twice, once in money and again in the proprietary knowledge it must reveal to make that intelligence useful. His framing is that the prompts, corrections, and evaluations a company feeds a model constitute an intelligence exhaust that leaks enterprise value to the model vendor. The model is the commodity, he argues, and the loop is the intellectual property. It is a genuinely good argument, grounded in Kenneth Arrow's 1962 observation about the paradox of disclosure in markets for information, and it resolves into a prescription that enterprises need a trust boundary they control. Microsoft sells Azure, Foundry, and exactly that trust boundary. He made the same point more bluntly at Davos in January, warning that firms risk leaking enterprise value to some model company somewhere. Every serious voice in this fight is at once an analyst and a merchant, and the analysis is downstream of the merchandise more often than anyone admits.

I do not raise the conflicts to dismiss the arguments. Arora, Karp, and Nadella are all partly right, which is what makes them worth reading. I raise the conflicts because they are the whole problem in miniature. When the people with the sharpest incentives are also the only ones producing the evidence, the evidence needs a discount before you can use it. Call it the self-report discount: the mental haircut you apply to any figure whose source has a position in the outcome. You already do this when a company reports its own adjusted EBITDA, or when a restaurant rates itself five stars. The enterprise-AI discourse has not built the reflex. That is why a hedge fund's internal benchmark, a venture firm's portfolio survey, and a software vendor's commissioned case study all get quoted as if they were readings off a shared instrument, when what they are is testimony from interested witnesses with no cross-examination.

Why the fight is happening now and not two years ago

The custom-versus-API question could not have been a live boardroom debate in 2023, and the reasons it can be one now also tell you where it goes next.

Three things changed at once. The base models converged, which is Model Convergence Pressure again. When the third-best model is 95 percent as capable as the best one for your task, the case for paying a premium to rent the best one weakens and the case for owning the third-best one strengthens. Fine-tuning got cheap and boring, moving from a research project to a Tuesday, as low-rank adaptation methods let a team specialize an open model for the price of a used car rather than a research grant. And a credible supply of open-weight bases finally appeared. For most of 2024 the strongest open models were Chinese, the DeepSeek and Qwen and Kimi families, which created a procurement problem for Western enterprises facing regulatory and reputational friction around Chinese-origin AI. Inkling's most underrated feature is its passport. A large, capable, US-built open-weight model under a permissive license is precisely the thing a regulated Western bank or hospital could not get from the open-weight frontier a year ago, and the market for it is real regardless of whether the ROI math works.

The passport point deserves more weight than it usually gets, because it is where the ROI question collides with something ROI cannot capture. Through 2024 and into 2025 the strongest open-weight models mostly carried Chinese origins, and for a regulated Western enterprise that origin is a hard constraint, sometimes a legal one, regardless of how the benchmarks read. A US bank cannot run its compliance workflow on a model it is barred from deploying for procurement and reputational reasons, however capable that model looks per dollar. So the real menu for a Western enterprise that wanted to own its stack stayed thin: capped models with restrictive licenses on one side, geopolitically awkward open weights on the other. Inkling widens the menu. A large, capable, permissively licensed, US-built open-weight model is a genuinely new option for exactly the buyers with the strongest control requirements. That is not a coincidence. The workloads that most want to own their model are the regulated, high-stakes ones, and those are precisely the workloads where nobody can cleanly measure whether the AI paid off, which is the whole problem compressed into a single sentence.

Two of this site's other frames sit directly underneath the fight. Inkling ships with a controllable thinking-effort dial, a setting a developer can move to trade accuracy against token spend, which is Reasoning as Billing Axis made concrete: test-time compute as a knob the buyer turns, a billing tier, and a research objective at the same time. And the ownership camp's strongest non-cost argument, the one Karp and Nadella both press, is really an appeal to Compliance as Differentiation: control over the model, the data, and the provenance becomes a regulatory and trust moat rather than merely a cost line, which is worth paying for in exactly the regulated industries where the ROI is hardest to measure. The fight is new because the technology and the geopolitics only recently made both options real for a serious buyer. It stays unresolved because nobody measures the thing that would resolve it.

The evidence everyone cites is broken in the same four ways

Here is the part that should bother anyone who has quoted an enterprise-AI statistic in the last year. Walk through the studies that anchor the discourse and the same structural flaws keep recurring. The source has a conflict. The method is a self-reported survey. The finding sits behind a paywall. Or the whole thing is a one-off with no reproducible methodology and no recurring panel. Most of the canonical numbers carry more than one of these at once, and the flaws point in directions you can predict from who paid for the work.

Start with the one that spread furthest. The MIT NANDA report from the middle of 2025, built around the idea of a generative-AI divide, produced the 95-percent-of-pilots-fail statistic that launched a thousand LinkedIn posts. Read the method and the confidence drains out of the headline. The finding rests on a review of a few hundred publicly disclosed initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected at four industry conferences over the first half of the year. Fifty-two interviews is not nothing. It is also not a representative sample of the global enterprise, and the study defined failure narrowly, as the absence of rapid, measurable profit-and-loss impact, which quietly excludes the long tail of value that shows up as avoided cost, faster cycle time, or a capability a firm simply did not have before. The report itself carried the nuance that outcomes depended heavily on approach. The number that escaped into the world did not. It mutated into "95 percent of AI fails," a claim the study does not make, and it is now the single most abused statistic in the field. It would be the first thing I retire.

McKinsey's State of AI, published late in 2025 off a survey of nearly two thousand respondents across more than a hundred countries, is a more serious instrument, and it lands in a similar place from the opposite direction. It finds that 88 percent of organizations use AI somewhere while only around 39 percent report enterprise-level earnings impact, with the high performers who actually capture value making up roughly 6 percent of respondents. The rhyme between that 6 percent and MIT's 5 percent is doing a lot of unearned work in the discourse, because the two numbers get cited together as mutual corroboration despite coming from different methods on different samples. Two survey instruments agreeing is not a fact. And McKinsey sells the workflow-redesign engagement that its own report prescribes as the cure, which means the diagnosis and the treatment come from the same shop. That is not a knock on the analysis. It is a reason to weight it accordingly.

Menlo Ventures publishes the most quoted spending figures, and they are useful, and they come from a venture firm. Its 2025 enterprise survey put generative-AI software spend at 37 billion dollars for the year, more than triple the prior year, with roughly 47 percent of pilots reaching production, far above the historical rate for ordinary SaaS. Good data, real signal, worth reading. Menlo is also an investor in Anthropic, and the report leads with the framing that this is a boom rather than a bubble, which is the framing an investor in the space would prefer the market to adopt. Sequoia, Bessemer, and Andreessen Horowitz all run comparable surveys with comparable tilts. When the people forecasting a market are long the market, the forecast is an act of advocacy in the costume of research, and the costume is convincing because the underlying numbers are often right.

Then there is Forrester's Total Economic Impact, which deserves special attention because it is the purest specimen of the pattern. TEI is a rigorous-looking framework, a four-part model of benefits, costs, flexibility, and risk built around a composite organization, and it produces the clean triple-digit ROI percentages that fill vendor case studies. It is also, by Forrester's own labeling, commissioned. The vendor whose product is being evaluated pays for the study. Forrester maintains, reasonably, that clients cannot buy a favorable conclusion and that TEI work is walled off from its syndicated research, and I take the firm at its word on process. The structure is still what it is: a paid analysis of the payer's product. Forrester's own surveys find that a large majority of technology buyers distrust vendor-produced content, which is the exact category a commissioned study occupies. The most polished ROI numbers in enterprise software carry the most direct financial relationship between the subject and the author, and buyers sense it even when they cannot name it.

The proxy-data players are better and still not clean. Ramp's AI Index, drawn from card and bill-pay data across more than seventy thousand businesses, is one of the more honest instruments in the field because it measures behavior rather than opinion, actual dollars spent rather than an executive's self-description. It reports adoption more than twice as high as the government's own count, which tells you how much self-report distorts the surveys. Ramp is also a fintech marketing its proprietary data, its sample skews toward the tech-forward companies that use Ramp, and the firm acknowledges the skew. The US Census Bureau's Business Trends survey is the least conflicted source in the entire landscape, a nationally representative government instrument, and it is close to useless for this debate because it is blunt, it lags, and the agency changed its AI question in late 2025 in a way that broke the time series just as the question got interesting. Gartner, for its part, put generative AI in the trough of its hype cycle two years running, pegged the average project near two million dollars with most chief executives dissatisfied, and forecast well over a trillion dollars of AI-infrastructure spend, all of it behind a subscription and shaped in part by the vendor briefings that fund the firm.

There is one genuinely honest attempt to name the problem from inside the consulting world, and it proves the rule by how alone it is. Tom Davenport and Laks Srinivasan run something they call the Return on AI Institute, and Davenport's long-running column distinguishes strategic AI returns, which are deep and narrow, from tactical ones, which are broad and shallow, and has criticized the NANDA methodology directly and by name. It is the closest thing to a neutral voice in the field. It is also a named-consultant operation with services to sell, which is not a criticism so much as an observation about how empty the truly independent corner is. When the most disinterested analysis available still comes from someone with a consulting practice attached, the referee's chair is not merely unoccupied. It is a chair nobody has thought to build.

The pattern, once you see it, is hard to unsee. Every incumbent source of enterprise-AI ROI data is compromised on at least one axis, and the compromises run in predictable directions. Vendors and their commissioned analysts skew high. Skeptics with short positions and consultancies selling remediation skew toward failure. Venture firms skew bullish. The one genuinely neutral source is too crude to answer the question. There is no independent, reproducible, recurring measurement of whether enterprise AI actually pays off, which is a remarkable sentence to be able to write truthfully about a technology absorbing hundreds of billions of dollars a year.

What the rigorous evidence actually shows is a jagged frontier, not a verdict

Leave the survey-industrial complex and walk into the peer-reviewed and working-paper literature, and the picture does not resolve into a clean answer. It resolves into a frontier so uneven that anyone claiming a single headline number for AI's productivity effect is either selling something or has not read the papers.

The strongest positive result is Brynjolfsson, Li, and Raymond's study of a Fortune 500 customer-support operation, more than five thousand agents, published through the National Bureau of Economic Research. Access to a generative-AI assistant raised issues resolved per hour by roughly 14 percent on average, and by around 34 percent for the least experienced agents, while barely moving the most experienced ones. That is a real deployment, a large sample, and a well-identified effect. It is also one firm and one narrow task, novices getting a leg up on a routine support workflow, and it does not travel to a portfolio manager or a litigator or a chip designer without a great deal of hand-waving. Noy and Zhang found something similar for mid-level writing tasks, faster work at higher rated quality, and again the tasks were bounded and the workers were not experts operating at the edge of a hard domain.

The most important study in the whole literature is the one that complicates the good news, and it comes from BCG's own researchers. Dell'Acqua and colleagues ran a large experiment with several hundred BCG consultants and found that inside what they named the jagged technological frontier, on tasks the model was suited to, consultants using AI did more work, did it faster, and produced higher-rated output, with gains north of 20 percent on speed and 40 percent on quality. On tasks that fell outside that frontier, the same consultants using the same AI did materially worse than the ones without it, on the order of nineteen points worse on getting to a correct answer, because the model produced confident, plausible, wrong work and the humans trusted it. That is the finding to internalize. AI does not raise or lower productivity as a scalar. It raises it steeply inside a boundary and lowers it inside a boundary, and the boundary is invisible and moves, which means the average effect for any real firm depends entirely on how much of its work sits on which side of a line nobody can see clearly.

The strongest cautionary result is methodologically the most honest thing in the field, and it cuts against the vendors hardest. METR, a nonprofit that runs controlled evaluations, ran a randomized trial in 2025 with experienced open-source developers working on real tasks in their own repositories. The developers expected AI assistance to speed them up by about 24 percent. After the study, they estimated it had sped them up by about 20 percent. The measured result was that AI assistance slowed them down by roughly 19 percent. Seasoned engineers, working on code they knew intimately, were worse off with the tool and could not tell they were worse off. That gap between perceived and actual productivity is the single most consequential finding in the entire debate, because nearly every survey number in the self-report economy is a measurement of perception, and here is a randomized trial showing perception was off by close to forty points in the wrong direction. METR later reported that a follow-up gave a noisier signal and redesigned the experiment, which is what rigor looks like and what you never see from a commissioned ROI study.

The developer-tools case is where this perception gap does the most damage, because coding is the domain everyone points to as AI's settled win. The proof is thinner than the enthusiasm. Large studies of real developers using AI assistants report a wide range of effects depending on task and seniority, from meaningful speedups on boilerplate and unfamiliar languages to measurable slowdowns on complex work in code the developer already knows cold, and the developers themselves are poor judges of which one they are living through in the moment. Deployments at the scale of a major software company show real gains in accepted code on some teams and negligible movement on others, and the aggregate keeps refusing to resolve into the clean figure the marketing wants. Coding is the strongest case, and even the strongest case is contested. If the domain with the clearest wins cannot produce an uncontested productivity number after two years of intense measurement, that should temper anyone's confidence about the harder domains, the legal review and the underwriting and the strategy work, where the frontier is jaggedest and the perception gap is widest. The self-report economy is not a failure of diligence. It is what you get when the thing being measured genuinely resists measurement, and every party doing the measuring has a reason to prefer a particular result.

At the macro scale, Daron Acemoglu has argued that AI adds at most something like two-thirds of a percentage point to total factor productivity over a decade, a rounding error against the trillions in forecast value from Goldman and others. He may be too pessimistic. The point is the spread. Serious economists modeling the same technology land anywhere between negligible and civilization-altering, which is not a range you can average into a business case. All of it rides on top of the oldest pattern in technology economics, the productivity paradox Robert Solow named in 1987 when he observed that you could see the computer age everywhere except in the productivity statistics, and the J-curve Brynjolfsson later formalized to explain it. Value from a general-purpose technology shows up in the numbers years after the investment, because the investment has to be paired with slow, expensive, unglamorous organizational change before it pays anything back. If that pattern holds for AI, and there is no strong reason to think it will not, then the failure studies and the triumph studies are both premature. They are measuring a J-curve near its bottom and mistaking the dip for the destination.

The one place the evidence is close to decisive is the exact question Inkling and Bridgewater put on the table, whether building a domain-specific model beats using a general one, and the historical answer is a cautionary tale the ownership camp would rather skip. BloombergGPT was the prestige case for domain pretraining, a fifty-billion-parameter model trained on a proprietary financial corpus at a cost of something like ten million dollars and well over a million GPU-hours. Within months, GPT-4, with no finance-specific training at all, beat it on most public financial benchmarks. A separate open-source effort then replicated much of BloombergGPT's public-benchmark value through lightweight fine-tuning for a few hundred dollars. The lesson landed hard: pretraining a domain model from scratch is a bet against the frontier, and the frontier moves faster than your model can pay itself back.

The lesson has a second half the skeptics skip, and it is precisely the crack Bridgewater is driving a truck through. BloombergGPT still beat GPT-4 on Bloomberg's own internal tasks, the ones GPT-4 had never seen and could not have been trained on. That is the whole game. The frontier wins on anything public and general. The custom model can win on the genuinely proprietary and specific, the tasks that live inside one firm's data and nowhere else on the internet. Bridgewater's 84.7 percent, if it survives independent replication, is a claim of exactly that shape: not that its fine-tuned model is smarter in general, but that it captures a specific, proprietary, hard-to-articulate judgment about financial information that no general model has any way to access. Whether that is a real edge or an artifact of a benchmark the same team designed is the question an independent referee would settle, and it is the question no independent referee exists to settle. The evidence, honestly read, does not crown either side. It says the answer is jagged, and that the shape of the jag depends on facts about your workload that you have to measure and that nobody is measuring for you.

The economics are contingent, and the contingencies are knowable

Strip out the rhetoric and the real total-cost comparison between building and renting is not mysterious. It is a function of a handful of variables, and the debate feels unresolvable only because the loudest participants keep asserting their preferred point on the curve as if it were the whole curve.

Fine-tuning itself is cheap now, far cheaper than most executives believe. Adapting a small-to-mid-size open model with low-rank methods runs from a couple of thousand dollars to a few tens of thousands in compute and data preparation, not the six or seven figures people still assume from the era of training from scratch. Full fine-tunes rarely earn their cost over the cheaper adaptation techniques. The expensive parts of ownership live elsewhere, and they are the parts nobody puts on the slide. Data collection, cleaning, and annotation is weeks of skilled labor before a single training run. Building an evaluation harness that can actually tell whether the fine-tuned model is better than the API is its own project with its own headcount. Security review, compliance sign-off, and the standing cost of retraining when the base model updates or the underlying data drifts are recurring, not one-time. There is a specific trap here worth naming, which is that a fine-tuned model bakes in the world as it was on training day, so the moment your product catalog or your policy or your regulatory environment changes, the knowledge you paid to embed is quietly wrong, and you are back in the training loop. The naive build case counts the adaptation run and forgets all of this. The naive rent case counts the per-token price and forgets that a well-tuned small model can internalize a policy or a house style, shorten every prompt, and cut inference spend by more than half at scale.

The crossover is real and estimable. Practitioner analyses converge on a break-even somewhere between a hundred thousand and a million model calls a month. Below it, the API almost always wins on total cost, because you are not amortizing a fixed investment across enough usage. Above it, a fine-tuned small model on owned or rented infrastructure starts to pull ahead, sometimes by a wide margin. An insurance team fine-tuning a small open model for contract-clause extraction moved accuracy from the high seventies to the mid-nineties while cutting per-inference cost by more than an order of magnitude against a frontier API call, which is the ownership case working exactly as advertised: high volume, a narrow and stable task, proprietary training data, a clear accuracy target. Change any one of those conditions and the arithmetic inverts. And there is a subtlety that keeps moving the line, which is that prompt caching on the frontier APIs now bills repeated context at a fraction of the standard input price, which quietly pushes the break-even back toward renting for any workload that reuses the same long context over and over.

Then there is the variable that keeps moving under everyone's feet, and it is the quiet assassin of the build case. The Stanford AI Index documented that the inference cost of a system performing at the level of GPT-3.5 fell more than 280-fold between late 2022 and late 2024, roughly from twenty dollars per million tokens to seven cents. Frontier output prices fell more modestly over the same window, something like fourfold, but the good-enough tier fell off a cliff. When you self-host, you lock in today's hardware economics and today's model, while the API you fled keeps getting cheaper and smarter underneath you every quarter. BloombergGPT did not fail because it was a bad model. It failed because it was a fixed point in a moving field, and by the time it could have paid itself back, renting had become both cheaper and more capable than the thing it built. Anyone signing a nine-figure custom-model deal in 2026 is betting that their workload is stable enough, private enough, and high-volume enough that ownership beats a rental price that has been falling by an order of magnitude roughly every eighteen months. For some workloads that bet is clearly right. For most it is clearly wrong. The whole skill is telling which is which, and that skill is worth more than any single model, because it is the thing that does not depreciate.

The market, for what it is worth, is voting with its checkbook against the ownership rhetoric even as the rhetoric gets louder. The buyers are not listening to the keynotes. Menlo's own data shows the enterprise pendulum swinging hard toward buying rather than building, roughly three-quarters bought against one-quarter built in 2025, a sharp move from a near-even split the year before. That does not settle the argument, because the crowd is often wrong at inflection points and the whole ownership case is that the crowd has not yet priced in the intelligence exhaust and the lock-in. It does mean that the people actually spending the money are, in aggregate, still choosing to rent, which is a data point the ownership evangelists tend to leave out of the keynote.

There is a longer historical rhythm underneath all of this that should keep everyone humble. Enterprises have run this build-versus-buy argument before, with mainframes, with databases, with the migration from on-premise servers to the cloud, and the pendulum swings on a recognizable schedule. Early in a platform shift, capability is scarce and expensive, so the sophisticated players build in-house because nothing off the shelf is good enough for them. As the platform matures, vendors commoditize the capability, the in-house builds curdle into expensive liabilities, and the pendulum swings hard toward buying. Then the cost of the commodity layer collapses far enough that a new frontier of custom advantage opens on top of it, and it swings back again. We are somewhere in the middle of that swing for AI, which is the deep reason both camps sound right at once. The frontier labs are selling the commodity-layer maturity that is genuinely arriving. The ownership camp is selling the new custom frontier that is genuinely opening on top of it. Neither is lying. They are describing different phases of the same pendulum, and each is betting you will mistake its phase for the whole arc.

There is a cost on the rental side that belongs in every one of these calculations and appears in almost none of them. Call it the verification tax: the labor a firm spends checking, correcting, and defending AI output before it can be used for anything that matters. A study out of BetterUp Labs with Stanford's Social Media Lab, published in Harvard Business Review, put a number on one slice of it. Across more than a thousand desk workers, 41 percent reported receiving what the researchers called workslop, plausible-looking AI output that shifts the burden of actually finishing the work onto whoever receives it, at a cost of nearly two hours per instance and an estimated hidden drag of about 186 dollars per employee per month. Roll that across a ten-thousand-person company and it is a nine-figure annual tax that never appears in the ROI model, because it is smeared across everyone's day as friction rather than concentrated in a line item anyone owns. The verification tax applies whether you build or rent, but it scales with how much unreliable output you generate, and it is the most underweighted number in enterprise AI. A productivity study that measures output but not the downstream cost of verifying that output is measuring half a ledger and calling it a result. The METR developers were paying the verification tax in real time and could not see it. Neither can the surveys.

What actually decides build versus rent

I have said twice that the answer depends on four variables. Here they are, because a claim that hides its own criteria is only an opinion in a lab coat, and the point of this piece is to replace opinion with something a reader can check against their own situation.

The first is volume. Below the break-even, which sits somewhere between a hundred thousand and a million calls a month for most workloads, renting wins outright, because you never amortize the fixed cost of building across enough usage to matter. Above it, ownership starts to pay. Most enterprise workloads sit below the line, and most of their owners are convinced they sit above it.

The second is task stability. A fine-tuned model is a photograph of the world on training day. If your task and your data and your rules hold still for a year, the photograph stays accurate and the investment compounds quietly in your favor. If they churn every quarter, you are paying to re-shoot the photograph on a loop, and the retraining bill eats the inference savings while you are looking the other way. Stable narrow tasks reward building. Anything that moves rewards renting.

The third is the data moat, and it is the only variable that can override the first two. Do you hold genuinely proprietary data that encodes a judgment the frontier models have never seen and cannot buy? Bridgewater does. Most companies do not, though most companies believe otherwise. Sitting on a large pile of documents is not a data moat. A data moat is a body of labeled, structured, hard-won judgment that lives inside your walls and nowhere on the public internet, and if you have one, fine-tuning is how you turn it into a product nobody can copy. If you do not, you are paying to make a general model marginally more like a general model, for no durable edge.

The fourth is control, and it is the one the ROI math cannot price. Some workloads carry a regulatory or competitive requirement to keep the model, the data, and the loop inside a boundary the firm controls, whatever the token arithmetic says. A bank that cannot legally send customer records to a third-party API, or a company that treats its prompt-and-correction loop as core intellectual property in the way Nadella describes, will pay a premium to own and should. The thing being bought there is not efficiency. It is sovereignty. And sovereignty never shows up in a per-token comparison, which is why the ownership case is strongest exactly where the ROI case is hardest to make.

Line those four up and the decision falls out of them. High volume, stable task, a real data moat, a control requirement: build, and Inkling with Tinker is aimed precisely at you. Low volume, a churning task, no proprietary data, no control constraint: rent, and skip the ownership keynotes. The uncomfortable fact for the ownership camp is that the second profile describes most enterprises. The uncomfortable fact for the frontier labs is that the first profile describes the highest-value workloads in finance, healthcare, and defense, which is exactly where the nine-figure deals are being signed. Both camps are right about a different quadrant of the same map, and both are selling their quadrant as the whole thing.

The measurement problem is one this field has already solved a layer up

Here is the structural insight behind my confidence. The referee's chair is empty, and it is also buildable, because this is not a new kind of problem. The field has watched it play out one layer up already, in model benchmarks, and the way it got handled there is the template for how it gets handled here.

This site has a name for what happened to model evaluation: Benchmark Contamination . SWE-bench, GAIA, WebArena, the standard yardsticks all saturated and got gamed the moment they became the thing everyone optimized against, and the response was not to publish a better static benchmark. It was to relocate measurement to places that could not be gamed as easily, live trajectory monitoring, dynamic professional-task evaluations, held-out data the model builders never see. The lesson was general: any metric that an interested party can both produce and be judged by will be corrupted, and the fix is to move measurement to something observable and independent of the party being measured.

Enterprise-AI ROI is Benchmark Contamination wearing a suit. The metrics that matter, adoption and value realized and payback, are being produced by the parties who are also being judged by them, and they are corrupted in exactly the directions the theory predicts. Vendors report high. Skeptics report low. Everyone reports their own number and cites it as ground truth. The fix is the same fix. You do not ask a company whether its AI paid off, because the answer is testimony from an interested witness. You measure the shadow the spending casts on the world, in card-spend data, in hiring and layoff patterns, in productivity statistics, in the disciplined language firms use about AI on earnings calls when a securities lawyer has vetted the sentence, in cloud consumption, and you build the reading from proxies the measured party cannot author.

There is deep precedent for this, and none of it comes from inside the industry being measured. The most trusted economic indicators in the world were built by outsiders who found a clean proxy and published it on a fixed schedule. The Economist created the Big Mac Index in 1986, using the price of one homogeneous good to read currency misvaluation, and it became a canonical reference precisely because a newspaper with no position in currency markets published a simple, transparent, repeatable number that anyone could check. The Baltic Dry Index tracks the cost of moving raw materials by sea and is prized for carrying almost no speculative content, because nobody books a freighter without cargo to move, so the number is expensive to fake. ADP built an independent read on US employment from the payroll records of tens of millions of workers, released ahead of the government's own figure and watched precisely because it is not the government's figure. Two economists at MIT and Harvard launched the Billion Prices Project in 2008, scraping millions of online prices a day, and used it to demonstrate that Argentina's official inflation numbers were fiction, then spun the method into a commercial data product a major financial-services firm now operates. The pattern is identical every time. An outsider. A clean observable proxy. A transparent method. A fixed cadence. Free or near-free access. None of them asked the interested party to grade its own homework.

The craft of the incumbents worth borrowing sits alongside the craft worth avoiding. The consulting flagships, McKinsey's global institute and BCG's Henderson institute and Gartner, know how to manufacture authority, and some of their moves are worth stealing: a named framework that becomes shorthand, the growth-share matrix and the hype cycle and the three horizons, an annual cadence people plan around, one big sizing number, one killer exhibit, a methodology appendix that signals rigor. What they get wrong is the two things that matter most here, which are conflict and access. They sell the remediation, and they paywall the finding. The independent analysts who actually won durable seats got both right. Ben Thompson built Stratechery on a single named theory, aggregation, and priced it for anyone rather than for procurement departments. SemiAnalysis won the semiconductor beat with proprietary supply-chain models and contrarian calls that kept being right. Epoch AI and Artificial Analysis won trust in model evaluation by being transparent, reproducible, and free, which is the entire reason people cite them over vendor benchmarks. The State of AI Report became an annual fixture by being comprehensive, free, and unafraid to make calls. The lesson across all of them is a short checklist: a named metric, an annual cadence, a transparent and reproducible method, free access, one defensible contrarian claim, and one chart people cannot stop sharing.

That is the opening, and it is precise. Not another survey. Not another commissioned case study. Not another venture firm's annual bull note dressed as research. An independent, transparently methodologized, freely published, recurring index of whether enterprise AI is actually paying off, built from observable proxies rather than self-report, run on the outsider's playbook rather than the consultancy's. The Big Mac Index for AI value. I think it is the most defensible unclaimed position in the field, and I think the window to claim it closes inside a year, because the absence is now obvious enough that someone will notice it and move.

What the index would measure and why the name has to stick

An index needs a headline number, and the number needs a name that can become shorthand, because shorthand is how a metric escapes its creator and enters the language. The indicators that lasted are the ones that turned into nouns people use without remembering who coined them.

The metric I would build the index around is the Value Realization Rate, the VRR: realized, attributable profit-and-loss impact divided by total AI spend, measured across a defined population over a defined period, and published with the method fully exposed. The phrase value realization already drifts around the consultancy vocabulary, used loosely and never measured independently, which is exactly why it is worth taking, defining sharply, and anchoring to a public methodology before someone uses it to mean nothing. The point of the VRR is that it forces back together the two things the self-report economy keeps carefully apart. Vendors report spend and adoption. Skeptics report failure rates. Almost nobody divides realized value by total cost and publishes the quotient on a transparent basis, because doing it honestly means admitting how hard attribution is, and admitting how hard attribution is undermines every tidy number in the market, including the ones the incumbents are selling.

Attribution is the hard part, and pretending otherwise is how the incumbents get their clean figures. The counterfactual problem is real. You cannot cleanly separate what the AI did from what the reorganization around the AI did, or from what the market did that quarter anyway. The time-horizon problem is the J-curve. Measure too early and every deployment looks like a failure, because the organizational change that unlocks the value has not happened yet, which is exactly the trap the 95-percent statistic fell into. The soft-benefit problem is that avoided cost and new capability do not land in the same ledger as headcount reduction, so they get dropped. A serious index does not solve these by ignoring them. It solves them the way the Billion Prices Project solved official-inflation doubt, by triangulating enough independent observable proxies that the reading holds up even though no single proxy is clean. Census adoption data corrects vendor optimism downward. Hiring and layoff data from public filings and labor trackers reads the employment effect directly rather than asking about it. Earnings-call language, mined at scale, reveals what firms are willing to tell investors about AI value under legal liability, which is a far more disciplined claim than what they tell an anonymous survey. Card and cloud-spend data read the actual dollars flowing. None of these can be authored by the company being measured, which is the entire point, and which not one incumbent index can say about its own inputs.

Then apply the discipline the whole field is missing, and make the discipline part of the published product. Every first-party number that enters the index carries the self-report discount, an explicit, stated haircut on any figure whose source has a position in the outcome, the way a credit analyst discounts a company's own adjusted metrics before using them. Every ROI figure the index publishes carries the verification tax as a named deduction, so the cost of checking AI output stops vanishing into distributed friction and starts showing up in the arithmetic where it belongs. And the treatment of the Bridgewater result becomes the model for the whole approach: it enters the index not as a fact but as a first-party claim at full self-report discount, and it stays there until an independent party reproduces it, at which point the discount comes off and it becomes the first credibly audited custom-model win, which would be real news rather than a press release. That asymmetry, holding a claim as a claim until it is verified and then upgrading it when it is, is what the referee's chair is for. It is the one thing no vendor and no consultancy and no venture firm can credibly offer, because each of them is a player in the game it would be scoring.

Picture the single chart this index would lead with, because I think it is the most clarifying image in the entire debate. Put every headline enterprise-AI number in circulation on one plot. On the horizontal axis, methodological rigor: sample size, whether the method is reproducible, whether it measures behavior or opinion. On the vertical axis, the funder's conflict of interest. Size each point by how often the number gets cited. What appears is a scatter with a dense, loud cluster in the low-rigor, high-conflict corner, where the most-cited numbers live, MIT's 95 percent and Menlo's spending figures and the vendor case studies all crowded together, large because everyone quotes them and positioned exactly where a skeptic would predict from who paid for them. The high-rigor, low-conflict corner, where an instrument-grade independent reading would sit, is empty. That empty quadrant is the vacant seat, drawn to scale. The argument of this whole piece compresses into that picture: the most-cited numbers in enterprise AI are inversely correlated with their own trustworthiness, and the trustworthy corner has nobody in it.

Building the number is only half the work. The other half is distribution, and distribution in 2026 has a shape that did not exist when the Big Mac Index was born. A metric now has two audiences, the humans who cite it and the reasoning engines that answer questions with it, and the engines gain weight every quarter. A chief technology officer asking an AI assistant whether enterprise AI pays off gets whatever number is most citable, which today means the assistant reaches for the dead 95-percent statistic, because that is the figure the web has canonized. Whoever publishes the trustworthy replacement on an open, machine-readable, transparently sourced basis captures those queries, and captured queries compound, because the third time an engine surfaces your index as the answer, your index becomes the default answer. The incumbents cannot win this race. Their best numbers sit behind paywalls the engines cannot read, and they carry conflicts the engines are already learning to flag.

Timing is the last piece, and it is not comfortable for anyone who likes to move slowly. The seat is open because the failure of the existing numbers became common knowledge only in the last few months, as the 95-percent statistic buckled under scrutiny and the launch-week ROI claims stacked up on top of each other. Common knowledge of a vacancy is the exact condition under which someone fills it. The Big Mac Index worked because a newspaper got there first and made the name stick before anyone else thought to. The window here is measured in months. The cost of waiting is not that the idea gets worse. It is that someone else runs it, and there is only one chair.

None of this is easy, and the honest version admits the ways it could fail. The proxies could prove too noisy to yield a stable signal, in which case the index degrades gracefully into something still valuable, a meta-rater that scores everyone else's ROI studies on rigor and conflict and tells you which numbers to trust, which is a defensible seat even if the data moat underperforms. The neutrality is the whole asset, which means a single sponsored placement or a single vendor deal would poison it, so the governance has to forbid vendor sponsorship structurally and visibly from day one. And an index is only as durable as its cadence, because a one-time report is an article and a recurring one is an institution. The Big Mac Index is forty years old. That is the bar. But the difficulty is the moat. If it were easy, an incumbent with more resources would already have done it, and the reason none of them has is not capability. It is that none of them can, because the value of the thing is precisely that its author has no stake in the answer, and every one of them has a stake in the answer.

Where this goes

The capital riding on the unmeasured question keeps climbing, which is what makes the empty chair worth more by the month. The bets keep getting bigger. Enterprise generative-AI software spend hit 37 billion dollars last year on Menlo's count, and the four largest cloud providers have guided to something on the order of 725 billion dollars of capital expenditure in 2026, most of it AI infrastructure, a number the Financial Times assembled from their own earnings guidance and which represents a jump of more than 70 percent in a single year. Michael Burry is publicly short the trade, arguing the hyperscalers are understating depreciation by assigning long useful lives to chips that stay current for two or three years, and reviving Galbraith's word for the undiscovered embezzlement that inflates a boom, the bezzle, to describe the gap between reported and real. Analysts at Apollo have noted that something approaching ninety percent of venture funding now flows to AI, a concentration that dwarfs even the dot-com peak, when the internet drew less than half of the money chasing new companies. The Federal Reserve has begun listing AI as a systemic-stability concern.

The structure of the spending is what worries the skeptics most, and it rhymes with prior manias in a way worth naming out loud. The money moves in a circle. Nvidia invests in the model labs, the labs commit to the cloud providers, the cloud providers buy Nvidia chips, and a wave of specialized compute financiers borrows against the whole arrangement, so a single dollar of demand can appear several times across several sets of books before an end customer has paid for anything real. It is the pattern that turned the late-nineties telecom buildout into the Lucent and Nortel collapse, vendor financing inflating apparent demand until the apparent demand turned out to be the vendors financing each other. One large bond manager has estimated that capital expenditure now consumes the overwhelming majority of the hyperscalers' operating cash flow, which is a brittle place to stand if the returns arrive late.

The bulls have a real answer, and it is not nothing. The cloud-AI revenue is actually showing up in the accounts. The largest providers report AI-driven cloud businesses growing at triple-digit rates off already enormous bases, tens of billions of dollars of run-rate that did not exist two years ago. That is not a bezzle. That is revenue. But revenue at the infrastructure layer is not the same thing as return at the enterprise layer, and the entire bull case rests on blurring the two. Nvidia's revenue proves that enterprises are spending on AI. It does not prove that those enterprises are getting their money back. The picks-and-shovels layer can boom for years on the strength of prospectors who never strike gold, and the only way to know which is happening is to measure the prospectors' returns, which is the one number nobody publishes.

Every one of those positions, long and short, top of the market and bottom, is a wager on a return that no independent party measures. Trillions of dollars are being allocated against a number that does not exist.

That is the opportunity. It is why the referee's chair is worth more than another model in a market about to drown in models. The base models are converging and the differentiation is migrating to the layer above, exactly as Model Convergence Pressure predicts, which means the scarce thing is not intelligence. Intelligence is getting cheaper by an order of magnitude every eighteen months, the one number in this whole piece that everyone actually agrees on. The scarce thing is a trustworthy reading of what all that intelligence is worth, and every reading on the market today was written by someone holding a position in the answer. The self-report economy will not correct itself, because the parties who could correct it are the parties who profit from its fog. The fog is the business model. Someone from outside has to walk in, pick a clean proxy, publish a transparent method on a fixed schedule, and start keeping score. The first person to do that credibly does not win the enterprise-AI argument. They become the scoreboard the argument is settled on, and in a self-report economy, the only position that compounds is the one nobody can pay to occupy.