What Can the Models Actually Do?
thoughts on capability, private markets, deployment, and firms
Model capability and economic performance in a business are not the same thing. They should be, but they’re not. Why?
As we progress from models that can chat to models that can literally solve novel math, it’s stupefying to realize we still know very little about model capability in end businesses across the right work, conditions, cost, and performance curves. We therefore also have almost zero idea around the economics of individual companies, much less industries as capability continues to increase.
Mind you, the current eval methods and datasets are terrible for grokking these issues.
As Sean Cai and others continually catalogue, public benchmarks are often too synthetic, too benchmaxxed, and detached from the actual conditions of work. Data providers mostly sell approximations of that work while the useful information, economic performance on real workflows, is held tightly by labs, deploycos, applied AI companies, and large enterprises that can justify measuring this themselves (who often still aren’t).
Everyone else is working from what lossy data, public evals, and esoteric prompts.
This state cannot last. Already others are beginning to ask similar questions as the sheer amount of deployment, buyout strategies, and more that could be run with this sort of data become evident:
Evals are underwriting instruments
Imagine AI has essentially solved 80% of the work in a particular business service. Let’s say accounting. How would I know that as an investor today?
Currently? Not much! There’s no good dataset you can access that shows performance over time, edge cases where models struggle, or even more minute areas of accounting that are vertically dependent. In fact, given the data state, there is probably more value in knowing what the best accountants already use AI than in even trying to do a formal evaluation.
If I had that information, what would I do with it? Do I target firms specializing in the remaining edge cases? Do those edge cases persist as the models improve, or are they simply next in line? Do I buy companies operating inside the 80%, knowing I can consolidate them and build a national brand on differentiated economics? Do I just buy Intuit stock?
These are the kinds of strategies you would expect investors to start running as models continue to improve along different capability axes.
But again, the underlying data you would want to start to form investment strategies against doesn’t exist. It’s only right now inferred.
When AI was a side event in the economy, this was fine. Now that we are talking about hundreds of billions in compute spend, financial instruments for buying compute, and more, this information asymmetry cannot last.
In fact, it’s unclear to me if the information asymmetry is even beneficial today.. My suspicion is that it benefits the labs and data markets first, the current sophisticated holding companies second1, followed by a few narrow market participants with access to real workflows.
Over time, financial actors are going to need far more detailed information on expected AI spend by function, cost and performance, capability trend assessments, company-specific evaluations, and more.
At the macro level, this matters to any serious compute-market trade. Downmarket, it becomes fundamental to acquisition diligence.
Soon the question will not be whether a model is generally capable, or even whether it is capable inside an industry. It will be whether a particular model, using a particular harness, can perform the exact work of an enterprise at the requisite quality and cost.
That information mostly does not exist in a form investors can use or even make a reasonable bet on.
The AI trade has a data problem
Why hasn’t this been solved? I’ve got some hunches. One is that this is only a very recent need. The value of evals grows as model capability grows. It isn’t worth precisely measuring something with zero aptitude for a task. Model capability was thin almost everywhere outside chat reponses and code, but as the models become plausibly capable of doing far more work, measurement starts to matter a hell of a lot more.
The existing data market has distorted this through its pursuit of data sales to labs. Synthetic evals that do not reflect the conditions of a real business make capability harder to assess. Most of the tasks you find in evals have near zero real world fidelity with tasks in a business.
The AI trade in private markets has also only recently become concrete enough to work. Current models can increasingly perform real work. Open models reduce dependence on a single lab. Post-training and better harnesses can create meaningful differences in performance.
Can you tolerate another two or three major model progressions without knowing what they will do to your portfolio?
Very quickly, I think financial markets are going to decide they can’t and asset strategies will move into one of two categories:
A firm can decide to avoid the AI trade and aggregate assets believed to be durable regardless of capability trends.
Or, AI capability trends can become a defining part of the firm’s underwriting.
This is admittedly very hard to do. There are few rigorous evals outside of a handful of industries. Even where they exist, a model completing a task in a controlled environment says very little on its own about whether a company can deploy it, whether it materially improves the firm’s economics, and whether it allows them to disproportionately win.
Those distinctions are important for deployment and arguably even more important when pricing an asset in the first place.
Alpha, beta, and the edge
AI wildly affects what constitutes alpha versus beta in an end market.
There’s tons of questions we would expect firms to take differentiated stances on with capability data more available. Does a capability advancement enhance the value of an asset or devalue it? Does it commodify an existing moat or does it enhance durability by allowing a company to win more?
Does the company possess out-of-distribution data that might be valuable for post-training?
Private markets will increasingly trade on proprietary answers to these questions.
If better data existed, I would expect public and private markets to quickly fill with strategies that sound exotic today. Compute-market intelligence is about as far as this has progressed, with hedge funds already spending heavily to understand supply, demand, and usage.2
But as more compute shifts toward inference, knowing where inference can be deployed successfully becomes at least as important as knowing total compute demand.
If a pharmaceutical company can run trillions of tokens through the right models, talent, and platform and materially increase the number of clinical trials it performs every year, wouldn’t that be important to know before buying the company?
Wouldn’t you expect pharmaceutical companies to heavily acquire compute in anticipation of that outcome?
These seem like fairly basic investment questions to me.
An ode to the deal guys
First things first: AI isn’t the whole game in firms. It never will be.
This is my felt frustration with collapsing every question about the future of business into a question about AI.
Firms exist to provide valuable goods or services. On almost any useful horizon, that is not reducible to AI.
AI isn’t everything. Data is jagged. Capability is jagged. Companies are not simply applied AI, nor will they ever be.
I do not think you can price an asset solely off its theoretical AI transformation.
You simply cannot tokenmaxx your way out of an inferior company.
Rollups remain mostly a function of market structure, multiple arbitrage, and operating dynamics. AI can enhance the operating dynamics. It has little to say about whether you should perform a rollup in the end market in the first place.
Deal structure matters. Risk matters. Management matters. The quality of the underlying company matters a hell of a lot.
We should expect the number of AI buyouts to increase rapidly, partly fueled by massive creativity in the deal guy segment.
But AI still needs to fit inside an actual investment thesis.
The more interesting question is how the best dealmakers will express theses that were not previously possible. If you possess a proprietary view of capability inside a function, you can target different assets, price them differently, and perhaps create operating models that do not make sense to anyone underwriting from historical performance alone.
A lot of supposedly AI-native private-market activity today is still a normal rollup with an AI transformation case attached to it.
This will get more interesting as the underlying capability data improves.
N of 1 companies
I firmly believe N of 1 companies matter more than ever.
They will never be N of 1 because of AI.
PSA remains one of my favorite examples. The company has something close to a monopoly on sports-card grading.
Can I imagine an ML model that grades cards more accurately than a human?
Easily.
Does that lessen the value of PSA’s brand in issuing the grade?
Not at all.
The grade matters because PSA issued it. Better models may improve the economics of grading without changing the source of PSA’s value.
If LLMs become a commodity business over a long enough time horizon, rarified assets and services should matter more. AI may let these companies produce more without commoditizing the reason customers care about them.
What would you give to own Rockstar in ten years if AI lets its talent make games faster and at 20% or 30% less cost?
AI is not the reason anyone cares about Rockstar. It may still have a massive effect on Rockstar’s output and economics.
N of 1 companies exist in nearly every vertical, or will be created.
They are in the brand, status, authority, access, or hits business. Their internal costs may fall. They may move faster. They may be able to win more territory without adding the same organizational weight.
AI becomes part of the internal logic of the company without becoming the reason the company matters.
This is also why I don’t think general statements about AI exposure are particularly useful.
Two companies performing similar work may respond completely differently to the same model progression. One may lose its reason to exist. The other may have the brand, distribution, data, or operating control to take advantage of the capability and expand.
You need to understand the company.
On deployment
Part and parcel of this is my frustration with the lack of creativity in deployment markets.
I’m sorry. Why the hell are you building another ERP or CRM?
If you have an asset worth a damn, the goal is to grow it. Your reason to exist as a deployco is that you can disproportionately grow assets.
You’re really telling me the most urgent thing you can do is rebuild a CRM?
The same goes for the endless conversation about token-cost optimization.
I understand it if you are automating a call center. I understand it in enterprise engineering. I understand that costs matter at scale.I get it.
I don’t understand cost reduction as the most compelling reason to post-train models right now.
Post-training should first be an exercise in leverage and alpha. If the answer is only that the company saves a little money per million tokens, I don’t find that particularly interesting. And my hunch is that by focusing on areas where post-training can move the needle, deployment companies are not thinking grand enough about the pace of frontier models and the number of things they are on track to materially aid. But again, they don’t have the capability data either to make this determination.
Operating control is still underweighted.
Even with the current surge of interest in AI buyouts, the dynamics around operating control and AI transformation cases is still underrated.
AI today often benefits the individual more readily than it benefits the firm.
An employee uses a model and finishes the work faster. Maybe the work is better. Maybe the employee can take on more of it but often the company does not capture the gain.
If the company wants the increased throughput, someone has to change how the work is organized.
This is extremely hard to accomplish without total buy-in from management.
Better models have not made this easier. They have mostly made it more obvious when the technical capability is ahead of the organization.
This is another reason assets will trade hands based on AI cases. Ownership creates the authority to change incentives, workflows, systems, and organizational structure at the same time.
AI is ultimately a sociotechnical revolution and will require historic patterns around deploycos, vendors, and often new management.
My gut sense is that it will become obvious that the labs themselves cannot solve this through their deployment companies. There are adversarial dynamics between labs and end businesses. Model routing, cost and performance tradeoffs, and the possibility of switching providers push companies toward mercenarial relationships with labs.
Sovereign talent
As knowledge work becomes more dependent on data, compute, and algorithms, it’s ironically become clearer that the best talent matters more than ever.
The best AI talent is increasingly sovereign.
By sovereign, I mean that these people possess a degree of independent productive capacity that employees historically did not.
Many theories of work model it as a process-driven and highly specialized craft. This misses the sheer creativity that categorizes the best knowledge workers.
Knowledge work is ultimately a creative endeavor with a useful model in the new world being the ability of naturally creative workers being able to manage increasingly more compute to surface real economic opportunity for their company. As model capability grows, the need for these people grows as well.
Already, the best AI-native firms see this incredibly clearly. Talent recruitment and retention are core to their businesses. Soon the wider market will operate under this reality as well.
How do you obtain the correct talent that is going to run the business over a 7 year time horizon? If the difference between mediocre and exceptional talent applying current models inside a middle-market company can easily be on the order of tens of millions of dollars in equity value, how do you identify the right people?
If the investment case increasingly depends on what a company can do with models, the people capable of answering that question and subsequently operating it need to become part of the underwriting case itself.
I wonder if capability data and talent questions are one and the same and I’m currently exploring with a wide range of partners on how to address this. If you’re a private markets investor asking these questions, feel free to reach out. I think it’s perhaps the most fundamental set of questions for the art of investing going forward.
But importantly only in their existing acquisitions. Nobody to my knowledge has built the underwriting and search motion around the dataset I’m describing.
On the order of tens of millions.



