On the unit economics of inference.
A category priced like software will, in time, earn the margins of compute. The investment question is not whether this is true, but where in the stack the exceptions live.
For most of the last decade, software was understood to be the highest-quality business model available to public and private capital. Gross margins above eighty percent, dollar-based net retention reliably above one hundred, marginal cost of distribution rounding to zero. These properties were so reliable that the SaaS multiple became a stand-in for quality itself.
The current generation of artificial intelligence businesses is being financed at multiples that assume these properties. We think a meaningful fraction of them will not have them.
The reason is mechanical. Software's gross margin is a function of distribution: once written, a line of code costs nothing to deliver to the marginal customer. Inference is not software. Every query consumes compute, every token costs energy, and the cost is paid in real time at the moment of use. The marginal cost of an AI product is not zero. It is the cost of GPU-hours billed against a depreciating asset, plus the energy to run it.
This makes the economics of an AI product look structurally closer to those of a hosting business than to those of a SaaS business. The gross margin ceiling is not eighty percent. It is whatever margin the underlying compute provider chooses to leave on the table — which, given that the compute layer is itself competitive and capital-intensive, is unlikely to be generous over a full cycle.
The two arguments against this view.
There are two thoughtful counterarguments, and both deserve serious treatment.
The first is that inference cost will fall faster than usage rises, restoring software-like margins through deflation. This is partly true and partly misleading. Inference costs per token are falling sharply. But the average query is becoming dramatically more expensive: longer context windows, more reasoning, more tool calls, more multimodal inputs. The net effect on cost per useful unit of work has been ambiguous. Companies that planned for monotonic deflation have been surprised.
The second is that model differentiation creates moats: a company with a better proprietary model can charge a price untethered from compute cost. This is true for a small number of frontier labs. It is not true for the long tail of application-layer companies, which depend on access to models they do not control, on terms they do not set, with switching costs that are technical rather than contractual.
Where the exceptions live.
This does not mean AI is uninvestable. It means the investable surface is narrower than the financing volume implies, and the durable businesses are concentrated in particular structural positions.
The first position is owning the infrastructure layer itself — compute, data, orchestration, the tools that make inference reliable at scale. These businesses look more like cloud than software, but cloud has been an extraordinary category for capital. The margin ceiling is lower, but the demand curve is structural and the switching costs are real.
The second position is at the application layer, where the company owns proprietary data, distribution, or workflow integration that the model cannot replicate. The model is a commodity input. The moat is everything around it. In these cases the margin profile can look like software because the customer is paying for the data and the workflow, not the inference.
The third — and the smallest — is at the model layer itself, for the very small number of labs at the frontier. These are not venture investments in the conventional sense. They are infrastructure bets sized accordingly.
What this means for venture returns.
The implication for fund construction is uncomfortable for a category that has absorbed an enormous amount of capital. If the median application-layer AI company has the gross margin profile of a hosting business, the multiple at exit will reflect that, and the math of the fund will not work at the prices being paid in 2024 and 2025.
The most useful question to ask an AI company in 2026 is not what its product does. It is what its gross margin will look like at scale, and which of the three positions above explains that margin.
We have not invested behind the consensus answer to this question. We have invested where we believed a company occupied one of the three positions defensibly, and we have passed on a number of well-regarded companies where we believed the gross margin story did not survive scrutiny. Time will tell whether we are right. The mechanics, we think, are clear enough.