AI’s likely profit stack is barbelled: scarce upstream infrastructure and downstream products that own workflows, data, distribution, trust, and transactions.
Generic model APIs and thin wrappers are vulnerable; most generic-model value should flow to customers through lower prices and productivity gains.
Value persists where workflow owners turn context, integration, reliable governance, and action into dependable outcomes; infrastructure rents remain large but cyclical.
This directional thesis could shift with persistent model differentiation, capacity oversupply, interoperability, or weak application retention.
The answer
As foundation models commoditize, durable AI profits should migrate to scarce complements: upstream compute bottlenecks and downstream products that own workflows, proprietary data, distribution, trust, and transactions. Generic model APIs and undifferentiated “wrappers” sit in the vulnerable middle. The largest durable software pool should be downstream, while the largest near-term cash pool remains upstream infrastructure.
A durable pool needs more than growth. It needs scarce supply, switching costs, recurring demand, and an ability to retain customer value rather than pass savings through.
1. Workflow owners should capture the broadest durable software pool
Applications can swap cheaper or better models while preserving the customer relationship. The strongest will control a complete workflow, connect proprietary data, execute actions, manage exceptions, and prove measurable outcomes. Examples include coding, clinical administration, legal work, customer support, and financial operations.
This market is already forming. Menlo estimates that applications received $19 billion of 2025 enterprise generative-AI spending, versus $18 billion for infrastructure. Departmental applications received $7.3 billion, including $4.0 billion for coding. Vertical applications received $3.5 billion, with healthcare near $1.5 billion. Moreover, 76% of surveyed use cases were purchased rather than built. Startups captured 63% of application revenue, showing that incumbents do not automatically own this layer (Menlo Ventures).
However, current revenue does not prove durability. A thin interface loses pricing power when rivals can access the same models. Defensibility comes from embedded workflows, accumulated customer context, regulated-domain expertise, integrations, distribution, and becoming a system of record or action. As model costs fall, these firms can retain some savings as margin. Firms without those assets must pass savings to customers.
Reliability and governance are part of the product, not overhead. In a randomized experiment with 758 consultants, GPT-4 users completed 12.2% more tasks and worked 25.1% faster on suitable tasks. Yet they were 19% less likely to solve a task outside the model’s capability frontier correctly (Organization Science). Therefore, valuable applications will identify task boundaries, evaluate outputs, preserve human escalation, and provide audit trails. Customers should pay for dependable outcomes, not tokens.
2. Scarce compute complements should retain large, but cyclical, rents
Model abundance does not eliminate demand for accelerators, high-bandwidth memory, networking, power, cooling, fabrication, and data-center capacity. Rents accrue to whichever component remains hardest to expand or substitute.
NVIDIA illustrates the present pool. Fiscal 2026 revenue reached $215.9 billion, up 65%, with a 71.1% gross margin. Data-center revenue rose 68%. Its advantage extends beyond chips into CUDA and hundreds of software libraries and frameworks (NVIDIA 10-K). The OECD reports that NVIDIA held 90% of the GPU market in 2023. It also identifies high fixed costs, scale economies, long development cycles, and switching costs as structural entry barriers in AI hardware (OECD).
This does not establish permanent NVIDIA rents. Its two largest direct customers represented 22% and 14% of fiscal 2026 revenue. Customer-designed chips, excess capacity, export controls, or a shift in the physical bottleneck could compress returns. The durable thesis concerns scarce infrastructure and its software ecosystem, not one supplier forever.
3. Clouds and distribution platforms can tax the ecosystem
Hyperscalers can capture value through enterprise distribution, security, integrated data, model marketplaces, and switching costs. Their advantage is stronger in the control plane than in commodity compute resale.
The FTC found that Microsoft–OpenAI, Amazon–Anthropic, and Google–Anthropic arrangements included equity, revenue sharing, cloud-spending commitments, and product integration. Developers committed billions to partner clouds. Migration can be time- and capital-intensive, while proprietary cloud chips can create both cloud and chip switching costs (FTC). These mechanisms let platforms earn from compute, model usage, and downstream distribution simultaneously.
Cloud rents face antitrust, interoperability, and customer concentration risks. Open standards and multicloud routing would shift value toward independent orchestration, security, evaluation, and data-governance vendors.
4. Frontier-model rents should be narrower and less stable
Model providers can still earn temporary premiums for frontier capability, reliability, brand, safety, or superior inference efficiency. Menlo estimates model APIs generated $12.5 billion in 2025. Yet share moved sharply: Anthropic reached 40%, while OpenAI fell from 50% in 2023 to 27% in 2025. That volatility weakens the case for durable provider lock-in.
Stanford reports that leading models are becoming nearly indistinguishable. The top closed model led the top open-weight model by only 3.3% in March 2026. Competition is therefore shifting toward cost, reliability, and domain performance (Stanford AI Index 2026). Generic intelligence should increasingly resemble a competitive input, although frontier advances may create episodic rents.
Bottom line and uncertainties
The likely stack is barbelled: scarce physical infrastructure upstream, and workflow, data, trust, distribution, and transaction ownership downstream. Most generic model value should flow to customers as lower prices and productivity gains.
The evidence is directional, not a profit census. Menlo’s figures are estimates from roughly 500 U.S. decision-makers and a bottom-up model. NVIDIA shows current economics, not long-run equilibrium. The FTC identifies risks, not guaranteed outcomes. Persistent model differentiation would preserve laboratory rents. Hardware diversification, capacity oversupply, interoperability mandates, or weak application retention would shift the conclusion.
Comments