As inference costs fall 10×, will demand for useful AI work rise by more or less than 10×?
1. The key distinction: tokens are not useful work
The question has three different quantities:
Inference consumption — tokens, calls, agent turns, or compute.
AI-assisted activity — the number of tasks, users, workflows, and business functions using AI.
Useful AI work — successfully completed, economically valuable tasks that produce an accepted outcome with tolerable human-review cost.
A 10× price decline can increase all three by different amounts. If demand has price elasticity greater than one, inference consumption can rise by more than 10×. But useful work is approximately:
useful work = attempted work × success rate × acceptance rate × workflow capacity
The last three factors need not rise when price falls. In fact, cheaper inference may initially increase experimentation and failed attempts, making token demand grow faster than successfully delivered work.
2. Why demand is likely to be elastic
A lower price changes the economics in several ways. It makes existing workloads cheaper, permits more attempts and verification, enables longer agentic runs, and makes entirely new applications viable. Andreessen Horowitz argues that each order-of-magnitude cost reduction opens use cases that were previously not commercially viable. Its analysis estimated roughly a 10× annual decline in the cost of an LLM at equivalent performance, while also identifying hardware, quantization, software, smaller models, instruction tuning, and open-source competition as cost drivers. [1]
Epoch AI finds that the decline is not uniform: the price of reaching specified capability thresholds fell between 9× and 900× per year across its benchmark and performance-threshold sample, with a 40× annual decline cited for reaching GPT-4 performance on a PhD-level science benchmark. Epoch also cautions that the fastest declines may not persist and that its analysis does not explicitly model the causes of price reductions. [2]
The most direct recent demand signal is suggestive of elasticity. Citadel Securities reports that usage-weighted token prices fell 40% from the end of June, while July AI spend per employee rose 49% month over month among the top 1% of firms, 25% among the top 10%, and 9% at the median. It interprets the combination as evidence that falling prices were expanding the quantity of intelligence consumed rather than reducing aggregate ecosystem value. This is useful evidence, but it is not a clean causal estimate of the response to exactly a 10× price shock: the observation is short-term, spending is concentrated among firms at different adoption stages, and other changes may have occurred simultaneously. [3]
OpenAI’s enterprise data points in the same direction. By June, firms in the top 10% of monthly AI usage generated 8.3× as many output tokens per active user as typical firms, up from a 2.6× gap in January. The measure is a proxy for depth of use, not a direct measure of completed economic output. OpenAI also reports that Codex represented 64% of combined Codex and ChatGPT enterprise output tokens, while weekly active enterprise Codex users grew 108× in legal, 41× in sales, 41× in recruiting, and 26× in marketing from February, compared with 5× in engineering. [4]
These observations support the following mechanism: once inference is cheap enough, the constraint shifts from “can we afford another call?” to “which additional tasks are worth attempting?” The answer can produce a very large increase in consumption, especially among organizations that already have data, tools, and repeatable workflows.
What a 10× cost decline can and cannot imply
| Transmission channel | Why demand can increase | Why the increase may be less than 10× | Sources |
|---|
| Lower unit cost | New use cases become commercially viable; existing users may run more inference, use longer contexts, or delegate more steps. | Potentially greater than 10× demand if the demand response is highly elastic. | [1][3] |
| Expansion of users and workflows | Demand can grow through new adopters, new departments, and a shift from assistance to delegated execution. | Raises useful-work demand, but adoption and workflow integration take time. | [3][4] |
| Conversion into useful work | Useful output depends on task success, reliability, autonomy, and whether the task is economically valuable. | This conversion layer makes useful work grow more slowly than raw tokens when failures and oversight remain material. | [5][7] |
| Supply-side bottlenecks | Human review, organizational processes, data quality, task specification, and domain-specific capability constrain deployment. | These constraints can make aggregate useful-work demand rise less than 10× even when inference capacity is abundant. | [5][7] |
3. Why useful work probably grows less than 10× initially
Reliability is a multiplicative bottleneck
Anthropic’s Economic Index shows that Claude’s success rate declines as tasks become longer and more complex. In its API data, success falls from around 60% for tasks estimated to take less than an hour to roughly 45% for tasks estimated to take humans five or more hours; the fitted line crosses 50% at approximately 3.5 hours. Claude.ai’s extrapolated 50% point is about 19 hours, but that estimate is based on a linear fit and should not be treated as a universal capability threshold. [5]
The same report estimates that speedups alone could imply a 1.8-percentage-point increase in annual US labor-productivity growth over the next decade. After adjusting for task reliability, the estimate falls to 1.2 percentage points for Claude.ai usage and 1.0 percentage point for API traffic. The adjustment is a concrete illustration of why more inference does not translate one-for-one into useful work: failed or correction-heavy attempts consume resources without delivering proportional value. [6]
Anthropic also reports that effective AI coverage is a weighted measure of the share of a worker’s day that can be performed successfully, not simply the fraction of task types that AI touches. Even 90% task coverage need not imply a large job impact if the system fails on key tasks or misses the most time-intensive work. [5]
Human and organizational capacity remains scarce
Useful AI work requires more than model calls. People must specify goals, provide context, connect tools and data, review outputs, resolve exceptions, and change business processes. OpenAI’s enterprise data explicitly says that access alone may not be enough to scale AI; the largest usage gaps occur alongside deeper adoption of capabilities that connect agents to company context, tools, and repeatable workflows. [4]
This creates a staged response to lower prices:
First, existing users run more calls, retries, context, and verification.
Second, organizations extend AI to additional teams and processes.
Third, new products and services are built around abundant intelligence.
Only then does the additional inference reliably become additional economic output—and the transition may be limited by review capacity, data readiness, regulation, customer demand, and the availability of well-specified tasks.
The task universe is not infinitely expandable
METR defines a task-completion time horizon as the duration of a task, measured by human expert completion time, at which an agent is predicted to succeed with a specified reliability. It emphasizes that this is a measure of task difficulty, not the time the AI takes to complete the task. Its benchmark covers more than one hundred software tasks, and METR warns that an eight-hour horizon does not mean an AI can perform eight hours of a high-context professional’s work. The benchmark is concentrated primarily in software engineering, machine learning, and cybersecurity, while many jobs involve people, ambiguous goals, and success criteria that are not algorithmically scored. [7]
Thus, lower cost can expand the set of attempted tasks without proportionally expanding the set of tasks that are reliably completed. It can also increase the value of using AI as a supervisor-assisted tool rather than as an autonomous substitute.
4. A simple elasticity framework
Let Q be demand for inference or attempted AI work and P be price. Define the absolute price elasticity of demand as:
ε = percentage change in Q ÷ percentage change in 1/P
For a 10× price decline:
If ε = 0.5, demand rises about 3.2×.
If ε = 1.0, demand rises 10×.
If ε = 1.5, demand rises about 31.6×.
Those are illustrative elasticities, not estimates established by the evidence above. The empirical record is sufficient to say that some segments look highly elastic, but not sufficient to assign one economy-wide elasticity to useful AI work.
The relevant elasticity for useful work is lower than the elasticity for tokens whenever additional calls are used for experimentation, retries, evaluation, or human-supervised iteration. Conversely, it can be higher in a new application whose economics cross a viability threshold after the price decline.
5. What the current evidence supports
Near-term aggregate answer: less than 10× useful work
The best base case is less than 10× growth in aggregate useful AI work after a 10× inference-cost reduction. The rationale is not that demand is inelastic. Rather, demand appears elastic but the conversion pipeline is constrained:
Existing users can consume much more intelligence.
New users and departments can enter.
Agentic workflows can turn short interactions into long, delegated tasks.
But reliability falls on longer tasks, human review remains necessary, and organizations must redesign workflows before cheap inference becomes completed work.
Anthropic’s reliability adjustment—from 1.8 to 1.2 percentage points of implied annual productivity growth for Claude.ai and to 1.0 percentage points for API traffic—is consistent with this “elastic consumption, slower useful-output conversion” view. [6]
Selected segments: potentially more than 10×
Some segments can exceed 10× because they start from low adoption or because the price decline unlocks a new use case. OpenAI reports especially rapid growth in non-developer and cross-functional agent use, including 108× growth in weekly active enterprise Codex users in legal and 41× in both sales and recruiting from February to June. Those figures are adoption measures, not controlled price elasticities, but they show how quickly usage can scale once capability and workflow fit improve. [4]
OpenAI also reports that by May 2026, 80.6% of sampled individual users had made at least one Codex request estimated to correspond to more than 30 minutes of human work, 70.2% had made one exceeding an hour, and 25.6% had made one exceeding eight hours. The estimates are directional: task horizons were inferred with an LLM judge, and the analysis used a random sample of 0.1% of users. [8]
Long-run possibility: more than 10× in economic value, but not guaranteed
Over a longer horizon, the answer could become more than 10× if cheaper inference enables products and services whose demand was previously suppressed by cost. The relevant effect is analogous to a Jevons-style response: efficiency lowers unit cost, which can stimulate enough new consumption to increase total demand. The conceptual mechanism is supported by the observed coexistence of falling token prices and resilient or rising compute demand, but the available evidence does not establish a stable, economy-wide causal coefficient. [3]
The long-run upside depends on whether complementary inputs scale: trustworthy data, integration with software systems, human operators, distribution, customer willingness to pay, and models that are reliable enough for the relevant domain. If those complements scale, useful work can compound faster than the original price decline. If they do not, inference becomes abundant but underutilized.
6. Scenario view
Illustrative useful-work response to 10× cheaper inference
Illustrative useful-work response to a 10× inference-cost reduction under three deployment scenarios.
0 — 20 · Scenario · Illustrative increase in useful work (×) · ×
View chart data
Illustrative scenarios, not forecasts or observed elasticity estimates. A 10× increase is the break-even threshold for unit elasticity after accounting for reliability and deployment frictions.
The scenario chart is not a forecast. It illustrates the threshold implied by the question: a 10× increase in useful work requires unit-price elasticity of at least one after accounting for reliability, acceptance, workflow capacity, and task supply. The evidence supports a wide range rather than a single point estimate.
| Scenario | Demand response to 10× lower inference cost | Interpretation |
|---|
| Constrained deployment | 2–5× | More attempts and existing-user usage, but reliability, review, and integration bottlenecks dominate. |
| Base case | 5–10× | Strong expansion in usage and workflows, with useful-output conversion lagging token growth. |
| Breakout applications | 10–30× | New products or processes cross a viability threshold; adoption and task supply expand rapidly. |
These ranges are analytical scenarios, not observed estimates. They should not be confused with the much larger historical declines in inference prices or the large usage multipliers reported for particular OpenAI populations.
Bottom line
A 10× fall in inference cost is likely to produce more than a proportional increase in some forms of AI consumption, and may produce more than 10× demand in newly viable applications. But for the economy-wide quantity that matters—useful, accepted, economically valuable AI work—the defensible central answer is less than 10× in the near term.
The decisive variable is not price alone. It is the product of price elasticity and conversion capacity: how many additional tasks people attempt, how reliably AI completes them, how much human review they require, and how quickly organizations can redesign workflows around abundant intelligence. Current evidence shows strong demand expansion but also substantial reliability and deployment frictions. It therefore supports a conditional, segmented answer, not a universal 10× rule.
Comments