Skip to content
⌘ K

By z· 2,698 words

ResearchAnalysisQuestion

Why compute might get 10x+ more expensive in coming years?

Working answer

Compute bills for ambitious AI projects could rise 10× or more by 2030, chiefly because the projects may use far more computation—not because an unchanged unit of compute costs 10× more. Reasoning, repeated attempts, verification and multi-step agents can improve valuable outcomes: OpenAI’s o1 evaluation solved 74% of competition-math problems with one sample versus 93% using a more extensive reranking procedure. If a project uses 20× the compute while its effective unit price halves, its bill rises 10×. That is a plausible scenario, not a forecast for a fixed task; scarce powered capacity could add local premiums.

Counter view

Efficiency may outrun added work. Stanford found a roughly 286-fold drop in the minimum price at a specified capability benchmark from 2022 to 2024, and AWS cut an existing H100 instance’s price 44% in 2025. Better models could also need fewer steps or retries. Evidence that harder projects consistently consume 20× more billable work despite halved effective prices would strengthen the 10× case; stable bills per successful project would weaken it.

Research note · September 25, 2026 · Outlook through 2030

How much computation would you buy for a better answer? In OpenAI’s evaluation of o1 on competition mathematics, a single sample solved 74% of problems. Consensus among 64 samples lifted that to 83%. Reranking 1,000 samples reached 93%. The improvement came from changing how much work the system did before delivering an answer. [1]

That experiment contains the strongest explanation for why AI compute bills might rise tenfold. As AI takes on more valuable work, developers can spend more computation on deliberation, alternatives and verification. The economic ceiling becomes the value of getting the job done.

Our view is that tenfold larger bills are plausible for increasingly ambitious AI projects by 2030. The arithmetic that matters is simple: 20 times as much computation at half the unit price produces a tenfold bill. Across scenarios with twofold to tenfold unit-price improvements, the required increase in work is 20–100 times, holding reliability constant. Those are the thresholds the workload must cross.

We give the greatest weight to growing workload intensity, followed by scarce ready-to-use capacity and the cost of recovering the capital behind it. We’re skeptical of a broad tenfold increase in the price of equivalent computation: matched hardware prices and fixed-benchmark inference costs have moved sharply downward. The investment question is whether customers’ appetite for more capable systems outruns those savings—and which suppliers get paid before the answer becomes clear. [2][3]

Better answers create an appetite for more computation

Reasoning changes the economics of inference. A model can generate alternatives, examine intermediate results and select among competing answers. OpenAI described the mechanism directly: performance improved with “more time spent thinking.” Its mathematics results show why a customer might accept a large increase in work for a smaller increase in success. Sample counts describe the evaluation procedures; translating them into dollars also requires the length and cost of each attempt. [1]

Exhibit 1. More inference work improved o1’s accuracy

OpenAI o1 accuracy on 2024 AIME problems, percent solved. Sampling and selection procedures differ; these observations do not specify billed-cost multipliers.

[1]
View data — original input
Original input data for Exhibit 1. More inference work improved o1’s accuracy; chart filters and transformations do not change this table.
procedureaccuracylabel
Single sample7474%
64: consensus8383%
1,000: rerank9393%

More extensive inference procedures improved accuracy, giving developers a reason to spend more computation on difficult answers.

Source: OpenAI. Procedures differ in sampling and selection; sample counts cannot be read directly as billed-cost multiples. [1]

Agents extend the same mechanism across a workflow. A coding system can inspect a repository, propose a change, run a test and revise its answer. Each step can require another model call and another reading of the context. More demanding work can also move toward more capable models: Anthropic reported that Opus served 54% of Claude Code sessions, against 10% of chat and Cowork conversations. That mix shift adds another route to a larger bill. [4]

Our inference is that high-value jobs can support much larger compute budgets when additional reasoning materially improves the outcome. The constraint is whether the extra work earns its cost. Longer context, parallel agents, repeated attempts and richer inputs all belong in the same accounting ledger: multiplying separate averages for each would count some of the work twice.

The scale of aggregate demand is already visible. At Google I/O 2026, Sundar Pichai reported more than 3.2 quadrillion monthly tokens across Google’s surfaces, against roughly 480 trillion the preceding year. Dividing those rounded observations gives approximately 6.7 times the activity. That expansion combines users, products and workloads; it shows how quickly total consumption can grow even while the economics of an individual request improve. [5]

We think cheaper computation can enlarge the set of applications worth attempting. Google’s growth is consistent with that mechanism. Whether it translates into a tenfold bill for a particular project depends on how much more work that project consumes after efficiency and reliability improvements.

The workload must outrun falling unit costs

The counterweight to rising ambition is powerful. AWS cut the on-demand price of its existing P5 instance, using H100 GPUs, by 44% in June 2025. Stanford’s AI Index found that the minimum query price at a specified GPT-3.5-equivalent MMLU score fell from $20 per million tokens in November 2022 to $0.07 in October 2024: about 286 times cheaper, calculated by dividing the two prices. That benchmark captures a particular capability threshold; complete workflows also depend on reliability and latency. [2][3]

A project’s bill therefore needs three terms: work per attempt, effective price per unit of work and attempts per successful completion. If the first rises twentyfold and the second halves, the bill rises tenfold at unchanged success. If success per attempt also falls from an assumed 80% to 64%, expected attempts per completion rise by 25%, taking the bill to 12.5 times its starting level.

This is an aggressive scenario for a more ambitious project. It requires much more work, and its retry assumption requires worse reliability. Better models can push the other way by completing a job in fewer steps. Five times the work at one-fifth the unit price leaves the bill unchanged. Our tenfold view applies to rising ambition; we would not apply it to a fixed basket of tasks.

Training budgets have a separate route upward. Epoch estimates historical frontier training-compute growth of fivefold annually and training-run cost growth of 3.5 times annually. A tenfold increase over four years requires 78% annual budget growth, calculated as the fourth root of ten minus one. We regard tenfold frontier budgets as plausible, although funding, power and research returns can interrupt the trend. [6]

The distinction matters to investors. More computation per model, more inference per project and more projects per customer can all increase spending. None requires the price of an unchanged computation to rise. The next question is whether the physical system can deliver the larger workload.

A contracted megawatt still has to become a working cluster

A useful AI cluster needs compatible accelerators, memory, networking, cooling and firm electricity at the same place. Relief at one stage can expose the next bottleneck. Our view is that permitted, powered and commissioned sites are the clearest downstream constraint in tight North American markets, while qualified memory and packaging can bind earlier for particular accelerator generations. [7][8][9][10]

CoreWeave gives the problem a concrete scale. At the end of the second quarter of 2026, it reported 1.5 GW of active power against approximately 3.7 GW contracted. The 2.2 GW difference, calculated by subtraction, is a development pipeline. CEO Michael Intrator described securing “power, sites, cooling, hardware, storage, networking and supply chain inputs well ahead of need.” The commercial commitment arrives before all the pieces are ready to earn revenue. [7][11]

Exhibit 2. CoreWeave’s power commitments run ahead of active capacity

CoreWeave power at June 30, 2026, in gigawatts. Approximately 3.7 GW contracted includes 1.5 GW active; the 2.2 GW difference is a development pipeline.

[11]
View data — original input
Original input data for Exhibit 2. CoreWeave’s power commitments run ahead of active capacity; chart filters and transformations do not change this table.
statusgwlabel
Active1.51.5 GW
Contracted3.7≈3.7 GW

CoreWeave’s contracted power exceeded active power by approximately 2.2 GW at June 2026.

Source: CoreWeave. Contracted capacity includes the active portion; the gap represents development commitments. [11]

Utility queues require the same discipline. Dominion’s Virginia territory recorded an approximately 4 GW coincident data-center peak in 2025. Its agreements totaled 47 GW across existing service, construction and engineering-study categories. Dominion’s demand forecast was 16.6 GW in 2046. The enormous queue includes projects at very different stages and probabilities of completion. [12]

Electricity infrastructure also moves slowly. The IEA puts new transmission lead times in advanced economies at four to eight years and estimates that around 20% of planned data-center projects could face delays without action. Its global all-data-center electricity scenario rises from 415 TWh in 2024 to 945 TWh in 2030, roughly 2.3 times by division. That is substantial growth to accommodate in a system whose hardest constraints are often local. [8][13]

There is relief upstream. SK hynix began HBM4 mass shipments in the second quarter of 2026 and planned a second-half ramp. Micron expected its Singapore advanced-packaging expansion to contribute meaningfully from the first half of 2027. More qualified memory can unblock systems, but a completed system still needs a site that can run it. [9][10]

We think this creates meaningful pricing power for immediate delivery in selected places. It can also force expensive site changes or leave capital waiting for commissioning. The size and durability of those premiums remain an open question. A queue by itself cannot tell investors what customers will ultimately pay.

Power access matters more than the electricity tariff

The obvious inflation story is a higher electricity bill. The cost stack points elsewhere. Using a published 72-GPU rack specification and a transparent set of project assumptions, we calculate electricity at about 5% of total cost per productive GPU-hour. A tenfold electricity tariff then raises the total by approximately 44%. The small share of electricity in this configuration makes broad tenfold cost inflation through tariffs alone difficult to sustain. [14][15][16][17]

Exhibit 3. Capital recovery dominates the modeled rack cost

InputValueBasisSources
Rack72 GPUs; 130 kWPublished 125–135 kW operating range; midpoint used[16]
System capital$5.0mNVIDIA illustrative investment[15]
External network/storage$1.0mAnalyst assumption
Facility capital$1.47m0.130 MW × JLL’s $11.3m/MW forecast; IT-MW basis assumed[14]
Recovery period / rate4 years / 10%Analyst project-recovery assumptions
Productive utilization / PUE70% / 1.30Analyst operating assumptions
Electricity tariff9.17¢/kWhJune 2026 US industrial average[17]

Capital recovery dominates this rack scenario, limiting the direct effect of higher electricity prices.

Sources: NVIDIA investment example, Supermicro specification, JLL construction outlook and EIA electricity data; our calculations. [14][15][16][17]

Calculation note: $7.469 million total capital is recovered over four years at 10%, across 441,504 productive GPU-hours annually. Maintenance is 5% of equipment and external network/storage capital. Power draw is held constant throughout the year. The four-year recovery period applies conservatively to the entire project, including the longer-lived facility. Software, tax, profit and exceptional demand charges are excluded.

The result is about $5.34 per productive GPU-hour for capital recovery, $0.68 for maintenance and $0.31 for electricity: $6.32 in total. Replacing the electricity component with ten times its value gives about $9.09. Across a wider assumed electricity share of 3–15%, the same tariff shock raises total cost by 27–135%, using one plus nine times the starting electricity share. The conclusion survives a considerably larger power share. [14][15][16][17]

Utilization has much greater force. In the same model, reducing productive utilization from 70% to 7% spreads unchanged annual costs over one-tenth as much output and raises allocated cost tenfold. That is a severe stranded-asset scenario. Owners may absorb losses, sell equipment or close capacity rather than pass the increase through to customers.

This is why a missing grid connection can matter more than an expensive kilowatt-hour. It prevents the capital from working at all. We think the more consequential risks are commissioning delays, weak utilization and replacement arriving before the original investment has paid back.

Suppliers collect before fleet owners recover their capital

The spending travels through several balance sheets. Customers buy applications or model access; model providers buy cloud capacity; operators buy accelerators, networks and facilities. Foundry and memory costs sit inside those systems. Adding revenue at every stage would count portions of the same customer dollar repeatedly, while current equipment purchases may be financing services delivered years later.

The clearest current earnings exposure sits with scarce inputs. NVIDIA reported $89 billion of Data Center revenue and a 75% companywide gross margin in fiscal Q2 2027. Those economics reflect its position in the buildout. Our view is that scarce chip and infrastructure suppliers have clearer current earnings exposure than an indiscriminate basket of compute owners, whose returns also depend on delivery, utilization and replacement costs. [18]

The opportunity extends beyond chips. Vertiv reported $3.27 billion of second-quarter 2026 sales and a 22.6% adjusted operating margin. Eaton’s Electrical Americas business reported a 27.5% operating margin. GE Vernova booked $2.7 billion of data-center Electrification orders and said first-half equipment orders were priced more than 20% above fourth-quarter 2025 orders, with configuration mix contributing to higher dollars per kilowatt. Scarcity is reaching the equipment needed to turn a site into usable capacity. [19][20][21]

The owner’s accounts tell a different part of the story. CoreWeave’s quarterly revenue more than doubled to $2.58 billion, but its $1.51 billion adjusted EBITDA coexisted with a $49 million GAAP operating loss and $640 million of net interest expense. Hardware consumption and financing separate a strong operating cash proxy from the return available to shareholders. [11]

Exhibit 4. CoreWeave’s EBITDA leaves a large capital burden

Q2 2026 measureAmountShare of revenueSources
Adjusted EBITDA$1.51bn58.6%[11]
GAAP operating result−$49m−1.9%[11]
Net interest expense$640m24.9%[11]

CoreWeave’s rapid revenue growth produced a large EBITDA figure alongside an operating loss and substantial financing costs.

Source: CoreWeave Q2 2026 results; margins are calculated by dividing each measure by $2.575 billion revenue. Interest is shown as an expense. [11]

This is not a judgment that every host loses money. AWS reported $42.2 billion of quarterly revenue and $16.6 billion of operating income; Google Cloud reported $24.8 billion and $8.81 billion, respectively. Their whole-cloud operating margins were approximately 39% and 36%. Established workloads and businesses help finance expansion, while the returns on incremental AI capacity remain embedded in broader accounts. [22][23]

Nevertheless, buyers are committing capital on a formidable scale. Microsoft said quarterly capital expenditure was $41 billion, “including the impact from higher component pricing,” with roughly two-thirds going to short-lived assets, primarily CPUs and GPUs. Alphabet guided to $195–205 billion of 2026 capex; Meta guided to $130–145 billion, including finance-lease principal. These companywide commitments put the recovery period at the center of the investment question. [24][25][26]

Price competition can shorten that recovery period. CoreWeave’s advertised eight-H100 node rates were $49.24 an hour on demand and $19.71 spot, reflecting different service terms. At an assumed 70% billable utilization, the on-demand rate generates about $302,000 annual gross receipts per node. A hypothetical 50% fall in the realized rate removes approximately $151,000 before any expense adjustment. The calculation is hourly price multiplied by annual billable hours. [27]

One contract reveals how the risk can travel back upstream. CoreWeave disclosed an initial $6.3 billion arrangement under which NVIDIA would purchase residual unsold capacity, subject to delivery and availability, through April 2032. The chip supplier can therefore become a buyer of its customer’s output. That contingent support is useful to the operator, but the ultimate economics still require valuable work for the capacity to perform. [28]

Scarcity could still overpower the cost curve

The strongest case against our skepticism is that market prices can detach from production costs. Customers competing for the last available energized cluster may pay a premium far larger than its electricity bill. Microsoft’s higher component costs, strong supplier margins and long infrastructure lead times give this argument substance. [24][18][8]

We give that risk substantial weight for selected locations, immediate-delivery contracts and new accelerator generations. It is why we would not promise every buyer lower prices. To support a marketwide tenfold increase, however, the evidence would need to move from queues and order books to sustained increases in realized prices for matched useful output across suppliers.

The opposing force is improving efficiency and reliability. Better systems can use fewer steps, custom accelerators can offer alternatives, and new memory capacity can relieve product shortages. Microsoft has claimed more than 30% better total cost of ownership for Maia 200 against its stated fleet comparator. Such vendor comparisons are workload-specific; their relevance is that buyers have several ways to contest supplier pricing power. [29][9][10]

Our view would strengthen toward tenfold project bills if billing and outcome logs showed twentyfold more work despite halved effective prices, or fiftyfold more work despite fivefold savings. It would weaken if models finished those harder jobs with fewer steps and retries. For physical scarcity, the decisive series is contracted capacity becoming active, firm, cooled and productively loaded capacity.

For investors, the final test is cash recovery: realized prices, utilization, collections, replacement spending and interest for a consistent fleet cohort. If those improve, compute owners can share more of the value now visible upstream. If customer receipts fall short, shareholders, lenders or suppliers must finance the difference. The asset that matters is a paid, useful computation.

Sources

  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.
  9. 9.
  10. 10.
  11. 11.
  12. 12.
  13. 13.
  14. 14.
  15. 15.
  16. 16.
  17. 17.
  18. 18.
  19. 19.
  20. 20.
  21. 21.
  22. 22.
  23. 23.
  24. 24.
  25. 25.
  26. 26.
  27. 27.
  28. 28.
  29. 29.

Comments

LatestPopular
Write a comment
Loading comments…

Request a Thesis

Tell us what you’d like Roadstar to investigate.

New question

What would you like to know?

Context guides the research and is not shown on the finished page.