OpenAI reported that o1 solved 74% of 2024 AIME problems with one sample, 83% using consensus among 64 samples, and 93% with reranking of 1,000 samples. Source
1
The results establish a reason to spend more computation on difficult answers. Sample counts do not establish corresponding dollar-cost multiples.
A fixed-benchmark inference price reaches $0.07 per million tokens
Stanford’s comparison records a minimum query price of $0.07 per million tokens at its specified GPT-3.5-equivalent MMLU threshold by October 2024, down from $20 in November 2022. Source
AWS cuts the on-demand price of an existing H100 instance
A September 2025 CoreWeave disclosure describes an arrangement obligating NVIDIA to purchase residual unsold capacity, subject to delivery and availability, through April 2032. Source
Google reports a surge in aggregate token activity
At Google I/O 2026, Sundar Pichai reported more than 3.2 quadrillion monthly tokens across Google’s surfaces, versus roughly 480 trillion at the preceding year’s I/O. Source
CoreWeave reports the financial burden of expansion
CoreWeave reported second-quarter revenue of $2.575 billion and adjusted EBITDA of $1.510 billion, alongside a $49 million GAAP operating loss and $640 million of net interest expense. Source
By z· 2,698 words
Why compute might get 10x+ more expensive in coming years?
OpenAI shows how additional inference improves answers
[1]OpenAI reported that o1 solved 74% of 2024 AIME problems with one sample, 83% using consensus among 64 samples, and 93% with reranking of 1,000 samples. Source
The results establish a reason to spend more computation on difficult answers. Sample counts do not establish corresponding dollar-cost multiples.
A fixed-benchmark inference price reaches $0.07 per million tokens
[3]Stanford’s comparison records a minimum query price of $0.07 per million tokens at its specified GPT-3.5-equivalent MMLU threshold by October 2024, down from $20 in November 2022. Source
AWS cuts the on-demand price of an existing H100 instance
[2]AWS announced a 44% on-demand price reduction for its P5 instance using H100 GPUs. Source
NVIDIA agrees to backstop unsold CoreWeave capacity
[28]A September 2025 CoreWeave disclosure describes an arrangement obligating NVIDIA to purchase residual unsold capacity, subject to delivery and availability, through April 2032. Source
Google reports a surge in aggregate token activity
[5]At Google I/O 2026, Sundar Pichai reported more than 3.2 quadrillion monthly tokens across Google’s surfaces, versus roughly 480 trillion at the preceding year’s I/O. Source
CoreWeave reports the financial burden of expansion
[11]CoreWeave reported second-quarter revenue of $2.575 billion and adjusted EBITDA of $1.510 billion, alongside a $49 million GAAP operating loss and $640 million of net interest expense. Source
Sources
Comments