There are liars, damned liars, and people who play silly buggers with scales.
This graph misrepresents the rate the price declines and the length of time it has been stable for, which throws off nearly all non-trivial conclusions.
Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.
What does "scarce" mean to you?
I'd wager that to most reasonable people it means "available in a supply that, measured against demand, is low". Market forces typically react to such a state by pushing prices up. Saying "they're not scarce, they're just expensive" is just silly.
The exception to this would be during e.g. supply chain hiccups (lots of compute sitting there, unable to get to its operational destination) or when other forces artificially push up prices. But that's not what we're seeing here. Compute is much scarcer than it's been for years.
I'm just grateful that I don't need to upgrade my computer for a while, and cross my fingers that I don't have a hardware failure in the next couple of years. I had an unpleasant laptop failure in 2021 that I don't wish to repeat.
I think it more resembles the "content" of our era.
https://refactoringenglish.com/blog/why-i-stopped-creating-c...
We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park)
Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam
It’s all choices
We just are making bad ones mostly
The choices now are would you like ads or more ads. Shall I waste compute seeing if the web page you clicked on can be summarised ?
The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out.
It might take the US stock market with it …
I don’t believe in 5-10yrs we’ll be in a shortage anymore
https://www.apollo.com/insights-news/pressreleases/2026/01/a...
https://www.apollo.com/insights-news/pressreleases/2025/11/a...
etc
Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range.
So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.
As long as they sell an iPhone 18 at a more accessible price point, I do not see the issue.
While it will be at a premium price point, people use their phones for hours every day, while the Vision Pro is a niche gadget.
With videos, photos, web browsing, and reading being much better on a foldable, I can see this have mass appeal.
1. There is a huge demand for compute, specifically GPU compute
2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills
3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying.
4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others).
All while advances in open weight models are making it appear that the major labs truly have no model moat.
Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.
If the massive demand is still present for compute at market rates (which I believe it is), then your second point is just investors spending cap ex to build out valuable and profitable assets, no problem there.
Time will tell.
Does it not make sense to rent out your compute if competitors have a better model and demand at higher prices?
About the moat, Mythos was first made available to customers at the beginning of April. Kimi K3 is still behind this.
Both OpenAI and Anthropic are expected to deploy significant upgrades in August.
No, it is not a moat.
The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over.
> At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027.
MS bought more GPU's than they had rack space for: https://www.datacenterdynamics.com/en/news/microsoft-has-ai-...
Open AI bought out the memory: https://x.com/kwharrison13/status/2029248559388746168 but they have no means to consume anywhere close to their order.
Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them.
Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows).
At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.
Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.
AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.
But.. "good design taste" has no compiler.
Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.
In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.
Before you even get the code for a compiler to do its thing, you need an AI to ingest and emit thousands of tokens autoregressively. That's what makes RLVR so expensive.
Likewise - I don't think your "design taste" argument holds? Areas where "taste" exists are far easier for AI to tackle than areas where no data exists. It's easier to teach an AI how to make an image that looks good than to teach an AI how to de-solder a BGA chip. "Tacit knowledge" yes, "taste" - not really?
I think the best way to fix the graph would be shading or line texture to indicate the scale. The biggest problem is that the inconsistent scale is a surprise you only see upon close reading, after you've visually digested the trend. So the scale needs to be apparent in this first visual digest. (Sort of like how many logarithmic charts include thin axis marker lines in the body of the graph itself, so as to immediately inform a quick glance.)
As such, saying “GPUs are scarce” is an empty statement from a strictly economic perspective - so are all other physical goods. The correct economic phrasing would be “there is a shortage of GPUs” or “GPUs are in high demand”.
Of course, in colloquial English we interpret “GPUs are scarce” as meaning the same thing.
Housing, collectors items, access to the time of a top of the line doctor... all relatively scarce. And You'd be laughed out pf most rooms if you claimed those things aren't scarce.
If it were true that they’re making money hand over fist they wouldn’t need to raise tens of billions of dollars every few months.
How do you know they are not?
It will be curious to see the cost of inference for these newly released open weight models and will help give an idea of the actual cost of inference. But for now, I think saying the $200 plans allows for "tens of thousands of dollars worth of inference" provides very little insight when you are measuring the inference cost in API pricing with an unknown margin.
The simple question to ask is, if it is so profitable, where is all the money going? If Anthropic have 90% margins on API usage and API usage is $50bn+ in revenue per year, where is the $45bn going? Why do they need to raise so much cash, constantly?
But I do wonder how a 60% margin would be realistic when Sonnet costs 3-6x more than GLM 5.2 hosted by third party providers.
Anthropic aren't building out the data centres themselves, they're renting/leasing/borrowing from companies that are doing the actual spend on building out infrastructure. And the data centre companies aren't spending their own money, they're borrowing too (hence Apollo investing in data centres). Anthropic are paying SpaceX ~$1.25bn/month right now for access to more compute, that's $15bn a year, more than what these supposed margins would require in total spend (based on current revenue estimates).
https://www.anthropic.com/news/higher-limits-spacex
The SpaceX deal is a great example of Anthropic creating demand, i.e:
> We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.
They committed to spending $15bn per year with SpaceX and then increased limits for customers on fixed cost plans, creating more demand without any increase in revenue.
So, sure, it's not necessarily that they are raising money because they are unprofitable, but no alternate explanation makes any sense. The argument that could maybe made in favor is based on announcements like this one:
https://www.anthropic.com/news/anthropic-invests-50-billion-...
> Today, we are announcing a $50 billion investment in American computing infrastructure, building data centers with Fluidstack in Texas and New York, with more sites to come. These facilities are custom built for Anthropic with a focus on maximizing efficiency for our workloads, enabling continued research and development at the frontier.
You might conclude from that, Anthropic are financing Fluidstack's build out, but they're not.
https://x.com/fluidstack/status/2079250004510728559
Just after that announcement, Fluidstack raised $830 million to build out data centres, none of the money coming from Anthropic. Fluidstack are currently rumored to be raising another $1bn. Anthropic's "$50 billion investment in American computing infrastructure" is just committed spend on renting compute from Fluidstack, a commitment that Fluidstack then use to raise money to actually deliver it. If Anthropic making money hand over fist, they wouldn't need to raise for committed spend.
And thus we return to the original question, how does future demand translate to spend? Actual handing over of dollars?
Meanwhile, OpenAI is at ~$25B ARR, but is likely not yet profitable.
https://newsletter.semianalysis.com/p/anthropic-growth-and-b...
Their link seems to claim semi analysis thinks it is 80%. It looks like it might be referencing this newer article from them, as the same picture is in both articles, but I didn't feel like paying to find out: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
The industry and users moves on from single chat-based to more and more "agentic" workflows that may generate longer workloads with multiple simultaneous agents (separate agents - separate contexts - separate KV caches).
My estimation is based on, say, running a Kimi3 on a 24 B200 GPUs - it is very easy to lose money when selling tokens at "market" prices.
Phones are a bad example in the US at least-not sure how it works in the EU.
I think most people in the US simply finance new phones through their provider who is already providing them the plan, so it just looks like an expensive phone plan to them for a few years. Apple could offer a deal where a $3000 phone is simply a $200/mo. add-on to your bill over 3 or 4 years.
I bet this causes more things to be tied to a cellular provider and offered the same way, like game consoles and maybe even TVs. Imagine signing on to your cellular provider for a 10 year contract to get a few new phones, a game console or two, and a couple TVs.
oezi•3h ago
no-name-here•3h ago
incognito124•3h ago
That's definitely one way to say "it doubled this year"
computerphage•3h ago
no-name-here•15m ago
Mistletoe•3h ago
https://www.bigtrends.com/education/lessons-from-the-past-10...
jcattle•3h ago
post-it•2h ago
> We fit a trend to the quarterly quantities of AI compute sold, as measured in H100-equivalents. We find that computing capacity has been growing by 3.3x per year, equivalent to a doubling time of 7 months.
This is just a graph of cash changing hands. I transferred a loonie from one hand to the other one trillion times this morning and have the largest data centre in the world.