VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.
Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.
The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...
There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.
If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.
The market is legally locked down and we're in hostage pricing mode.
And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...
Nobody is coming to save us. That's our job.
Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There's no need to posit cartel behavior and a fair amount of evidence that there is none.
wonder what voting would be like?
gamer vote ++
datacenter hater vote --
datacenter lobby ++
micron lobby --
Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.
Plus expanding other existing facilities.
These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.
Samsung and HK Hynix also have fabs under construction and planned.
CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.
Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.
> The Micron CEO just recently said this is the exact plan
CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
1660 ti, 4790k, 16gb ddr3
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
Is OpenAI coming in $20B under a sign of "freaking out"?
People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.
The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.
And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at end.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
Would be cool if they added it.
There are some quirks if your harness use unsupported features of course.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
but DS 4.1 Flash is good enough for most tasks
Can't you just say "shrank to 1/437th the size"? It's not that hard.
see https://artificialanalysis.ai/models/releases/comparisons?co...
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
verdverm•18h ago