It also misses most cards being used for inference, only limiting it to relatively recent Nvidia cards. Even if this was meant purely for datacenters, it would be useless for AMD customers.
And as for different models being different, it also doesn't understand kv cache size, how many concurrent sessions you need to manage, the compute cost overhead of kv cache quantization, nor the compute cost overhead of model quantization. It also doesn't know what MTP nor DFlash is, and cannot increase your effective tps to match.
As an example: compare Unsloth's quant of Qwen 3.8 Q4_K_M (4.5 BPW), a decent smaller model, without MTP, you'd get a baseline of, say, around 3-6 TPS per 100W. With the built in MTP model, you're closer to 6-12 TPS per 100W; then, switch quants to Byteshape's IQ4-XS (3.84 BPW) and use their suggested Dflash drafter, you're now in the realm of 15-30 TPS
... yet if I scribble into the calculator "27B, 27B active, 4-bit, 1B per one day", it seems to be the low end of my non-MTP figure: 1B per day = 11574 per second, it recommends 3x B200s, each B200 is 1200W, so 11574 / (1200*3) = 3.215, yet real world results would be 2x to 10x higher.
Also, quoting from the website, "A100 and V100 results for the large models are theoretical.". The small scale inference people absolutely know how these perform and it is well-enough documented.
The whole thing seems like AI slop.
dannyw•36m ago
For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.
Secondly, it seems to base all benchmarks off a batch size of 1.
Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.
And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.
I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.
cwmoore•20m ago
Those units…
Also “is” is load-bearing
b112•4m ago
This is how viruses swap DNA, and it's how AGI and sentience will accidentally happen. A mix of code here, a combine of cache data there, and AGI inadvertently appears!
I don't let anyone in my company do this, and you should not do so either.
It's all fun and games until you spark the literal apocalypse, dannyw!
So yes, the website is 100% accurate for safe LLM usage.