frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

How many GPUs is 1M/B/T tokens?

https://cedana.com/resources/tokens-to-gpus/
12•kmavm•21h ago

Comments

dannyw•36m ago
I’m sorry but this looks like vibe coded marketing slop that’s highly inaccurate.

For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.

Secondly, it seems to base all benchmarks off a batch size of 1.

Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.

And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.

I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.

cwmoore•20m ago
Thank you. For me the hint was: “How many GPUs is 1M/B/T tokens?“

Those units…

Also “is” is load-bearing

b112•4m ago
What! If you start running multiple models through the same GPU, you'll end up with cache issues like spectre on Intel CPUs!

This is how viruses swap DNA, and it's how AGI and sentience will accidentally happen. A mix of code here, a combine of cache data there, and AGI inadvertently appears!

I don't let anyone in my company do this, and you should not do so either.

It's all fun and games until you spark the literal apocalypse, dannyw!

So yes, the website is 100% accurate for safe LLM usage.

DiabloD3•3m ago
This calculator is kinda useless. Different models work differently, you cannot generically define them by parameter count, and the models it does list seem to just be it prefilling the param counts for you.

It also misses most cards being used for inference, only limiting it to relatively recent Nvidia cards. Even if this was meant purely for datacenters, it would be useless for AMD customers.

And as for different models being different, it also doesn't understand kv cache size, how many concurrent sessions you need to manage, the compute cost overhead of kv cache quantization, nor the compute cost overhead of model quantization. It also doesn't know what MTP nor DFlash is, and cannot increase your effective tps to match.

As an example: compare Unsloth's quant of Qwen 3.8 Q4_K_M (4.5 BPW), a decent smaller model, without MTP, you'd get a baseline of, say, around 3-6 TPS per 100W. With the built in MTP model, you're closer to 6-12 TPS per 100W; then, switch quants to Byteshape's IQ4-XS (3.84 BPW) and use their suggested Dflash drafter, you're now in the realm of 15-30 TPS

... yet if I scribble into the calculator "27B, 27B active, 4-bit, 1B per one day", it seems to be the low end of my non-MTP figure: 1B per day = 11574 per second, it recommends 3x B200s, each B200 is 1200W, so 11574 / (1200*3) = 3.215, yet real world results would be 2x to 10x higher.

Also, quoting from the website, "A100 and V100 results for the large models are theoretical.". The small scale inference people absolutely know how these perform and it is well-enough documented.

The whole thing seems like AI slop.

GitHub Incident with Git Operations, Pull Requests and Actions

https://www.githubstatus.com/incidents/djlmxz2zd0j7
106•gagan2020•45m ago•60 comments

Shipping JPEG XL in Chrome

https://developer.chrome.com/blog/jpeg-xl-in-chrome
306•AshleysBrain•4h ago•172 comments

Google Playground

https://labs.google/playground
78•trollied•2h ago•49 comments

Why Were Victorian Elites So Effective?

https://worksinprogress.co/issue/the-seven-vices-of-highly-effective-victorians/
25•karakoram•53m ago•31 comments

A font recreated from photographs of classic Commodore 64 keycaps

https://github.com/szabadkai/c64-keyboard-font/
263•sohkamyung•6h ago•46 comments

Nobel Prize in Chemistry 2026 to Henri B. Kagan and Kenso Soai

https://www.nobelprize.org/prizes/chemistry/2026/press-release/
194•sasvari•6h ago•30 comments

Show HN: A walkable 3D art history museum built from Wikipedia

https://artmuseum.artfrompixels.com/
68•jasontr•3h ago•38 comments

AI-assisted proof of optimal packing for 11 squares

https://github.com/Queuingtheorydotcom/11SquaresFormalized
27•bluepeter•1h ago•19 comments

Write Like It's 1866: LLMs Relearn Telegraphese

https://fiveminutesforward.com/post/2026-10-04-telegraph-test/
53•Theory42•3h ago•37 comments

Rust's derive often implies inline

https://yossarian.net/til/post/rust-s-derive-often-implies-inline/
76•woodruffw•3d ago•11 comments

Mallet Head Angle

http://www.timberframe-tools.com/tools/mallet-head-angle/
38•frogulis•1d ago•9 comments

Animated ASCII Art for Web Pages

https://ascii.rest/
14•turrini•57m ago•1 comments

Across the Globe, People Increasingly Say Social Media Is Harming Democracy

https://www.pewresearch.org/global/2026/10/01/across-the-globe-people-increasingly-say-social-med...
16•karakoram•39m ago•3 comments

Show HN: Procinsh – A 3D Linux process inspector

https://github.com/akawashiro/procinsh
23•a_kawashiro•3h ago•6 comments

How many GPUs is 1M/B/T tokens?

https://cedana.com/resources/tokens-to-gpus/
12•kmavm•21h ago•4 comments

Show HN: AstroHelm – Use your phone camera to aim a telescope or telephoto lens

https://astrohelm.app/
67•HeavenFox•3d ago•16 comments

Strands Decider 2B: a small, open-source, decision model

https://strandsagents.com/blog/introducing-strands-decider/
247•gmays•14h ago•69 comments

Anti-Patterns in Software Blogging

https://refactoringenglish.com/blog/anti-patterns-software-blogging/
23•ilreb•2h ago•4 comments

Sharing AI progress in mathematics

https://openai.com/index/sharing-ai-progress-in-mathematics/
1123•OfficialTurkey•17h ago•1212 comments

God of War on PSP, recompiled to WebAssembly and running in the browser

https://github.com/snuri00/psp-web-recomp
22•sn001•4h ago•9 comments

House with 15m underground tunnels for sale for 300k

https://www.readingchronicle.co.uk/news/26612080.house-15m-underground-tunnels-sale-300k/
97•librasteve•3h ago•106 comments

SynthID Detector

https://synthid.com/
52•ilreb•1h ago•53 comments

Wood Tape (2004)

http://gamesbyemail.com/WoodTape/Default.htm
21•NaOH•1d ago•3 comments

Mistral Large 4

https://mistral.ai/news/mistral-large-4/\
1949•Philpax•1d ago•1165 comments

ESP32-C3 Adblock

https://github.com/M-Abozaid/esp32-c3-adblock
165•jayhoon•14h ago•63 comments

Tell HN: GitHub refuses to remove cracked copies of my software after a month

450•IvanK_net•21h ago•243 comments

What is Codemode

https://lucumr.pocoo.org/2026/10/6/codemode/
137•Tomte•1d ago•59 comments

Gallery of Processor Cache Effects (2010)

https://igoro.com/archive/gallery-of-processor-cache-effects/
42•porridgeraisin•2d ago•4 comments

The cost of lies: A Mineserver story

https://www.jeremyreimer.com/rockets-item.lsp?f=true&p=272
165•luu•3d ago•69 comments

Google Playground: Create and play custom games

https://blog.google/innovation-and-ai/technology/ai/playground-experimental-gaming-platform/
78•acossta•3h ago•71 comments