frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Muse Spark 1.3

https://developer.meta.com/ai/models/muse-spark/
201•bvaldivielso•1h ago•110 comments

I wanna live an NPC life

https://signalundefied.bearblog.dev/i-wanna-live-an-npc-life/
84•conferza•1h ago•48 comments

Gemini 3.8 Flash and 3.8 Flash Cyber

https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-c...
692•bratao•6h ago•407 comments

Google avoids a breakup of its ad tech business

https://www.nytimes.com/2026/09/02/technology/google-ad-tech-remedies.html
132•donohoe•6h ago•65 comments

Fable 5.1 World Modeling

https://github.com/PhiloLabs/fable51-worlds
46•surreal_•1h ago•11 comments

Wendell Berry has died

https://www.nytimes.com/2026/08/31/us/wendell-berry-dead.html
82•Curiositry•1d ago•37 comments

Nango (YC W23) is hiring across eng, product and GTM (SF and remote)

https://nango.dev/careers
1•bastienbeurier•17m ago

Can I opt out of my input or output data being used for training?

https://help.mistral.ai/en/articles/455207-can-i-opt-out-of-my-input-or-output-data-being-used-fo...
329•teekert•8h ago•140 comments

AI Agents and the Refactoring That Never Happens

https://www.rosenfeld.page/articles/programming/2026_09_02_ai_agents_and_the_refactoring_that_nev...
27•rosenfeld•1h ago•35 comments

Qantas Airbus A380 engine failure in 2010 (2023)

https://admiralcloudberg.medium.com/a-matter-of-millimeters-the-story-of-qantas-flight-32-bdaa62d...
38•gumby•2h ago•14 comments

Embedded Rust RTOS vs. C RTOS

https://tweedegolf.nl/en/blog/65/async-rust-vs-rtos-showdown/
36•kooi•2h ago•15 comments

Biggest dark matter detector spots a single weird particle

https://www.science.org/content/article/world-s-biggest-dark-matter-detector-spots-single-weird-p...
211•randycupertino•7h ago•60 comments

Reverse Engineering Unknown File Formats with ImHex

https://werwolv.net/posts/file_format_reverse_engineering/
8•carlos-menezes•2d ago•0 comments

Exit the Cave

https://turtlespace.blog/p/exit-the-cave
172•akkartik•7h ago•45 comments

Altair Basic Interpreter Source Code (1975) [pdf]

https://images.gatesnotes.com/12514eb8-7b51-008e-41a9-512542cf683b/34d561c8-cf5c-4e69-af47-3782ea...
12•Eridanus2•1h ago•2 comments

Aging Brains Blend Memories Together Instead of Just Forgetting Them

https://studyfinds.com/aging-brains-blend-memories-together-instead-of-forgetting-them-study-finds/
158•mdp2021•8h ago•75 comments

A practical guide to running 8x RTX PRO 6000's

https://www.gpupartner.com/blog/a-practical-guide-to-running-8x-rtx-pro-6000s
17•Retell15•2d ago•27 comments

Commodore 64 released September 1, 1982

https://dfarq.homeip.net/commodore-64-released-september-1-1982/
300•giuliomagnifico•12h ago•156 comments

We could save petabytes of cache storage with Zstandard and Pingora

https://blog.cloudflare.com/cache-transcoding/
41•torutofu•1d ago•13 comments

SteamdDB Joins Nexus Mods

https://www.nexusmods.com/news/15597
86•HelloUsername•5h ago•46 comments

Introducing Muse Spark 1.3

https://research.meta.ai/blog/introducing-muse-spark-1-3
51•scrlk•1h ago•10 comments

Using Cloudflare Workers and reCAPTCHA v3 for a Static Site Contact Form

https://nooshu.com/blog/2026/03/09/using-cloudflare-workers-and-recaptcha-v3-for-a-static-site-co...
16•speckx•2h ago•3 comments

A Selection of Los Alamos Rolodex Business Cards

https://clui.org/collections/los-alamos-business-cards/selection-cards
121•1970-01-01•2d ago•26 comments

Making the Internet Boring

https://cemrehancavdar.com/2026/08/30/making-the-internet-boring/
45•zdw•3d ago•25 comments

Poisson Disk Sampling

https://stripeacross.com/posts/poisson-disk-sampling/
104•vismit2000•7h ago•16 comments

Paint.net 5.2 alpha now runs on Linux

https://forums.paint.net/topic/134562-paintnet-52-alpha-build-9739/
154•judah•3h ago•125 comments

A note on subscription prices from LWN

https://lwn.net/Articles/1090585/
650•rwky•8h ago•128 comments

WebLLM: high-performance in-browser LLM inference engine

https://github.com/mlc-ai/web-llm
67•saikatsg•7h ago•16 comments

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/
255•jakobgreenfeld•7h ago•119 comments

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

https://frontierharness.org
61•shiqimei•5h ago•37 comments
Open in hackernews

A practical guide to running 8x RTX PRO 6000's

https://www.gpupartner.com/blog/a-practical-guide-to-running-8x-rtx-pro-6000s
17•Retell15•2d ago

Comments

varispeed•56m ago
> (~9.2 million tokens node-wide at 4k context).

stopped reading after that. What 4k context would be usable for?

erdaltoprak•21m ago
Not long-horizon coding but for a lot of other things like batch processes with structured outputs, quick checks/fixes, making sense of unstructured data etc..
monster_truck•12m ago
So much more than you'd realize!

That's a solid 2200 words to spend on operating parameters and conveying state, leaving a generous 700 word window for them to decide and respond in.

When the bonsai/prism 1bit models dropped and I saw how many prompts a minute I could get from a dusty m2 mini I started hooking it up to all sorts of shit, like a traffic simulator that translates the car state/surroundings/immediate goal into text, it responds with a seqeunce of actions defined in the system prompt, which then get translated back into NPC input.

What I was hoping for here was that it would result in fucking chaos, all sorts of stupid decisions and epic car accidents. I cannot overstate my disappointment (and terror) when they were perfectly reasonable, safe drivers. I had to cut the tire grip by 75% without telling them and make them control twice as many cars to delay their ability to respond before I saw anything resembling an enjoyable traffic accident.

schaefer•56m ago
> We currently have 14x nodes of CG480-S6053 ready to ship.

Oh, okay, so this is an ad.

I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)

baron3dl•53m ago
You must be the only one 'round here without a stack of RTX PRO 6000s, 8 high, that you're not sure how to use.
schaefer•48m ago
what I have is a DGX Spark to play and learn on, and an OpenRouter account for actual work.

I mostly rent my tokens.

baron3dl•39m ago
that was an attempt at monocle and top hat irony.

edit: how do you feel about paying 5% on every token's cost to stripe? it made me cancel, and sign up directly with a few providers.

schaefer•20m ago
> attempt at ... irony

For the record, I did laugh, I just didn't type it out.

> how do you feel about paying 5% on every token's cost to stripe?

Life is compromise. In a perfect world, I'd love to buy a linux native system comparable to the Mac Studio that could run something in the ballpark of Deepseek V4 flash "fast" at home.

Not only does that not exist, but even if it did I couldn't justify the cost. Between work and home my annual token spend is maybe 3k.

So... paying a 5% fee on a service that allows me to not shell out 10k on hardware at home still seems like a pretty sweet deal.

--

Plus, I still am learning a ton and playing lots. It's hard to imagine anything beating openrouter for that. So much exposure to all the latest LLMs...

christkv•53m ago
600W * 8 just for the GPUs when maxed out (besides the cost). Def nothing for my home lab.
touisteur•9m ago
I'm curious whether actual inference workloads actually push to 600W (and not 350W) and what the last 250W get you. Rare is the (generic gpu) workload where I get >5%, some rare light inference benchmarks up to 10%...
RachelF•52m ago
For those who can't afford RTX 6000's you can unlock around 20% increased card to card speed on consumer GPUs using this library:

https://github.com/aikitoria/open-gpu-kernel-modules

The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.

robotnikman•33m ago
Now if only I could afford 8 RTX PRO 6000's
CamperBob2•26m ago
Start with 4, see my other comment. The recent GLM, Qwen and DeepSeek releases are amazingly promising.
estebarb•19m ago
Now, if I could afford 4...
jplusequalt•9m ago
>Start with 4

Sure, let me just buy $60,000 worth of GPUs to run a *quantized non-frontier model*.

For that price you could:

- put a down payment on a home in a large % of the US

- buy a brand new car in cash (possibly two!)

- take a long sabbatical and travel the world

- pay all 4 years or your child's college tuition

CamperBob2•30m ago
4x RTX 6000 Blackwell cards is a good place to be if you can't swing 8 of them, or if you don't have the power or cooling to run that many. A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision [1] from a US-standard 120V 20A circuit when derated to 300W, and give you a better pelican than Fable 5.1 [2]. What's not to like?

(Edit: I'm mistaken here, the pelican didn't come from Flash on 4 cards but from the full GLM 5.3 model on 8. But the Flash model is still crazy good for its size.)

1: https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4

2: https://crimson-jeri-74.tiiny.site/

nojs•7m ago
> A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision

It actually runs fine at FP8 on this hardware too, with the full 1M context.

jimmoores•28m ago
These people have zero idea what they're doing. Not a single mention of pipeline parallelism that would actually make the setup useful to run a big model.
kelmoran•23m ago
I feel like you would want to run 8 smaller models separately for quantity of raw output. 1 big model is slow and isnt guaranteed to make no mistakes.
monster_truck•1m ago
That's not quite how it works. Throwing Deepseek V4 Flash on 4 of these would net you something like >200tk/s for 16 concurrent requests, that's 600 million _output_ tokens a month. Guess what happens when you use 8
schaefer•15m ago
can you point to a write up that discusses what you're talking about?

because I would read it.

itkovian_•10m ago
I can’t stand it. Very engineering-y over specified formal language around a complete lack of core understanding. Is damaging other people read this and try to learn things from it.
inventor7777•23m ago
This will really help when my 8 RTX PRO 6000s ship. /s
xyst•17m ago
SLI is relevant again
srcreigh•11m ago
Makes you realize how insane the M5 Ultra Mac Studio is. 1.2GB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.
teaearlgraycold•2m ago
These GPUs are extremely inflated in price because Nvidia effectively has a monopoly on hardware that is used to train models. Apple Silicon tends to have good inference software available but as soon as you want to train even a YOLO model bits and pieces fall back to software implementations. Try to train an LLM and it'll get even worse.

The M3 Ultra's GPU performance is around a 4070 Ti. The M5 Ultra more like a 5080. They're both amazing deals compared to Nvidia for local inference because of their massive pool of high bandwidth memory. But a single RTX PRO 6000 should be 2 or 3x the compute of an M5 Ultra.

kmike84•6m ago
Pass. When articles keep mentioning models like DeepSeek R1, or Llama 3.1, or Qwen3 32B, it is a pretty robust indicator of AI slop. LLMs love to suggest DeepSeek R1, etc. - training data cut-off?

No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.