frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Show HN: Simple algorithm and color space to generate diverse skin tones

https://toneyalexander.github.io/inclusive-color-space/
193•automatoney•2h ago•47 comments

Ray Bradbury's "There Will Come Soft Rains" is set today (2026-08-04)

https://short-stories.co/@raybradbury/there-will-come-soft-rains-6k8vr4xxlnmj
470•askvictor•7h ago•211 comments

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

https://arxiv.org/abs/2602.16763
30•doppp•1h ago•31 comments

DeepSeek V4 Flash on a Single AMD MI300X

https://github.com/ryanzhou/deepseek-v4-flash-mi300x
283•zhoutong•7h ago•62 comments

Keyv and friends compromised in active Shai-Hulud supply chain attack

https://www.aikido.dev/blog/keyv-and-friends-compromised-in-npm-supply-chain-attack
150•cimi_•6h ago•72 comments

Germany Records Historic 12B KWh Solar Feed-In in July 2026

https://solarquarter.com/2026/08/03/germany-records-historic-12-billion-kwh-solar-feed-in-in-july...
102•johnbarron•3h ago•102 comments

Truemetrics (YC S23) Is Hiring in Berlin – GTM Lead

https://www.ycombinator.com/companies/truemetrics/jobs/bIQQ7tP-founding-gtm-lead
1•truemetricsIngo•34m ago

LLMs reward expertise

https://www.seangoedecke.com/llms-reward-expertise/
1257•MaxMussio•20h ago•517 comments

Online ad giant Adform was hacked, proving once again why ad blockers are needed

https://this.weekinsecurity.com/online-advertising-giant-adform-was-hacked-proving-once-again-why...
103•speckx•2h ago•26 comments

Apple says more ex-employees may have taken confidential data to OpenAI

https://techcrunch.com/2026/08/04/apple-says-more-ex-employees-may-have-taken-confidential-data-t...
107•thewebguyd•1h ago•73 comments

Buckminster Fuller: everything I know

https://www.bfi.org/about-fuller/everything-i-know/
92•simonebrunozzi•6h ago•29 comments

Dates That Don't Exist (2015)

https://blog.yossarian.net/2015/06/09/Dates-That-Dont-Exist
70•EndXA•3d ago•47 comments

First Came the DOGE Cuts, Then Came the Wildfires

https://www.outdoorlife.com/conservation/doge-cuts-forest-service-wildfires/
8•hn_acker•27m ago•2 comments

Harness Engineering for Self-Improvement

https://lilianweng.github.io/posts/2026-07-04-harness/
228•tosh•11h ago•44 comments

Xbox goes down. You can't play games you own on disc

https://birchtree.me/blog/xbox-goes-down-you-cant-play-games-you-own-on-disc/
407•surprisetalk•5h ago•472 comments

Looking inside a 1970s PROM chip that stores data in microscopic fuses (2019)

https://www.righto.com/2019/07/looking-inside-1970s-prom-chip-that.html
26•Jimmc414•3d ago•4 comments

The AI Demand Bubble

https://www.wheresyoured.at/the-ai-demand-bubble/
60•7777777phil•1h ago•16 comments

Why Large Language Models Fail at Tabular Prediction

https://arxiv.org/abs/2608.02412
76•sbulaev•7h ago•30 comments

Show HN: Fine-tune an 8B model on a 4 GB laptop GPU

https://github.com/MakazhanAlpamys/Soup
92•MakazhanAlpamys•6h ago•14 comments

RCade: The Arcade Cabinet with CI/CD Deployment, Custom Graphics Card for CRT [video]

https://www.youtube.com/watch?v=W-OpIbLUOU0
28•evakhoury•5d ago•7 comments

Why etymologies matter: How tracing words can illuminate history (2024)

https://resobscura.substack.com/p/why-i-love-etymologies
72•benbreen•3d ago•16 comments

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

https://github.com/leonickson1/Swiftlet
278•leonickson•1d ago•128 comments

Rebuilding and analysing 4 years of Wordle stats from WhatsApp chat logs

https://blog.omgmog.net/post/rebuilding-wordle-stats-from-whatsapp/
17•surprisetalk•1d ago•4 comments

Ten advances in mathematics and theoretical computer science

https://openai.com/index/ten-advances-in-mathematics/
605•milkshakes•1d ago•888 comments

Devtools must be open source

https://blog.exe.dev/devtools-must-be-open-source
694•bryanmikaelian•1d ago•227 comments

Amazonian civilization had estimated 3M people in 3% of forest area

https://www.science.org/content/article/odd-shapes-hidden-dense-amazon-rainforest-reveal-sprawlin...
248•marojejian•6d ago•180 comments

Homebench – Benchmark local LLMs for speed, memory, and quality

https://github.com/david-g-3654/homebench
48•davai-g•7h ago•3 comments

There Will Come Soft Rains (1950) [pdf]

https://users.wpi.edu/~zrbutzke/Docs/BradburyStories(1).pdf
291•pmg101•18h ago•106 comments

Webb telescope finds signs of ancient disaster for Neptune's moons

https://www.reuters.com/science/webb-telescope-finds-signs-ancient-disaster-neptunes-moons-2026-0...
17•Teever•1h ago•3 comments

Learning-Rust.Github.io: Rust Programming Language Tutorials for Everyone

https://learning-rust.github.io
37•dumindunuwan•7h ago•7 comments
Open in hackernews

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

https://arxiv.org/abs/2602.16763
26•doppp•1h ago

Comments

otterdude•46m ago
This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.

If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence.

I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm

“The seeker after truth is not one who studies the writings of the ancients and, following his natural disposition, puts his trust in them, but rather the one who suspects his faith in them and questions what he gathers from them, the one who submits to argument and demonstration and not the sayings of human beings whose nature is fraught with all kinds of imperfection and deficiency. Thus the duty of the man who investigates the writings of scientists, if learning the truth is his goal, is to make himself an enemy of all that he reads, and, applying his mind to the core and margins of of its content, attack it from every side. he should also suspect himself as he performs his critical examination of it, so that he may avoid falling into either prejudice or leniency.” - ibn al-Haytham

graemep•38m ago
I am wondering whether the reason he needed to say it was because he was arguing with those who did out their trust in the writings of the ancients.
scotty79•37m ago
Do you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?
freejazz•30m ago
Can't call it AI like that without discrediting yourself. You mean LLMs?
logicchains•25m ago
Talk about moving the goalposts. Pray tell, exactly what must an LLM do before you're willing to consider it AI? Be specific, otherwise you're just woo-mongering.
otterdude•24m ago
Jumping in here, frankly I hate the trend of calling every type of automation intelligence.

Most "AI" is really an optimization algorithm in software tools, same as its always been. This really isnt anything new, aside from adding a chatbot / MCP interface to the same tools.

scotty79•9m ago
When a Big Killing Robot comes to murder you be sure to always call it BKR and don't discredit yourself by calling it AI.
otterdude•28m ago
Its not really that surprising when models are trained on the exams
adrianN•32m ago
I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
tsunamifury•14m ago
I'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work.

We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.

warkdarrior•9m ago
If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example)

I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?

Jensson•12m ago
You aren't contradicting the person.
runarberg•6m ago
[delayed]
0xdeadbeefbabe•27m ago
The seeker of truth must also hold his breath.
logicchains•26m ago
>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.

It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.

Jensson•7m ago
Yes, and good senior software engineer is ahead of fable, but benchmarks can't capture that either.

We already know from testing humans that test scores don't correlate that well with how effective a person is at work. Same applies here, we just aren't that great at making good tests.

heaney-555•15m ago
>This seems to be the end of the road for LLM's

This is an amazingly ignorant thing to say given the current pace of progress.

Jensson•10m ago
There is high rate of progress in specific domains, not high rate of progress in generalness. The models haven't gotten generally smarter, for things they didn't focus on the models are just as bad as a year ago.
hiddencost•15m ago
Weird moment for this take. We're seeing some of the fastest and most impressive progress ever right now.

Frontier labs have categorically different & better set ups for evaluation, they're fine. It's work but it's not a crisis.

buckle8017•37m ago
Slop

> We find that nearly half of the our bench- marks exhibit saturation

joeyagreco•35m ago
This leads me to believe it's NOT slop lol
jdiff•34m ago
It's a grammatical error, sure, where is the indication of slop?
gertlabs•21m ago
I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better than any non-aggregator benchmark, and likely at lower cost to run.

Data at https://gertlabs.com/rankings

nwienert•9m ago
[delayed]
behnamoh•20m ago
This is AI slop. They didn't even change the plots default template.
hagen8•13m ago
Check out https://agents-last-exam.org/ there is still room for improvements!
tsunamifury•9m ago
I think its been pretty clear that in abnsense of clear use cases that are monetizable many model providers have been benchmaxxing on abstract or low utility average user performance.

This results in a lot of "oh wow it can do math I dont care about" and "it can't code a lot, but not well" outcomes instead of the core needs:

1) Cheaper faster and real time 2) Long walk capable without losing attention while rescoring goals over updated enviroment 3) Specific domain knowledge that can be trained quickly into the model (how we do work in this specific case)

astro1234•24m ago
I work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see.

I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the trend in capabilities show no sign of slowing down.

So I think people’s feelings clash with reality, and that’s because releases are more frequent and the jumps between releases are smaller, but the growth in capabilities _over time_ has not changed for the better or worse over a very very long period of time.

otterdude•12m ago
Benchmarks saturate around 80-90%?

This is not "Acing" a test, this is hitting a wall.