frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Vote on which of Hacker News' challenges for AI have been met

https://stoppels.ch/goalposts/
21•stabbles•1h ago

Comments

ben_w•50m ago
Very pleased one of my predictions was totally wrong: https://news.ycombinator.com/item?id=23252711

Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.

FabCH•44m ago
Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.

In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.

tripleee•30m ago
> An LLM today sure can do many many many business-speak conversion tasks

Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)

You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing

Dylan16807•25m ago
You can't ignore the rest of the sentence. "every other task their business does" "everyone will be out of a job"

This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.

ben_w•17m ago
I'm not always precise with my language, but business tasks can be pretty broad, I think "arbitrary new tasks" is not an unreasonable rephrasing on my part?

Consider I was replying to this:

> So are we all going to be out of a job?

While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.

If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.

People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...

vlyan•33m ago
so the conditions for your prediction simply haven't been met yet.

if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.

ben_w•27m ago
The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
tripleee•32m ago
> reliably convert business-speak into efficient bug-free code

I actually think this would take AGI to solve, which makes me optimistic about the future of software development.

All the benchmarks are currently testing against automated tests the AI can use as an oracle

simianwords•35m ago
I made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I'm on track to winning.

https://news.ycombinator.com/item?id=48517353

I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic

https://news.ycombinator.com/item?id=48500827

I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.

mrweasel•13m ago
You're very lucky that market value and actual value isn't the same thing.
ngruhn•7m ago
I love it! Kinda wholesome how that heated discussion ended with that bet.
delichon•32m ago
If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
Retr0id•30m ago
Alternatively, you can bet on your predictions. If you're wrong, you lose money.
bmenrigh•32m ago
At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
happytoexplain•26m ago
Right - people on HN are generally reasonable about objective things. The vast majority of comments (outside those chosen for this website) are not "AI will never ..." but rather, "AI does not currently ...". Of course the further you go back (I'm seeing a lot of comments from ten years ago!) the more skeptical they get, obviously. That's a funny thing to go back and see with modern context, but it doesn't really call for snideness/mockery (something I think is sadly increasing on HN).
jerf•21m ago
Well, I can answer this one: https://stoppels.ch/goalposts/?c=40662140

jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".

"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.

"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."

The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."

Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.

travisgriggs•26m ago
How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
ErrantX•18m ago
What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.

And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.

But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.

That alone tells you a lot IMO

simianwords•16m ago
Here's a compilation of HN commenters that never believed that AI could solve Millennium problems (or same in spirit)

> I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon https://news.ycombinator.com/item?id=38433655

> Let's talk when we've got LLMs proving the Riemann Hypothesis (or any mathematical hypothesis) without any proofs in the training data. I'm confident in my belief that an LLM can't do that, and will never be able to. LLMs can barely solve elementary school math problems reliably.

https://news.ycombinator.com/item?id=42331654

> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up. It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.

The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.

https://news.ycombinator.com/item?id=41522605

> Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.

https://news.ycombinator.com/item?id=38435909

> LLMs cannot reason or use mathematics - in a way, they don't know what they are talking about. Why would such technology lead to superhuman smarts?

https://news.ycombinator.com/item?id=35752293

> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.

> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all. We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.

> It is like expecting a real parrot to say words it has never heard before.

> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM

https://news.ycombinator.com/item?id=41525962

mrweasel•11m ago
The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.

Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.

Retr0id•9m ago
Heh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943

"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."

The vote is currently 64% yes, 18% no.

Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...

Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).

pitched•17m ago
> cannot do precise things like coding software since humans will never be able to use natural language to specify their requirements.

To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.

Clef: Open-source decision models, and new RL fine-tuning platform

https://blog.cloudflare.com/clef-decision-models/
239•jasondavies•2h ago•90 comments

RIP, vector database

https://turbopuffer.com/blog/rip-vector-database
157•razin•2h ago•43 comments

StreetComplete on iOS is now in public beta

https://github.com/streetcomplete/StreetComplete/issues/5421
427•Snowly•7h ago•100 comments

RacketCon Is Saturday

https://con.racket-lang.org/
91•spdegabrielle•3h ago•23 comments

Show HN: Vote on which of Hacker News' challenges for AI have been met

https://stoppels.ch/goalposts/
21•stabbles•1h ago•22 comments

Ask HN: Who is hiring? (October 2026)

71•whoishiring•3h ago•71 comments

Cloudflare K2: serverless event streams

https://blog.cloudflare.com/cloudflare-k2-streams/
110•elffjs•4h ago•37 comments

Show HN: Open-source model routing for coding agents at Astra-level performance

24•adchurch•1d ago•3 comments

Various Projects Find Hidden SDR Capabilities in ESP32 Microcontrollers

https://www.rtl-sdr.com/various-projects-independently-find-hidden-sdr-capabilities-in-esp32-micr...
74•nkw•3h ago•8 comments

How to speed up the Rust compiler in September 2026

https://nnethercote.github.io/2026/09/30/how-to-speed-up-the-rust-compiler-in-september-2026.html
189•trickypr•6h ago•93 comments

Gemini 4 Argon

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
1611•bradleyg223•22h ago•1077 comments

Ask HN: Who wants to be hired? (October 2026)

42•whoishiring•3h ago•160 comments

Context Language Models

https://arxiv.org/abs/2609.37725
44•emersonmacro•4h ago•9 comments

Identity Management for Agentic AI [pdf] (2025)

https://openid.net/wp-content/uploads/2025/10/Identity-Management-for-Agentic-AI.pdf
52•cgeier•3h ago•9 comments

Lightweight PDF parser with layout, tables, formulas and bounding boxes

https://github.com/beatrizalmeidaf/papero-pdf-text-extractor
35•beatrizalmeidaf•2h ago•7 comments

ParadeDB Search Performance Improvements

https://www.paradedb.com/blog/opening-a-closed-tin
34•craigkerstiens•1h ago•7 comments

GPT-Synopsys: Frontier Intelligence to Revolutionize Chip Design

https://news.synopsys.com/2026-09-30-OpenAI-and-Synopsys-Announce-GPT-Synopsys-Frontier-Intellige...
142•giuliomagnifico•8h ago•76 comments

Polyedergarten: Garden of Paper Polyhedron Models

https://www.polyedergarten.de/e_index.htm
32•isaacimagine•3h ago•4 comments

Bez: Generating a browser engine from specs and tests

https://tangled.org/burrito.space/bez
5•nerdypepper•47m ago•0 comments

Red Hat being phased out of existence?

https://techrights.org/n/2026/10/01/Red_Hat_Being_Phased_Out_of_Existence_Like_Many_Other_Compani...
116•amcclure•3h ago•55 comments

OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network

https://github.com/maanHimself/OpenDLSS-NR
224•sagacity•1d ago•104 comments

Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

https://www.techpowerup.com/353296/micron-ceo-says-memory-supply-will-be-much-tighter-in-2027-and...
251•speckx•6h ago•294 comments

Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones

https://www.404media.co/cops-can-bypass-iphone-automatic-inactivity-reboot-graykey/
173•speckx•4h ago•126 comments

Truemetrics (YC S23) Is Hiring a GTM Founder's Associate

https://www.ycombinator.com/companies/truemetrics/jobs/THLEzXI-gtm-founder-s-associate
1•truemetricsIngo•9h ago

How to set up SPF, DKIM, and DMARC for your sending domain

https://mailfully.com/blog/spf-dkim-dmarc-setup
19•spy888•4h ago•4 comments

Figma restricts MCP access to whitelisted clients, excluding Pi

https://twitter.com/GayaniFigma/status/2105295629941350454
125•thdr•3h ago•68 comments

Book of Shapes – Collection of minimal, generative and customizable SVG-patterns

https://bookofshapes.com/
223•eustoria•2d ago•16 comments

FTC is investigating OpenAI, Anthropic and other AI companies over product risks

https://www.cnbc.com/2026/09/30/ftc-ai-probe-openai-anthropic.html
170•dgellow•5h ago•113 comments

Returning from vacation? The government can search your phone without a warrant

https://arstechnica.com/tech-policy/2026/09/immigration-advocate-sues-border-agents-for-demanding...
304•rbanffy•7h ago•287 comments

Why the Bronze Age Collapsed

https://www.worksinprogress.news/p/why-really-caused-the-bronze-age
377•AnodicElegy•2d ago•259 comments