frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Can AI automate AI R&D yet?

https://epoch.ai/publications/innovationeval
5•merksittich•3h ago

Comments

janalsncm•34m ago
In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.

For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.

Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.

And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.

rmunn•33m ago
Short version of the article: no, not even close.

Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.

My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.

janalsncm•23m ago
A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.

If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.

Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.

rmunn•15m ago
I'd classify that as an entirely different category than AI self-training. What you're describing could have been done with a short script, though which parameter to tweak and how to tweak it would be difficult to automate with a non-LLM script, so the LLM's being able to parse the error message and base the tweak on the content of the error is a definite improvement to the process there.

But I'd classify this as LLM being used to automate a sysadmin task, rather than calling that self-training.

charcircuit•21m ago
I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.
thoughtpeddler•18m ago
How much of this can change if subsequent training runs produce models that are much better at abduction?

REA Reverse – Engineer Anything

https://rea.tools/
58•modinfo•1h ago•10 comments

Cloudflare acquires Deno

https://deno.com/blog/cloudflare
1068•ilreb•12h ago•558 comments

11 of 23 Core Open Source Projects Run on 1 or 2 People

https://linuxstans.com/11-of-23-core-open-source-projects-run-on-1-or-2-people/
46•dxs•1h ago•15 comments

Triple-A Minesweeper

https://minesweeper.mikelacher.com/
659•robin_reala•9h ago•122 comments

Our $445M Series D

https://oxide.computer/blog/our-445m-series-d
593•ahlCVA•12h ago•265 comments

Show HN: Carrier-Explode: iPhone, Pixel and Galaxy carrier settings decoded

https://carrierexplode.com/
215•simplyalec•7h ago•26 comments

Typesafe AI raises $870M at $7.5B

https://typesafe.ai/blog/series-ai
273•tosh•8h ago•209 comments

Sorry, I'm in a meeting

https://iminafleeting.com/
766•splintersio•16h ago•241 comments

Anthropic AI model submits false tip on unsolved Philly murder

https://www.nbcphiladelphia.com/news/local/anthropic-ai-model-submits-false-tip-on-unsolved-phill...
42•Zambyte•3h ago•22 comments

YouTuber Says Cops Visited Him After He Built a Flock-Style Camera to Track Cops

https://gizmodo.com/youtuber-says-cops-paid-him-a-visit-after-he-built-flock-style-camera-to-trac...
415•gumby•4h ago•228 comments

Compiling Rust to readable C with Eurydice

https://lwn.net/Articles/1055211/
15•peter_d_sherman•2h ago•2 comments

Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

https://jessewaites.com/blog/post/i-pointed-ai-at-400-years-of-archives/
105•piratebroadcast•14h ago•53 comments

Can you use autoregressive diffusion to generate market data?

https://blog.janestreet.com/can-you-use-autoregressive-diffusion-to-generate-market-data/
9•jsomers•10h ago•0 comments

Rewriting Prime Agent in Rust

https://www.primeintellect.ai/blog/prime-agent-rust
20•piotrgrabowski•2h ago•5 comments

Show HN: Let your AI agents paint big arrows, boxes and text on your screen

https://github.com/franzenzenhofer/big-arrow-on-the-screen
378•franze•14h ago•165 comments

Show HN: Proton Drive for Linux

https://oss.lsantos.dev/proton-drive-linux-fs/
33•khaosdoctor•1d ago•13 comments

Has the Autonomous Trucking Revolution Arrived?

https://phenomenalworld.org/analysis/has-the-autonomous-trucking-revolution-arrived/
6•demonsthenes•1h ago•0 comments

Atari Falcon

https://atarimuseum.nl/atari-falcon/
15•debo_•2h ago•3 comments

The role of cat eye narrowing movements in cat–human communication (2020)

https://www.nature.com/articles/s41598-020-73426-0
39•bushwart•3d ago•22 comments

Taxing Entrepreneurial Wealth: Evidence from Norway, 2021–2025

https://www.nber.org/papers/w35854
16•PLenz•3h ago•0 comments

Show HN: The rarest tech books and docs you've probably never read

https://readrare.com/
76•miletus•8h ago•11 comments

'Wallace and Gromit,' 90% Alone

https://animationobsessive.substack.com/p/wallace-and-gromit-90-alone
160•vinhnx•12h ago•21 comments

M7.6 Earthquake in Panama

https://earthquake.usgs.gov/earthquakes/eventpage/us6000u18k/executive
129•gslin•7h ago•35 comments

What mathematicians should know about the Lean Theorem Prover: reliability & AI

https://terrytao.wordpress.com/2026/10/09/what-mathematicians-should-know-about-the-lean-theorem-...
35•matt_d•8h ago•4 comments

Microsoft-Decision-1, our model for fast decision-making

https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
144•lisajaloza•7h ago•53 comments

Scam American companies are using to manipulate ingredient lists

https://twitter.com/WallStreetApes/status/2108594998656807078
47•bilsbie•8h ago•57 comments

How to Fix autoconf-style Configuration Probing

https://build2.org/blog/fix-autoconf.xhtml
12•boris•1d ago•5 comments

The logarithms of rational numbers have irrationality exponent 2 [pdf]

https://jdb19937.github.io/log-irrationality-measure/logmeasure.pdf
3•jdb1729•2h ago•1 comments

Ideas aren't getting harder to find (2022)

https://www.experimental-history.com/p/ideas-arent-getting-harder-to-find
123•rafaelc•7h ago•54 comments

Programming Isn't Special

https://blog.glyph.im/2026/10/programming-isnt-special.html
173•ingve•18h ago•193 comments