frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Auto-research with codex: How I achieved a 232x Faster Kernel

https://sankalp.bearblog.dev/autoresearch/
63•tosh•2h ago

Comments

Almondsetat•54m ago
In the last couple of days I wanted to try out the new definitive DeepSeek v4 releases. I gave it the repository of a semi-abandoned video compression codec and I told it to perform the usual benchmark -> profile -> verify -> research -> improve loop. I specifically chose this codec because the authors include a verifier for the bitstream to make sure you don't break stuff if you want to try your own implementation. I gave the agents access to the compiler's profiler and also Intel's VTune, which has fantastic output. In a couple of hours the LLM generated SSE and AVX implementations of the compression and decompression algorithms that almost doubled performance with a single core. Then I asked it to create a CUDA implementation using NVIDIA's NSIGHT profiler as a guide and it also started doing some good work.

Personally, I believe that LLMs should be treated like an advanced version of Prolog or linear programming: you give the constraints, you have a way of verifying correctness, and you give it a clear goal. If the LLM can verify itself and course-correct you can basically leave it on autopilot

rrhjm53270•41m ago
I tried kernel autoreasearch using DeepSeek-V4-Flash as well. It spent about 1-2 hours to complete the FlashAttention optimization job (https://github.com/fengwang/FA5090/tree/main/v7) and cost me only $0.2. I believe we are ready to offload a lot of this kind well-defined constrained optimization problems to AI Agent autoresearch.
_zoltan_•40m ago
This is exactly how I use it. I mean not on abandoned repos, but in a benchmark - profile - verify - research - improve loop.
eterm•19m ago
I did something similar recently with Google's C# protobuf library. I had spotted I was getting CPU bound rather than memory bandwidth bound when doing streaming of uint32 buffers in dotnet gRPC.

I then asked claude to compare the C#/.NET implementation in the library with the C++ version, and it quickly identified that the C# library was missing a couple of fairly cheap optimisations that were present in the C++ version.

If I can help get a PR merged, then it'll be by far the biggest impact of any work I've ever done.

I also compared the Rust version, it had this specific optimisation. The far more popular Tokio/Prost library did not.

Given appropriate guardrails, LLMs are impossibly fast at iterating to find root causes and specific performance bottlenecks.

Jackobrien•45m ago
Damn! If a solo engineer can do this, it makes the most around OAI/Anthropic start to look pretty weak.
dzbarsky•13m ago
This was nowhere near the top submission. But even if a solo engineer could get a top kernel, you don't think that having thousands of engineers, infinite tokens, and stronger models than are available to the public would give the labs a significant edge?
tosh•44m ago
Training material seems to be especially rich re GPU kernels and SIMD.

I wonder if there is extra effort put into this because they are useful for the researchers working on the models or just a sub-domain that language models are a great fit for and humans have trouble with?

spacemanspiff01•41m ago
This is really cool - I really like the beam search idea,
oinoom•25m ago
this is the first time ive heard of beam search. i would have reached for a genetic algorithm of some sort, although it seems like some stochastic versions of beam search exist to avoid local minima. i wonder if there are any good frameworks for building these that agents can construct and use.
shken•38m ago
Every step here has an oracle: wall-clock, the profile, pass or fail from the verifier. I had an agent-built app audited task by task, 10 came back done and 7 worked, and the three misses were the ones needing a credential or a setting on someone else's dashboard. Nothing in the loop could tell the agent it had failed, so it said done and moved on.
ramon156•37m ago
some of these submissions seem to be omitting the actual rules. the #1 on edinh has a line that says "bypass ban check"
spwa4•10m ago
People are always going to hate auto-research and "loop engineering". Because it's got 2 properties:

1) it's the only way to get something out of models (or people for that matter) that they don't know yet.

2) it's harder to do with an LLM than without. Not easier.

3) and when you fuck it up, half the time the LLM (or other ML technique) makes a fool out of you and you spent $1000 to find the quickest way to get a robot leg on the ground is just to crash it into the ground.

Auto-research with codex: How I achieved a 232x Faster Kernel

https://sankalp.bearblog.dev/autoresearch/
65•tosh•2h ago•13 comments

The other Sean Byrne doesn't exist

https://conic.al/writing/the-other-sean-byrne-doesnt-exist/
249•rdl•8h ago•123 comments

Qwen 3.8 27B

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
1212•erdaltoprak•22h ago•722 comments

The mathematical beauty of hyperbezier curves

https://linebender.org/blog/hyperbezier/
49•raphlinus•5d ago•2 comments

The Color of White Light

https://ludens.cl/photo/spectra/spectra.html
17•xk3•3d ago•9 comments

Show HN: Eigendrum - Draw any shape and hear what it sounds like as a drum

https://baselashraf81.github.io/eigendrum/
45•BaselAshraf81•4d ago•10 comments

Using GCC's Nested Functions with Wide Pointers and No Trampolines II

https://uecker.codeberg.page/2026-07-14.html
38•uecker•5h ago•16 comments

Going Dark, and the era of law enforcement hacking

https://blog.cryptographyengineering.com/2026/08/14/everything-is-about-to-go-dark/
373•vslira•16h ago•174 comments

In 1962, Egypt's Missile Program Lost Its Key Scientist Without a Trace

https://www.popularmechanics.com/military/a73358518/nazi-rocket-scientist-disappearance/
67•bookofjoe•3d ago•37 comments

Understanding WCAG 2.2 as ePub and PDF

https://doeken.org/wcag-ebook
36•doekenorg•3d ago•2 comments

Google is making private AI practical with homomorphic encryption

https://blog.google/security/how-google-is-making-private-ai-practical-with-homomorphic-encryption/
428•u1hcw9nx•21h ago•257 comments

Show HN: ThoughtDAG – An editable context graph for LLM conversations

https://chenxiachan.github.io/thoughtdag/
59•chatchan•8h ago•15 comments

Firefox is now the last major browser that still supports uBlock Origin

https://www.pcworld.com/article/3212428/firefox-is-now-the-last-major-browser-that-still-supports...
1248•DemiGuru•17h ago•473 comments

RustDesk now supports true unattended remote access on Wayland

https://rustdesk.com/blog/unattended-remote-access-wayland/
319•rustdesk•20h ago•129 comments

Magnitude 7.7 Earthquake – 68 km NNW of Ende, Indonesia

https://earthquake.usgs.gov/earthquakes/eventpage/us6000tkt2/executive
207•Bender•11h ago•51 comments

eigendrum

https://eigendrum.com/#p=circle
189•bookofjoe•14h ago•54 comments

Coin-sized device can hack a Boeing 737

https://www.wired.com/story/this-coin-sized-device-can-hack-a-boeing-737/
104•_tk_•2d ago•72 comments

AI by Hand

https://www.byhand.ai/
331•sans_souse•21h ago•24 comments

Geometric Reasoning

https://sophontic.ai/
3•6510•3d ago•1 comments

Maximizing the value of your Claude Code sessions

https://claude.com/blog/maximizing-the-value-of-your-claude-code-sessions
260•twapi•20h ago•140 comments

Simplifying and Refactoring Introductory Calculus (2018)

https://arxiv.org/abs/1811.03459
107•E-Reverance•12h ago•54 comments

Show HN: Silent Shark – tactical map-based WWII submarine sim

https://silentshark.app/
51•epaga•2d ago•17 comments

Unearthing a 31 year old Easter egg in Ecco the Dolphin

https://32bits.substack.com/p/under-the-microscope-ecco-the-dolphin-98c
105•bbayles•2d ago•23 comments

Debian has begun voting on the future of AI/LLM contributions

https://lists.debian.org/debian-devel-announce/2026/08/msg00002.html
43•matheusmoreira•3h ago•27 comments

This Hi-Fi Tape Recorder Changed Radio Forever

https://spectrum.ieee.org/magnetophon-laugh-track
54•Jimmc414•3d ago•16 comments

Super Mario Derivations

https://fzakaria.com/2026/08/05/super-mario-derivations
119•domenkozar•1w ago•16 comments

Racket v9.3

https://blog.racket-lang.org/2026/08/racket-v9-3.html
112•privong•18h ago•7 comments

I turned my RSS feeds into an e-ink newspaper to stop reading on my phone

https://heyjonny.dev/posts/rss-to-eink-newspaper/
204•speckx•22h ago•87 comments

GLM-5.3: Frontier coding with emergent cyber capabilities

https://z.ai/blog/glm-5.3
1108•pella•1d ago•543 comments

Introducing Toast 1

https://www.mixedbread.com/blog/toast-1
210•mplappert•21h ago•64 comments