frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Show HN: RULER – Easily apply RL to any agent

https://openpipe.ai/blog/ruler
34•kcorbitt•4h ago
Hey HN, Kyle here, one of the co-founders of OpenPipe.

Reinforcement learning is one of the best techniques for making agents more reliable, and has been widely adopted by frontier labs. However, adoption in the outside community has been slow because it's so hard to implement.

One of the biggest challenges when adapting RL to a new task is the need for a task-specific "reward function" (way of measuring success). This is often difficult to define, and requires either high-quality labeled data and/or significant domain expertise to generate.

RULER is a drop-in reward function that works across different tasks without any of that complexity.

It works by showing N trajectories to an LLM judge and asking it to rank them relative to each other. This sidesteps the calibration issues that plague most LLM-as-judge approaches. Combined with GRPO (which only cares about relative scores within groups), it just works (surprisingly well!).

We have a full writeup on the blog, including results on 4 production tasks. On all 4 tasks, small Qwen 2.5 models trained with RULER+GRPO beat the best prompted frontier model, despite being significantly smaller and cheaper to run. Surprisingly, they even beat models trained with hand-crafted reward functions on 3/4 tasks! https://openpipe.ai/blog/ruler

Repo: https://github.com/OpenPipe/ART

Comments

someoneontenet•2h ago
Love these write ups!
kcorbitt•1h ago
Thank! If there are any topics that you'd find particularly interesting, let me know and I can try to find time. :)
sadiq•1h ago
Excellent, look forward to giving this a go.

I was looking at: https://arxiv.org/abs/2506.18254 but your approach is even more general.

spmurrayzzz•45m ago
Might end up being some confusion with the RULER benchmark from NVIDIA given the (somewhat shared) domain: https://github.com/NVIDIA/RULER

EDIT: by shared I only mean the adjacency to LLMs/AI/ML, RL is a pretty big differentiator though and project looks great

ndgold•30m ago
Dope

OpenAI's Windsurf deal is off, and its CEO is going to Google

https://www.theverge.com/openai/705999/google-windsurf-ceo-openai
144•rcchen•1h ago•71 comments

ETH Zurich and EPFL to release a LLM developed on public infrastructure

https://ethz.ch/en/news-and-events/eth-news/news/2025/07/a-language-model-built-for-the-public-good.html
247•andy99•3h ago•35 comments

jank is C++

https://jank-lang.org/blog/2025-07-11-jank-is-cpp/
151•Jeaye•5h ago•47 comments

Upgrading an M4 Pro Mac mini's storage for half the price

https://www.jeffgeerling.com/blog/2025/upgrading-m4-pro-mac-minis-storage-half-price
269•speckx•8h ago•168 comments

Allen G. Hassenfeld, former CEO of Hasbro, dies at 76

https://abcnews.go.com/Business/wireStory/allen-hassenfeld-former-ceo-hasbro-family-founded-iconic-123624319
14•Bluestein•2d ago•1 comments

Andrew Ng: Building Faster with AI [video]

https://www.youtube.com/watch?v=RNJCfif1dPY
125•sandslash•1d ago•33 comments

A software conference that advocates for quality

https://bettersoftwareconference.com/
5•leoncaet•44m ago•3 comments

Astronomers race to study interstellar interloper

https://www.science.org/content/article/astronomers-race-study-interstellar-interloper
84•bikenaga•6h ago•48 comments

Activeloop (YC S18) Is Hiring AI Search and Python Back End Engineers(Onsite,MV)

https://careers.activeloop.ai/
1•davidbuniat•1h ago

Bill Atkinson's psychedelic user interface

https://patternproject.substack.com/p/from-the-mac-to-the-mystical-bill
339•cainxinth•11h ago•183 comments

Monorail – Turn CSS animations into interactive SVG graphs

https://muffinman.io/monorail/
24•stanko•3d ago•2 comments

Show HN: RULER – Easily apply RL to any agent

https://openpipe.ai/blog/ruler
34•kcorbitt•4h ago•5 comments

Lead pigment in turmeric is the culprit in a global poisoning mystery (2024)

https://www.npr.org/sections/goats-and-soda/2024/09/23/nx-s1-5011028/detectives-mystery-lead-poisoning-new-york-bangladesh
264•perihelions•7h ago•135 comments

Repaste Your MacBook

https://christianselig.com/2025/07/repaste-macbook/
141•speckx•9h ago•86 comments

Pa. House passes 'click-to-cancel' subscription bills

https://www.pennlive.com/news/2025/07/pa-house-passes-click-to-cancel-subscription-bills-as-court-throws-out-federal-rule.html
183•bikenaga•6h ago•62 comments

At Least 13 People Died by Suicide Amid U.K. Post Office Scandal, Report Says

https://www.nytimes.com/2025/07/10/world/europe/uk-post-office-scandal-report.html
511•xbryanx•10h ago•437 comments

Preliminary report into Air India crash released

https://www.bbc.co.uk/news/live/cx20p2x9093t
27•cjr•2h ago•40 comments

I'm more proud of these 128 kilobytes than anything I've built since

https://medium.com/@mikehall314/im-more-proud-of-these-128-kilobytes-than-anything-i-ve-built-since-53706cfbdc18
84•mikehall314•2h ago•22 comments

In a First, Solar Was Europe's Biggest Source of Power Last Month

https://e360.yale.edu/digest/solar-biggest-power-source-europe-june-2025
166•Brajeshwar•6h ago•101 comments

Show HN: Pangolin – Open source alternative to Cloudflare Tunnels

https://github.com/fosrl/pangolin
438•miloschwartz•1d ago•98 comments

LLM Inference Handbook

https://bentoml.com/llm/
279•djhu9•19h ago•15 comments

Introduction to Digital Filters

https://ccrma.stanford.edu/~jos/filters/
5•ofalkaed•3h ago•0 comments

The ChompSaw: A benchtop power tool that's safe for kids to use

https://www.core77.com/posts/137602/The-ChompSaw-A-Benchtop-Power-Tool-Thats-Safe-for-Kids-to-Use
273•surprisetalk•4d ago•188 comments

OpenFront: Realtime Risk-like multiplayer game in the browser

https://openfront.io/
177•thombles•16h ago•44 comments

Google nerfs Pixel 6a batteries following fire hazard

https://arstechnica.com/gadgets/2025/07/a-mess-of-its-own-making-google-nerfs-second-pixel-phone-battery-this-year/
34•fffrantz•3h ago•34 comments

Overtourism in Japan, and how it hurts small businesses

https://craigmod.com/ridgeline/210/
179•speckx•9h ago•339 comments

Show HN: Vibe Kanban – Kanban board to manage your AI coding agents

https://github.com/BloopAI/vibe-kanban
140•louiskw•7h ago•91 comments

The day someone created 184 billion Bitcoin (2020)

https://decrypt.co/39750/184-billion-bitcoin-anonymous-creator
78•lawrenceyan•17h ago•87 comments

'123456' password exposed chats for 64M McDonald's job applicants

https://www.bleepingcomputer.com/news/security/123456-password-exposed-chats-for-64-million-mcdonalds-job-applicants/
6•nan60•46m ago•0 comments

Postgres LISTEN/NOTIFY does not scale

https://www.recall.ai/blog/postgres-listen-notify-does-not-scale
546•davidgu•4d ago•282 comments