frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Show HN: RULER – Easily apply RL to any agent

https://openpipe.ai/blog/ruler
54•kcorbitt•10h ago
Hey HN, Kyle here, one of the co-founders of OpenPipe.

Reinforcement learning is one of the best techniques for making agents more reliable, and has been widely adopted by frontier labs. However, adoption in the outside community has been slow because it's so hard to implement.

One of the biggest challenges when adapting RL to a new task is the need for a task-specific "reward function" (way of measuring success). This is often difficult to define, and requires either high-quality labeled data and/or significant domain expertise to generate.

RULER is a drop-in reward function that works across different tasks without any of that complexity.

It works by showing N trajectories to an LLM judge and asking it to rank them relative to each other. This sidesteps the calibration issues that plague most LLM-as-judge approaches. Combined with GRPO (which only cares about relative scores within groups), it just works (surprisingly well!).

We have a full writeup on the blog, including results on 4 production tasks. On all 4 tasks, small Qwen 2.5 models trained with RULER+GRPO beat the best prompted frontier model, despite being significantly smaller and cheaper to run. Surprisingly, they even beat models trained with hand-crafted reward functions on 3/4 tasks! https://openpipe.ai/blog/ruler

Repo: https://github.com/OpenPipe/ART

Comments

someoneontenet•7h ago
Love these write ups!
kcorbitt•6h ago
Thank! If there are any topics that you'd find particularly interesting, let me know and I can try to find time. :)
sadiq•6h ago
Excellent, look forward to giving this a go.

I was looking at: https://arxiv.org/abs/2506.18254 but your approach is even more general.

kcorbitt•4h ago
I really like RLPR for when you have a known-good answer to compare to as well!
spmurrayzzz•6h ago
Might end up being some confusion with the RULER benchmark from NVIDIA given the (somewhat shared) domain: https://github.com/NVIDIA/RULER

EDIT: by shared I only mean the adjacency to LLMs/AI/ML, RL is a pretty big differentiator though and project looks great

kcorbitt•4h ago
Dang, hadn't seen that. Namespace collision strikes again.
ndgold•5h ago
Dope
maxrmk•4h ago
Very cool. Do you do anything to mitigate ordering bias in the evaluation function, or do you just expect it to average out over time?
kcorbitt•4h ago
No, we don't do anything. Theoretically we could judge several times with different ordering.

We could measure order bias really easily though; we just need to look at the average score by rollout position across many runs. I'll add that to my list of experiments!

Show HN: RULER – Easily apply RL to any agent

https://openpipe.ai/blog/ruler
54•kcorbitt•10h ago•9 comments

Show HN: Transition – AI Triathlon Coach

https://www.transition.fun/
3•awormuth•1h ago•0 comments

Show HN: VibeKin – Gated Discord Tribes via Personality Matching

https://tgc.fly.dev
2•madebywelch•1h ago•0 comments

Show HN: Vibe Kanban – Kanban board to manage your AI coding agents

https://github.com/BloopAI/vibe-kanban
150•louiskw•12h ago•97 comments

Show HN: Interactive pinout for the Raspberry Pi Pico 2

https://pico2.pinout.xyz
129•gadgetoid•4d ago•28 comments

Show HN: Pangolin – Open source alternative to Cloudflare Tunnels

https://github.com/fosrl/pangolin
460•miloschwartz•1d ago•109 comments

Show HN: Open source alternative to Perplexity Comet

https://www.browseros.com/
274•felarof•1d ago•112 comments

Show HN: Cactus – Ollama for Smartphones

https://github.com/cactus-compute/cactus
213•HenryNdubuaku•1d ago•80 comments

Show HN: CXXStateTree – A modern C++ library for hierarchical state machines

https://github.com/ZigRazor/CXXStateTree
47•zigrazor•4d ago•35 comments

Show HN: I built a playground to showcase what Flux Kontext is good at

https://fluxkontextlab.com
69•Zephyrion•2d ago•16 comments

Show HN: Director – Local first, open source MCP Gateway

11•bwm•12h ago•5 comments

Show HN: OffChess – Offline chess puzzles app

https://offchess.com
366•avadhesh18•3d ago•163 comments

Show HN: FlopperZiro – A DIY open-source Flipper Zero clone

https://github.com/lraton/FlopperZiro
353•iraton•2d ago•73 comments

Show HN: MCP server for searching and downloading documents from Anna's Archive

https://github.com/iosifache/annas-mcp
250•iosifache•2d ago•78 comments

Show HN: An Improvisational Web Server

https://github.com/jasonthorsness/ginprov
4•jasonthorsness•11h ago•1 comments

Show HN: Typeform was too expensive so I built my own forms

https://www.ikiform.com/
180•preetsuthar17•1d ago•94 comments

Show HN: asyncmcp – Run MCP over async transport via AWS SNS+SQS

https://github.com/bh-rat/asyncmcp
31•bharatgel•1d ago•4 comments

Show HN: Helices Create a New Model of Deterministic Computation [pdf]

https://lambdalord.github.io/Twin-Helix-Geometric-Computation/TwinHelix_rough_fin.pdf
2•bkaminsky•12h ago•1 comments

Show HN: BreakerMachines – Modern Circuit Breaker for Rails with Async Support

https://github.com/seuros/breaker_machines
44•seuros•5d ago•19 comments

Show HN: Petrichor – a free, open-source, offline music player for macOS

https://github.com/kushalpandya/Petrichor
196•kushalpandya•2d ago•105 comments

Show HN: Multiple barcodes can be generated on single page

https://ddddddo.github.io/barcode/
2•ddddddO•14h ago•0 comments

Show HN: NodeLoop – Hub for electronics design knowledge and tools

https://nodeloop.org/
5•eezZ•15h ago•0 comments

Show HN: NYC Subway Simulator and Route Designer

https://buildmytransit.nyc
197•HeavenFox•4d ago•32 comments

Show HN: TUI personal monthly budget planner

https://github.com/eliasdorneles/moomoolah
3•eliasdorneles•16h ago•0 comments

Show HN: Ten years of running every day, visualized

https://nodaysoff.run
29•friggeri•1d ago•10 comments

Show HN: A decentralized command line key-value store on Nostr

https://github.com/chr15m/nkv
2•chr15m•18h ago•0 comments

Show HN: AI Movie Finder – I created a way to find movies by describing

https://www.aimoviefinder.com
4•mosbyllc•19h ago•2 comments

Show HN: I wrote a "web OS" based on the Apple Lisa's UI, with 1-bit graphics

https://alpha.lisagui.com/
513•ayaros•5d ago•141 comments

Show HN: Code is all you need – Sherlog MCP

https://github.com/GetSherlog/Sherlog-MCP
4•randomaifreak•20h ago•0 comments

Show HN: Jukebox – Free, Open Source Group Playlist with Fair Queueing

https://www.jukeboxhq.com/
121•skeptrune•3d ago•43 comments