frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

fp.

Open in hackernews

Show HN: Legal RAG Bench

https://isaacus.com/blog/legal-rag-bench
2•beowa•1h ago
Hey HN, This is Legal RAG Bench, the first benchmark for legal RAG systems to simultaneously evaluate hallucinations, retrieval failures, and reasoning errors.

The key takeaways of our benchmark are: 1. Embedding models, not generative models, are the primary driver of RAG accuracy. Switching from a general-purpose embedder like OpenAI's Text Embedding 3 Large to a legal domain embedder like Kanon 2 Embedder can raise accuracy by ~19 points. 2. Hallucinations are often triggered by retrieval failures. Fix your retrieval stack, and, in most cases, you end up fixing hallucinations. 3. Once you have a solid legal retrieval engine, it doesn’t matter as much what generative model you use; GPT-5.2 and Gemini 3.1 Pro perform relatively similarly, with Gemini 3.1 Pro achieving slightly better accuracy at the cost of more hallucinations. 4. Google's latest LLM, Gemini 3.1 Pro, is actually a bit worse than its predecessor at legal RAG, achieving 79.3% accuracy instead of 80.3%.

These findings confirm what we already suspected, that information retrieval sets the ceiling on the accuracy of legal RAG systems. It doesn’t matter how smart you are, you aren’t going to magically know what the penalty is for speeding in California without access to an up-to-date copy of the California Vehicle Code.

Even still, to our knowledge, we’re the first to actually show this empirically.

Unfortunately, as we highlight in our write-up, high-quality open legal benchmarks like Legal RAG Bench and our earlier Massive Legal Embedding Benchmark (MLEB) are few and far between.

We point out, for example, that the popular Vals AI CaseLaw (v2) benchmark yields LLM rankings inexplicably and starkly different from ours while also failing to properly evaluate end-to-end RAG performance. Because CaseLaw (v2) is a private and proprietary benchmark, we are unable to confirm the source of the discrepancies we discovered, though we suspect they lie in a seriously flawed evaluation and labeling methodology.

In the interests of transparency, we have not only detailed exactly how we built Legal RAG Bench, but we’ve also released all of our data openly on Hugging Face here: https://huggingface.co/datasets/isaacus/legal-rag-bench/. We will also soon be publishing our write up as a paper.

My side project got banned from the internet

https://trysound.io/how-my-side-project-got-banned-from-the-internet/
1•speckx•2m ago•0 comments

Show HN: Remote-OpenCode – Control your AI coding assistant from Discord

https://github.com/RoundTable02/remote-opencode
1•remocode•4m ago•1 comments

The Market for Marriage

https://worksinprogress.co/issue/marriage-customs-very-different-from-ours/
1•bensouthwood•4m ago•0 comments

Building an Elite AI Engineering Culture in 2026

https://www.cjroth.com/blog/2026-02-18-building-an-elite-engineering-culture
1•cmsefton•5m ago•0 comments

Show HN: A Self-Paced Exercise to Build a CLI Coding Agent from Scratch

https://github.com/primaprashant/alduin
1•primaprashant•6m ago•1 comments

Show HN: AgentBouncr – Governance layer for AI agents

https://github.com/agentbouncr/agentbouncr
1•Soenke_Cramme•6m ago•1 comments

Fixing macOS window chaos: display-aware layout restore with Hammerspoon

https://ai.rundatarun.io/Practical+Applications/fixing-macos-window-chaos-hammerspoon-karabiner
1•RyeCatcher•6m ago•0 comments

Themes and plugins for Claude Code's status bar

https://github.com/npow/oh-my-claude
1•elwebmaster•8m ago•0 comments

Ask HN: Who is hiring? (February 2026)

1•ssunboyy•10m ago•0 comments

CryptPad Funding Status January 2026

https://blog.cryptpad.org/2026/02/18/CryptPad-Funding-Status-2026/
1•Timshel•10m ago•0 comments

Amazon's cloud was hit by two outages involving AI tools in December, FT says

https://www.marketscreener.com/news/amazon-s-cloud-was-hit-by-two-outages-involving-ai-tools-in-d...
1•_____k•12m ago•0 comments

Reassessing European Contact: Insights from Spanish America [pdf]

https://isonomiaquarterly.com/wp-content/uploads/2026/02/iq-4.1-spring-2026-bassi-conquest-and-th...
1•brandonlc•13m ago•0 comments

Epoll and Kqueue: How Operating Systems Learned to Wait Efficiently

https://thecodinggopher.substack.com/p/epoll-and-kqueue-how-operating-systems
2•syntacticbs•15m ago•1 comments

The U.S. and China Are Pursuing Different AI Futures

https://spectrum.ieee.org/us-china-ai
2•oldnetguy•16m ago•0 comments

Nvidia DGX Spark: Is DGX Spark Blackwell?

https://www.backend.ai/blog/2026-02-is-dgx-spark-actually-a-blackwell
1•ResearchAtPlay•17m ago•1 comments

Pg-here: Run a local PostgreSQL instance in your project folder with one command

https://github.com/mayfer/pg-here
1•birdculture•17m ago•0 comments

Private Equity Debt Left a Leading VPN Open to Chinese Hackers

https://www.bloomberg.com/news/features/2026-02-19/vpn-used-by-us-government-failed-to-stop-china...
1•pvachon•23m ago•0 comments

Exercise has 'similar effect' to therapy, study on depression shows

https://medicalxpress.com/news/2026-01-similar-effect-therapy-depression.html
4•PaulHoule•25m ago•0 comments

Show HN: Lobster Sauce – the most recent OpenClaw news in one central place

https://www.lobstersauce.news/
1•Tjerkienator•26m ago•0 comments

NASA recalls dysfunction, emotions during Boeing's botched Starliner flight

https://www.reuters.com/business/aerospace-defense/nasa-chief-slams-boeing-agency-failures-botche...
1•JumpCrisscross•28m ago•0 comments

IEEE 802.19.3:Coexistence Recommendations for sub 1-GHz IEEE 802.11 and 802.15.4

https://ieeexplore.ieee.org/document/10148905
1•teleforce•30m ago•0 comments

Fault tolerant message passing C# with NATS.io: Use Distributed Object Store

https://nats-io.github.io/nats.net/documentation/object-store/intro.html
1•northlondoner•31m ago•1 comments

Show HN: Google Drive CLI for LLMs / Coding Agents

https://github.com/NmadeleiDev/google-drive-cli
1•Gregoryy•32m ago•0 comments

Ente Locker

https://ente.io/blog/locker/
1•sylens•33m ago•0 comments

The Software Factory: When No Human Writes or Reviews the Code

https://www.thepragmaticcto.com/p/the-software-factory-when-no-human
1•allanmacgregor•35m ago•0 comments

Coding Agents Are at Stage 5, Everything Else Is Stuck at Stage 1

https://kanyilmaz.me/2026/02/19/five-stages-of-ai-agents.html
1•thellimist•35m ago•0 comments

Show HN: LinkedRecords – A Server-Sovereign Alternative to Firebase

https://linkedrecords.com/
1•WolfOliver•38m ago•0 comments

Django ORM Standalone⁽¹⁾: Querying an existing database

https://www.paulox.net/2026/02/20/django-orm-standalone-database-inspectdb-query/
1•pauloxnet•40m ago•0 comments

Gentoo Linux moves away from GitHub due to AI

https://www.pcgamer.com/software/linux/after-microsoft-couldnt-keep-its-ai-hands-to-itself-a-noto...
3•majkinetor•41m ago•0 comments

Is there a way to sort/filter by score on Hacker News?

1•7gorillaz•41m ago•1 comments