frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Benchmarking LLM query generation across SQL, Cypher, and TypeQL

https://typedb.com/blog/benchmarking-llm-query-generation-across-sql-cypher-and-typeql
1•flyingsilverfin•59m ago

Comments

flyingsilverfin•48m ago
Hi all, CTO of TypeDB here. The question I actually wanted to answer was "does a type system help agents with correctness?" Everyone's intuition says yes, especially as complexity grows, but evidence is thin and contested, and I haven't found much that isolates the effect for LLMs beyond constrained-decoding work.

A database is a smaller surface than a programming language, so we turned it into a benchmark that holds the model constant and swaps the language: SQL, Cypher, TypeQL, same questions, same data. It's not a clean apples-to-apples comparison and the post's limitations section outlines that further.

One result that points us in the direction of a strong type system is that SQL wins first-shot, but 93% of TypeQL's wrong queries fail with an error, versus 5% for SQL and 40% for Cypher. Add in a retry loop TypeQL ends up ahead, because it can only fix mistakes it can see.

Of course there's a lot of variables here, such model, skill, token usage, question selection, etc. I actually think it's quite a hard thing to study!

If anyone knows of work on LLMs driving systems with stronger vs weaker verifiers, with the model held fixed, I'd love to see it.

Accident: JSX E145 Houston 2024, gear collapse and runway excursion on landing

1•r2sk5t•53s ago•0 comments

Meta's New AI Agent Is an Instant Hit–and the Backlash Has Begun

https://www.wsj.com/tech/ai/meta-ai-agent-muse-reactions-5bf236af
1•thm•1m ago•0 comments

Parsing JSON Objects without intermediate ASTs

https://arthi-chaud.github.io/posts/json-ir/
1•ibobev•1m ago•0 comments

Show HN: Fractalysis – Online interactive infinite-zoom Mandelbrot viewer

https://www.erwindegroot.nl/fractalysis/
1•erwindegroot•1m ago•0 comments

Do not let your type system reason about aliasing in your programming language

https://futhark-lang.org/blog/2026-09-22-aliasing.html
2•ibobev•2m ago•0 comments

The Zig Journey

https://kristoff.it/blog/the-zig-journey/
1•ibobev•3m ago•0 comments

Trump's 1,156 July Stock Trades Involved AI, Big Oil, Weapons-Makers, and More

https://www.commondreams.org/news/donald-trump-stock-trades
3•cirelli94•3m ago•0 comments

LinkedIn is a terrible quality website

1•skyfantom•3m ago•1 comments

Slashing Our AWS Bill at Levels.fyi, Part 2

https://www.levels.fyi/blog/slashing-our-aws-bill-at-levelsfyi-part-2.html
1•ghostfoxgod•6m ago•0 comments

Hacker News Sentiment Analysis Using Laya (System One)

https://github.com/skhaz/hackernews-sentiment-analysis
1•mococa•6m ago•0 comments

Nikclas – Compare the prices of different AI models

https://github.com/NikclasTech/nikclas
1•Cr12dev•6m ago•1 comments

Help with Advise

1•saventiiy•8m ago•0 comments

Agentic Coding for Builders Who Ship

https://github.com/Kuberwastaken/claurst
1•Bluestein•8m ago•0 comments

Show HN: Ax – Let Claude, Codex and OpenCode talk to each other locally

https://useax.dev/
1•rcdexta•9m ago•0 comments

Laya-mlx and 12 more trending open-source AI repos · week 39, 2026

1•thezakulo•10m ago•0 comments

Building a Custom Harness with Jev and Pi

https://academy.dair.ai/resources/jev-decisions-in-a-pi-sdk-harness
1•omarsar•11m ago•0 comments

Cambridge Coincidences Collection

http://understandinguncertainty.org/coincidences/
1•momentmaker•11m ago•0 comments

Show HN: HackDigest – AI-summarized daily news digest via email

https://minjungsung.github.io/hackdigest/
1•minjungsung1994•12m ago•0 comments

Show HN: Karada.ai – CI/CD to turn APIs into MCP servers with 1-click plugins

https://karada.ai
3•shashtag•12m ago•0 comments

Montreal adopts bylaw banning insults against police, municipal employees

https://www.cbc.ca/news/canada/montreal/montreal-city-council-police-9.7352920
4•MC995•14m ago•0 comments

You Don't Need a %Frontier LLM%

https://rakshazi.me/blog/you-dont-need-frontier-llm
1•aine•14m ago•0 comments

PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from LLMs

https://arxiv.org/abs/2109.05093
1•Bluestein•14m ago•0 comments

Volt, A compiled reactive web language targeting WASM-GC (No VDOM)

https://github.com/bshea-1/volt/
2•bshea1•15m ago•0 comments

Streaming 500M rows into Apache Arrow in 2.3 seconds

https://questdb.com/blog/streaming-500-million-rows-into-apache-arrow/
2•jinqueeny•15m ago•0 comments

Is China's Power Advantage About to Trigger an 89% Crash in U.S. AI Stock

https://oilprice.com/Energy/Energy-General/Is-Chinas-Secret-Power-Advantage-About-To-Trigger-An-8...
1•thelastgallon•17m ago•0 comments

Cleaning up after Matt Parker

https://leancrew.com/all-this/2026/09/cleaning-up-after-matt-parker/
1•surprisetalk•17m ago•0 comments

Show HN: Open benchmark for STT when a second person is talking (265 recordings)

https://krisp.ai/blog/voice-isolation-benchmark/
1•davitb•17m ago•0 comments

Synchronous Control Monitoring: Preventing Harmful Agent Actions in Real Time

https://max.ax/writing/synchronous-control-monitoring/
1•k5hp•17m ago•0 comments

How to Use T-SNE Effectively

https://distill.pub/2016/misread-tsne/
1•cagrie•19m ago•0 comments

Ask HN: What job boards are good these days?

2•phendrenad2•20m ago•1 comments