frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Gemini Robotics 2 brings whole body intelligence to robots

https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
40•ai2027•24m ago•14 comments

Why DNA damage from smoking and UV rays cause cancer in some but not others

https://www.cam.ac.uk/research/news/study-reveals-why-dna-damage-from-smoking-and-uv-rays-may-cau...
35•gmays•1h ago•38 comments

SDL_GPU minimal, single-header, high-performance 2D graphics painting library

https://github.com/n67094/sdl_gp
21•n67094•1h ago•4 comments

The Lost Civic Life of Movie Rental Stores

https://thereader.mitpress.mit.edu/the-lost-civic-life-of-movie-rental-stores/
31•facundo_olano•1h ago•19 comments

Launch HN: Prized (YC S26) – Let non-engineer staff build secure internal tools

https://prized.dev
33•marinoseliades•2h ago•17 comments

Paging Through a Parquet File in DuckDB: File_row_number or Offset?

https://rusty.today/blog/paging-parquet-duckdb-file-row-number-vs-offset/
10•rustyconover•41m ago•0 comments

RFC 8890 – The Internet is for End Users (2020)

https://mnot.net/blog/2020/for_the_users
50•notarobot123•2h ago•12 comments

Hacker Public Radio

https://hackerpublicradio.org/
31•bmacho•1h ago•1 comments

Git worktrees are not an isolation boundary for coding agents

https://fletch.sh/blog/git-worktrees-vs-clones-for-ai-agents/
16•alchaplinsky•1h ago•12 comments

Ron Gilbert started production on Thimbleweed Park 2

https://www.grumpygamer.com/twp2_announce/
157•alberto-m•7h ago•63 comments

Why Don't People Use Formal Methods? (2019)

https://www.hillelwayne.com/post/why-dont-people-use-formal-methods/
74•Thom2503•3h ago•47 comments

Trusted URLs via Cryptographic Signatures

https://blog.certisfy.com/2026/04/trusted-urls-via-cryptographic.html
8•Edmond•1h ago•5 comments

Mbodi AI (YC P25) Is Hiring Robotics/Research Engineers

https://www.ycombinator.com/companies/mbodi-ai/jobs
1•chitianhao•3h ago

Gpiozero Flow

https://bennuttall.com/blog/2026/07/gpiozero-flow/
101•benn_88•5h ago•30 comments

Building a native C# implementation of CEL engine

https://bsid.io/writing/building-a-cel-engine-for-net
13•jackedEngineer•2d ago•0 comments

Why Is Everyone Trying to Build a Solid-State Battery?

https://www.construction-physics.com/p/why-is-everyone-trying-to-build-a
74•crescit_eundo•3h ago•77 comments

Are We Stuck with Lean?

https://mathoverflow.net/questions/513742/are-we-stuck-with-lean
44•jjgreen•3h ago•14 comments

Upper stage impacting the moon on 2026 August 5

https://www.projectpluto.com/25010d.htm
18•ryannevius•2h ago•1 comments

The Alice and Bob After Dinner Speech (1984)

https://hex.ooo/library/alicebob.html
16•kamma4434•3d ago•1 comments

The Economic Benefit of Refactoring

https://martinfowler.com/articles/exploring-gen-ai/refactoring-economic-benefit.html
5•javaeeeee•29m ago•0 comments

Go LLM SDK for streaming, tool-calling AI backends (plus frontend React lib)

https://github.com/grafana/ai-sdk
33•matryer•3h ago•12 comments

3D Pinball for Windows (1995)

https://98.js.org/programs/pinball/space-cadet.html
40•mushstory•4h ago•18 comments

The Glass Famine

https://edconway.substack.com/p/the-glass-famine
69•baud147258•3d ago•31 comments

You can't solve computer use by ignoring the interface

https://steelmanlabs.com/blog/computer-use-is-far-from-solved
37•mpavlov•4h ago•11 comments

CosmosEscape: Taking over Every Database in Azure Cosmos DB

https://www.wiz.io/blog/cosmosescape-taking-over-every-database-in-azure-cosmos-db
28•uvuv•3h ago•7 comments

Azulejo

https://en.wikipedia.org/wiki/Azulejo
117•Amorymeltzer•1d ago•36 comments

Google will expand age checks on Android worldwide till the end of the year

https://android-developers.googleblog.com/2026/07/google-play-age-signals-api-safer-experiences.html
241•dmantis•5h ago•271 comments

The first watch featuring computer functions

https://by.seiko-design.com/140th/en/topic/58.html
63•stefanv•4d ago•25 comments

Agent-Manager: A Tmux TUI for Running Claude Code, Codex and OpenCode

https://github.com/YoanWai/agent-manager
73•yoanwaidev•6h ago•56 comments

The Productivity Mirage

https://frantic.im/mirage/
328•msephton•16h ago•142 comments
Open in hackernews

Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points

https://agiranker.com/
4•baraklaniado•1h ago

Comments

baraklaniado•1h ago
I run AGI Ranker, an AI benchmark aggregator that calculates an AGI score for every benchmarked frontier model, and have just released v2 earlier today. Auditing my own scale has led to a lower AGI score, 6-15pts less, yet the models kept their shape. Only one of the ten benchmarks has actually measured a human ceiling, GPQA-Diamond 81%, so the other 9 are anchored to benchmark-max, not to human-parity. Coverage has greatly improved, from ~43% to ~80%, and that means the score leans less on shrinkage and more on measurement. Single evaluator dependency has been successfully dropped from ~45% to ~26%, with the new offsets published on the site. Agency was rebuilt on 4 benchmarks from 3 independent evaluators. Speciality tabs now include Coding and Knowledge, with Reasoning temporarily dropped because it was leaning on only one benchmark, but will make a comeback once AIME 2026 is published. The Value tab shows cost-per-capability. The Corrections log makes sure that all my mistakes are openly reported on the site, including a recent apology to Deepseek, for having published an unsourced ARC-AGI-2 number. No lab money, no paid placements, every cell traceable, open data (CC-BY). It's a solo project that I maintained while having only 450 monthly visitors, determined to offer those who do visit, the most accurate and useful information about AI model prowess. Known limitations are listed on the site - happy to answer questions about the methodology or the functionality of the site. Cheers!