frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Make your first edit to OpenStreetMap

https://high5apps.github.io/josm-plugin-website-wizard/
225•juliantigler•5h ago•63 comments

Nvidia is the central bank of AI

https://www.economist.com/interactive/briefing/2026/09/03/nvidia-is-the-central-bank-of-ai
309•tolugenius•6h ago•211 comments

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

https://withspecific.com/benchmarks/real-swe
31•theanonymousone•1h ago•17 comments

Stabilizing Rust's Never Type

https://lwn.net/SubscriberLink/1091015/d9e48318ed242b41/
91•cjd8•3d ago•10 comments

Benchmark: CadQuery vs. OpenSCAD for agentic CAD work

https://modelrift.com/blog/cadquery-vs-openscad/
19•jetter•2h ago•22 comments

Apple iPod Engraver (2019)

https://dunstanorchard.com/apple-ipod-engraver/
33•NaOH•3d ago•1 comments

LG denies TV spying claims, says tracking and snooping concerns 'not true'

https://www.tomshardware.com/tech-industry/big-tech/lg-strongly-denies-tv-security-claims-says-tr...
333•datakan•2d ago•285 comments

Will There Be a 7G?

https://arxiv.org/abs/2609.01877
63•Betelbuddy•4h ago•107 comments

I made a build visualizer to understand Bun's compile times

https://lalitm.com/post/buildprof/
65•lalitmaganti•7h ago•13 comments

IKEA made a mod for Skyrim [video]

https://www.youtube.com/watch?v=iZODN0QUgjI
519•kegenaar•2d ago•134 comments

We must pace the frontier

https://darioamodei.com/post/we-must-pace-the-frontier
445•apsec112•7h ago•614 comments

Microcode in Intel's 8087 floating-point chip: the scale instruction

https://www.righto.com/2026/09/8087-microcode-reverse-engineering-fscale.html
69•pwg•6h ago•21 comments

How Trail of Bits helps verify the integrity of Signal chats

https://blog.trailofbits.com/2026/08/11/how-trail-of-bits-helps-verify-the-integrity-of-your-sign...
38•dgroshev•10h ago•14 comments

Linux Zoom client proactively reading everything written to X11 clipboard

https://hachyderm.io/@simontatham/117201594980991062
72•encyclopedism•3h ago•19 comments

A Mathematical Framework for Transformer Circuits (2021)

https://transformer-circuits.pub/2021/framework/index.html
69•Bluestein•8h ago•16 comments

LG Says We're Fake News [video]

https://www.youtube.com/watch?v=ToP9xfLDSME
43•HelloUsername•2h ago•5 comments

I fixed a tractor using John Deere's self-repair service. Farmers aren't sold

https://www.wired.com/story/i-fixed-a-tractor-john-deere-self-repair-service/
83•sbulaev•1d ago•97 comments

Retrospectively Reverse-Engineering Apple's Neural Engine

https://eiln.github.io/posts/ane.html
207•zdw•14h ago•29 comments

Android NAT-T keepalive offload bypasses VPN lockdown

https://supuk.ch/papers/android-natt-keepalive-vpn-bypass
158•mhitza•1d ago•42 comments

The Magic Behind Cubacadabra

https://andrewarrow.dev/2026/moon/2/day/19/the-magic-behind-cubacadabra/
6•andrewfromx•2d ago•1 comments

Performance of WebAssembly Runtimes in 2026

https://00f.net/2026/06/23/webassembly-runtimes-2026/
79•fagnerbrack•3d ago•18 comments

Eating Fruit Skins

https://pgadey.ca/blog/eating-fruit-skins/
59•surprisetalk•3d ago•143 comments

Navier-Stokes Announcement

https://www.claymath.org/news/navier-stokes-announcement/
296•rvz•17h ago•236 comments

LRU is harder to beat than the KV-cache papers suggest

https://github.com/gauravapiscean/agentic-kv-cache
90•gauravapiscean•2d ago•41 comments

OpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026

https://techcrunch.com/2026/09/12/openais-sam-altman-says-it-would-be-ill-advised-to-go-public-in...
12•humanlion87•1h ago•7 comments

The worst spam emails: iLands AI agent hustle

https://tedium.co/2026/09/11/ilands-agents-email-spam-kaixin-tang/
98•ColinWright•10h ago•42 comments

λ Snap – An inviting programming language for kids and adults for CS study

https://snap.berkeley.edu/
169•dr_kiszonka•1d ago•107 comments

How we manage and engage with our horses shapes their personality

https://www.utu.fi/en/news/press-release/how-we-manage-and-engage-with-our-horses-shapes-their-pe...
32•thunderbong•2d ago•12 comments

A Design Space Exploration of Async/Await

https://cel.cs.brown.edu/blog/design-space-async-await/
412•wcrichton•3d ago•119 comments

google.com/goto: Google's anti-scraping update

https://www.autom.dev/blog/google-search-goto-links
618•1e1a•18h ago•477 comments
Open in hackernews

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

https://withspecific.com/benchmarks/real-swe
31•theanonymousone•1h ago

Comments

dgellow•39m ago
A bit of a meta question: what are the most relevant benchmarks by now?
andriy_koval•23m ago
Nvidia and OpenAI claimed AGI, but you still have a job.
tetec1•16m ago
Epoch.ai has a global score and tracks many benchmarks: https://epoch.ai/benchmarks
jcmontx•37m ago
I’ve been able to offload most tasks (coding or eles) to Codex since 5.3-codex with extra high thinking
riddlemethat•11m ago
Astra lets me offload entire projects without worrying about individual tasks…
jeffybefffy519•4m ago
Do you review the outputs?
traceroute66•26m ago
So TL;DR benchmarking in a completely non-reproducible manner ?

"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".

So basically pinky-promise benchmarking ?

I'm not sure I follow the value here ?

kadoban•17m ago
If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
traceroute66•12m ago
> You're giving up transparency for it being harder to game

But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?

demibabs•13m ago
Doesn’t it ultimately have to be this way, to prevent saturation?
bix6•25m ago
Wake me up when September ends or when I can do this locally.
demibabs•12m ago
> Each task comes from a private production codebase that we licensed from a real-world company

How does that work?

traceroute66•9m ago
> How does that work?

My gut feeling is that any serious real-world company with a proprietary codebase worth looking at would not be handing out the crown jewels to a third party. License or not.

I don't doubt somebody licensed their codebase to them, I just have my doubts about who the "who" could be.

IshKebab•12m ago
I think these benchmarks are not that useful, e.g. this suggests Fable is better than Astra, but in practice Astra is waaaaaay faster (like 5x; it's not even close), and also waaaay less annoying to talk to.

There's only two or three sane options here - you can easily try them all and pick yourself.

bdlowery•4m ago
The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark.

Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change.

lmeyerov•3m ago
My intuition is that many of the better & bigger 'private' code bases, at least in terms of claude code and codex... are not in fact private at this point.

One lesson of running botsbench.com, in a slightly different domain, is to measure for model contamination every time.