frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Claude Opus 5

https://www.anthropic.com/news/claude-opus-5
1170•alvis•6h ago•638 comments

Postgres LISTEN/NOTIFY actually scales

https://www.dbos.dev/blog/postgres-listen-notify-scalability
150•KraftyOne•4h ago•29 comments

My security camera shipped a GitHub admin token in its login page

https://hhh.hn/hanwha-github-token/
480•hhh•11h ago•167 comments

Show HN: I simulated closing the Strait of Hormuz on real oil trade data

https://globaloilnetwork.staffinganalytics.io/
38•eliotho•1d ago•16 comments

India's first privately-developed rocket reaches orbit on debut launch

https://arstechnica.com/space/2026/07/indias-first-privately-developed-rocket-reaches-orbit-on-dr...
449•sohkamyung•4d ago•132 comments

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

https://artificialanalysis.ai/models
35•aarondong•3h ago•19 comments

Designing an Ethernet Switch ASIC

https://essenceia.github.io/projects/ethernet_switch_asic/
56•random__duck•4d ago•15 comments

Nvidia, Microsoft, Meta warn against overregulating open-weight models

https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html
444•louiereederson•9h ago•207 comments

Fil-C: Garbage In, Memory Safety Out [video]

https://www.youtube.com/watch?v=5F-2Y1LPRek
84•Bootvis•1d ago•79 comments

An old patent inspired the new "Y-zipper", a three-sided fastener

https://news.mit.edu/2026/three-sided-y-zipper-design-0504
90•crescit_eundo•2d ago•25 comments

Gsxui – Shadcn-style components for Go

https://ui.gsxhq.dev/
50•jackielii•5h ago•8 comments

The Secret Origins of Amazon's Alexa

https://www.wired.com/story/how-amazon-made-alexa-smarter/
25•Anon84•3d ago•4 comments

Marimo now runs in PyCharm

https://marimo.io/blog/pycharm
58•cantdutchthis•2d ago•11 comments

If coding has been solved, why does software keep getting worse?

https://ptrchm.com/posts/nothing-works-and-everyone-is-euphoric/
410•pchm•13h ago•339 comments

Half-Life 2 running natively on HaikuOS

https://discuss.haiku-os.org/t/haiku-nvidia-porting-nvidia-driver-for-turing-gpus/16520?page=18
244•m0do1•10h ago•43 comments

Kimi K3 exploited the latest Redis server

https://twitter.com/fried_rice/status/2080059356322918777
105•Alifatisk•1d ago•31 comments

Don't Take the Black Pill [video]

https://www.youtube.com/watch?v=zLZwpH5lCD4
98•signa11•6h ago•64 comments

Firefox Containers Preview

https://blog.mozilla.org/en/firefox/firefox-containers-preview/
186•twapi•3d ago•70 comments

Show HN: Max Studio Tools – C++ DSP Modules for Max and Ableton Live

https://github.com/apresta/max-studio-tools
10•apresta•2h ago•0 comments

Unitree As2-W

https://www.unitree.com/As2-W/
83•MehrdadKhnzd•6h ago•37 comments

Be skeptical of OpenAI's rogue hacker agent story

https://www.theguardian.com/technology/2026/jul/24/openai-rogue-hacker
366•rwmj•6h ago•195 comments

Flux 3 X Mimic: The Next Generation of Video-Action Models

https://bfl.ai/blog/flux-3-mimic
305•kensai•13h ago•48 comments

The footprints of every building in NYC

https://www.beautifulpublicdata.com/the-footprints-of-every-building-in-nyc/
35•jonathanmkeegan•4d ago•4 comments

Government orders GitHub to remove Bluetooth-based chat app Bitchat: Jack Dorsey

https://www.thehindu.com/news/national/government-orders-github-to-remove-bluetooth-based-chat-ap...
338•rootkea•8h ago•254 comments

Future euro banknote design proposals

https://www.ecb.europa.eu/euro/banknotes/future_banknotes/html/all-design-proposals.en.html
105•robin_reala•13h ago•101 comments

IRGC claims it destroyed Amazon's Bahrain data center

https://houseofsaud.com/irgc-claims-destroyed-amazon-bahrain-data-center/
219•thisislife2•13h ago•274 comments

Programming language file extensions that match ISO 3166-1 alpha-2 country codes

https://www.bruh.ltd/blog/programming-language-file-extensions-that-match-an-iso-3166-1-alpha-2-c...
30•speckx•10h ago•17 comments

How to Write a Quine

https://czterycztery.pl/slowo/quine-EN.html
4•mci•1d ago•0 comments

The case for MUDs in modern times (2018)

https://www.andrewzigler.com/feed/the-case-for-muds-in-modern-times
64•bw86•11h ago•64 comments

I got into YC Startup School by hacking it

https://obaid.wtf/jotbook/2026/07/18/how-i-got-into-yc-by-hacking-it.html
92•speckx•5h ago•57 comments
Open in hackernews

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

https://artificialanalysis.ai/models
35•aarondong•3h ago

Comments

aarondong•3h ago
Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
midnightbobarun•3h ago
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
Schiendelman•1h ago
This must be on API costs, not counting the $100/200 tiers, right?
giancarlostoro•48m ago
Probably because they made ASICs to run inference for less.
brookst•46m ago
Are those actually deployed at scale yet?
brcmthrowaway•37m ago
Yes.
wmf•23m ago
I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.
nijave•34m ago
I think on swebench verified luna was only like 3% points lower for 1/5 the cost

Like 96% vs 93% or something

impulser_•22m ago
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
scrlk•14m ago
Plus GPT-5.6 is more token efficient across the board vs the Anthropic equivalents: https://artificialanalysis.ai/models?intelligence-index-toke...

No wonder why Tibo can afford to hit the reset button liberally.

charcircuit•5m ago
I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).
eli•22m ago
Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.
claude-ai•2h ago
On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb).

Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").

reilly3000•20m ago
Are you using Claude Code/CoWork or an API client? I’m curious if it has different training that makes it more effective with specific instructions/ tool calling methods that are only implemented in official harnesses.
pixelesque•10m ago
I'm curious about this too, and it's difficult to get any information about this given everyone has different setups, workflows and use-cases.

I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).

It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.

It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...

firasd•15m ago
Very interesting that one of the components is "AA-Omniscience Index"

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)

I've thought for a while that Gemini 3.x has 'big model smell'

sggyamg•7m ago
It's new, normal.
andy99•4m ago
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
chmod775•4m ago
The more interesting finding here is that it's still the second most expensive model (after Fable 5) by a wide margin, with plenty of models essentially matching its score (~1-2% diff) for half the cost.