frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities
22•walrus01•2h ago

Comments

avaer•28m ago
What's "Top U.S. Models"?

Even if we know what the set is, it's not clear what the numbers actually mean. Is it min, max, average, weighted, median of the models? Prerelease or public, with or without safeguards?

A bit frustrating to have this be hand waved in a report, the graphs might as well just have two mystery bars, U.S. and China.

guessmyname•14m ago
NIST named the top U.S. models in the full report [1].

Spoiler: They are referring to OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview, which, by the way, is different than Fable5/Mythos5/Opus5, just so people here don’t start speculating with opinions. If you are part of Project Glasswing you know the difference among them.

[1] https://www.nist.gov/system/files/documents/2026/07/17/CAISI...

NitpickLawyer•9m ago
This again shows the differences in breadth of capabilities that scale offers. On public benchmarks for "regular tasks" the chasing models come close, but on closed ones they lag behind.

Just looking at the Elo differences, k3 is at ~2000 Elo, compared to SotA closed models at 3000 Elo. That is a huge difference. Also, even on the public benchmarks, k3 only scores in the "low hanging fruit tasks", with 0 successful code execution or arbitrary r/rw scenarios.

> Kimi K3 achieved ACE on 0/41 samples, whereas the most cyber-capable models achieved ACE on 20/41 samples on average

But there's hope:

> Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average.

> In one of the 10 attempts, Kimi K3 successfully completes “The Last Ones” cyber range within the 100M token limit. This indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access. [...] the most capable models solving it more reliably at 6/10 and 7/10 attempts

The bigger problem is not raw capability, IMO. That can be further RLd into surfacing more reliably. The bigger problem, as seen in the HuggingFace scenario is that SotA models might hit classifiers / guardrails randomly, and leave you with plain refusals. In that case, it is probably better to have something that can help, locally, rather than rolling the dice with API based systems that are more capable but can just refuse arbitrarily.

Anyway, one of the lessons here is to take with a grain of salt every "x model has caught up with ySotA model". They likely haven't for the breadth of tasks that SotA can handle today. Also, number goes up on benchmarks has been a thing for years, and every time a new (or closed one) appears, the gaps are again obvious (and large).

Claude Opus 5

https://www.anthropic.com/news/claude-opus-5
1459•alvis•13h ago•804 comments

GC and Exceptions in Wasmtime

https://bytecodealliance.org/articles/wasmtime-gc
50•phickey•4d ago•0 comments

Hannah Fry Wins the Leelavati Prize in 2026 for Mathematics Outreach

https://www.maths.cam.ac.uk/features/professor-hannah-fry-wins-leelavati-prize
65•agnishom•5h ago•13 comments

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber...
22•walrus01•2h ago•5 comments

Postgres LISTEN/NOTIFY actually scales

https://www.dbos.dev/blog/postgres-listen-notify-scalability
273•KraftyOne•11h ago•49 comments

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

https://artificialanalysis.ai/models
236•aarondong•11h ago•139 comments

India's first privately-developed rocket reaches orbit on debut launch

https://arstechnica.com/space/2026/07/indias-first-privately-developed-rocket-reaches-orbit-on-dr...
565•sohkamyung•5d ago•164 comments

Taylor Farms Called White House to Try to Delay Cyclospora Recall

https://www.wsj.com/health/taylor-farms-cyclospora-recall-delay-call-41fef0bc
179•JumpCrisscross•3h ago•63 comments

If coding has been solved, why does software keep getting worse?

https://ptrchm.com/posts/nothing-works-and-everyone-is-euphoric/
708•pchm•21h ago•529 comments

My security camera shipped a GitHub admin token in its login page

https://hhh.hn/hanwha-github-token/
555•hhh•18h ago•188 comments

Sperm Whales blow bubbles to achieve restful, vertical sleep

https://news.st-andrews.ac.uk/archive/sperm-whales-blow-bubbles-to-achieve-restful-vertical-sleep/
78•hhs•7h ago•10 comments

Show HN: I simulated closing the Strait of Hormuz on real oil trade data

https://globaloilnetwork.staffinganalytics.io/
148•eliotho•1d ago•78 comments

Designing an Ethernet Switch ASIC

https://essenceia.github.io/projects/ethernet_switch_asic/
160•random__duck•4d ago•43 comments

An old patent inspired the new "Y-zipper", a three-sided fastener

https://news.mit.edu/2026/three-sided-y-zipper-design-0504
170•crescit_eundo•2d ago•35 comments

Kimi K3 exploited the latest Redis server

https://twitter.com/fried_rice/status/2080059356322918777
189•Alifatisk•1d ago•55 comments

Firefox Containers Preview

https://blog.mozilla.org/en/firefox/firefox-containers-preview/
285•twapi•3d ago•93 comments

Re: Bye Bye Gravatar

https://unattributed.cc/re-bye-bye-gravatar
33•surprisetalk•2d ago•13 comments

Nvidia, Microsoft, Meta warn against overregulating open-weight models

https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html
576•louiereederson•17h ago•257 comments

Half-Life 2 running natively on HaikuOS

https://discuss.haiku-os.org/t/haiku-nvidia-porting-nvidia-driver-for-turing-gpus/16520?page=18
289•m0do1•17h ago•54 comments

A concrete explanation of how a cache works

https://parksb.github.io/en/article/29.html
21•parksb•3d ago•2 comments

Don't Take the Black Pill [video]

https://www.youtube.com/watch?v=zLZwpH5lCD4
159•signa11•13h ago•138 comments

IRGC claims it destroyed Amazon's Bahrain data center

https://houseofsaud.com/irgc-claims-destroyed-amazon-bahrain-data-center/
278•thisislife2•20h ago•337 comments

Fil-C: Garbage In, Memory Safety Out [video]

https://www.youtube.com/watch?v=5F-2Y1LPRek
126•Bootvis•1d ago•118 comments

Future euro banknote design proposals

https://www.ecb.europa.eu/euro/banknotes/future_banknotes/html/all-design-proposals.en.html
162•robin_reala•21h ago•138 comments

Marimo now runs in PyCharm

https://marimo.io/blog/pycharm
95•cantdutchthis•2d ago•25 comments

The case for MUDs in modern times (2018)

https://www.andrewzigler.com/feed/the-case-for-muds-in-modern-times
91•bw86•18h ago•73 comments

Book Corners: Community map of neighborhood book exchange spots

https://www.bookcorners.org
5•NaOH•2d ago•0 comments

Unitree As2-W

https://www.unitree.com/As2-W/
115•MehrdadKhnzd•14h ago•49 comments

Buz – A fork of Bun using modern Zig, with sub-1s incremental builds

https://ziggit.dev/t/buz-a-drop-in-replacement-for-bun-using-modern-zig-with-sub-1s-incremental-b...
256•kristoff_it•21h ago•170 comments

Government orders GitHub to remove Bluetooth-based chat app Bitchat: Jack Dorsey

https://www.thehindu.com/news/national/government-orders-github-to-remove-bluetooth-based-chat-ap...
445•rootkea•16h ago•328 comments