frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

https://enclave.ai/blog/deepseek-v41-flash-is-now-our-best-hacking-model
70•talhof8•3h ago

Comments

fwip•44m ago
Edit: Comment deleted - no longer useful.
tyingq•33m ago
Apparently they updated it perhaps based on your comment?

> . The accepted runs cost $4.65. Failed attempts and replacement runs increased the complete cost to $5.14.

fwip•25m ago
Oh - I see that now. It's possible I missed it originally - I did read the article but I was skimming quickly. Mea culpa, if so.
Aldipower•32m ago
You can have 100 runs for the price of the Claude Max plan?
TuxSH•29m ago
I find this - or perhaps the title - a bit surprising.

I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

Perhaps DS works better where targets have low-hanging fruits than can be found fast?

simlevesque•23m ago
In your example, couldn't you parallelize DS's work more ? You could have 11 times as many agents for the same price.
mariopt•8m ago
Makes sense, GLM is a lot better than DS v4.1, I found the same results in other domains.

Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good with if provide custom tools for the task, on it's own it's mediocre.

allie1•7m ago
I think it really depends on what the data DS was fine tuned on. If your use case is very specific, it wouldn’t have distilled that knowledge well.
nickysielicki•16m ago
When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.
EGreg•11m ago
Win what? The race to the bottom always has this competitive language.

“If we ban CFCs now the Chinese will win!”

“If we ban chemical weapons, nuclear weapons, etc etc our enemies will triumph! They won’t stop!”

“If we switch to biodegradeable plastic then our rivals will have an advantage.”

“If we dont externalize the costs to our population, then they will, and then will win!”

I think workflows can do the job agents do, 20x cheaper and more predictably and safely. They can completely displace agents, just as HFCs displaced CFCs and then we were able to ban CFCs and phase them out through international COOPERATION. The language of COOPERATION is what saves us vs COMPETITION is all about cutting corners and externalizing costs. Google the Montreal Protocol, Geneva Conventions, Nuclear Non Proliferation Treaty, Unleaded Gasoline etc etc.

Agents have got to be marginalized. They are just popular because the labs need to make a ton of money for their investors and recoup their massive spending on training models.

airstrike•10m ago
[delayed]
jrflo•13m ago
Seems pretty bold to claim deepseek is the "best hacking model" while providing zero comparisons to other models...

PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand"

https://frvr.com/blog/news/ps5-linux-lead-quits-as-open-source-projects-have-become-a-bunch-of-no...
193•alexjplant•1h ago•101 comments

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

https://enclave.ai/blog/deepseek-v41-flash-is-now-our-best-hacking-model
75•talhof8•3h ago•14 comments

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

https://arxiv.org/abs/2609.14858
55•bananaflag•1h ago•15 comments

Mistral X Mozilla: Private, Multilingual AI Browsing

https://mistral.ai/news/mistral-x-mozilla/
306•vertigoruntime•7h ago•98 comments

Salesforce Global Outage

https://status.salesforce.com/products/all
194•mabil•5h ago•108 comments

Tell the speakers that you liked their talks

https://ohhelloana.blog/tell-the-speakers/
25•whisper2020•1d ago•6 comments

Show HN: Free WhatsApp MCP (+UI) – Give Your AI Agents Access to WhatsApp

11•fabian_shipamax•40m ago•15 comments

The Google Play app review process now regularly takes longer than a week

https://gultsch.social/@daniel/117280438824908947
233•inputmice•4h ago•203 comments

Introducing System One Models and Jev

https://typesafe.ai/blog/introducing-system-one-models-and-jev
1650•albelfio•20h ago•455 comments

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations

https://github.com/arnegiacomo/fugleramme
1832•arnemunthekaas•1d ago•218 comments

Hackers Got Inside a Flock Camera. Its Data Shows How the System Works

https://www.wired.com/story/hackers-flock-camera-data-shows-how-system-works/
176•driverdan•2h ago•78 comments

How Big Are Factorials?

https://eli.thegreenplace.net/2026/how-big-are-factorials/
17•ibobev•1d ago•8 comments

Kyber (YC W23) Is Hiring a Forward Deployed Engineer

https://www.ycombinator.com/companies/kyber/jobs/eturrAR-forward-deployed-engineer
1•asontha•3h ago

Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models

https://stale.jock.pl/
23•joozio•2h ago•17 comments

Apple Reference Image: A New Approach for Verified Photography

https://security.apple.com/blog/apple-reference-image/
394•imwally•13h ago•256 comments

Scaling Golang CI by Replacing actions/setup-go

https://www.cloudx.ai/posts/setup-go
38•peterldowns•2h ago•3 comments

Douglas Adams and the exterminated Doctor Who adventure

https://www.bbc.co.uk/news/articles/c8jdp38z4jgo
50•6LLvveMx2koXfwn•2d ago•31 comments

Original Sony PlayStation 2 security chip 'broken wide open' after 26 years

https://www.tomshardware.com/video-games/playstation/26-year-old-sony-ps2-security-chip-broken-wi...
137•rbanffy•3h ago•37 comments

An update on Wayback Machine access

https://blog.archive.org/2026/09/15/an-update-on-wayback-machine-access/
628•ChrisArchitect•21h ago•338 comments

ImpactGate: A merge gate that scores the structural decay AI adds

https://github.com/officefloor/ImpactGate
33•sagenschneider•2h ago•38 comments

Microsoft says AI rival Anthropic could have 'disastrous impact' on humanity

https://www.bbc.co.uk/news/articles/c6n07ypqz8kzo
23•pluc•1h ago•22 comments

Doing Everyone Else's Job

https://yosefk.com/blog/doing-everyone-elses-job.html
160•luu•1d ago•74 comments

Show HN: I made a flight simulator, except you're just a passenger

https://inflightsimulator.com
312•rkotcher•2d ago•172 comments

Gemini 3.8 Live and 3.8 Live Extended Thinking

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-...
463•leumon•22h ago•309 comments

Anatomy of a Texture

https://agentlien.github.io/texture/
6•Agentlien•1h ago•1 comments

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

https://arxiv.org/abs/2511.07885
116•pythonic_hell•2d ago•36 comments

Why I'm still bearish on LLMs after Navier-Stokes

https://dank.systems/posts/2026-09-15-ai-bear.html
350•jaykru•22h ago•444 comments

A software thing I built: GPS on a 25MHz 486-SX

https://forum.vcfed.org/index.php?threads/a-software-thing-i-built-gps-on-a-25mhz-486-sx.1258966/
57•JPLeRouzic•10h ago•19 comments

German Rheinmetall open-sources its Battlesuite connected weapon system protcol

https://rheinmetall.github.io/onboardapi-documentation/9.10.0/index.html
259•summarity•18h ago•97 comments

Recreating Voodoo Graphics and a Late-1990s Gaming PC on an FPGA

https://nand2mario.github.io/posts/2026/zsst-voodoo/
165•zdw•16h ago•49 comments