frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

DeepSeek-V4-Flash Update

https://api-docs.deepseek.com/updates/
54•dnhkng•1h ago

Comments

dnhkng•1h ago
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

Iolaum•47m ago
Since they did this with their own harness I m not sure it's apples to apples comparison.
NitpickLawyer•44m ago
> not sure it's apples to apples comparison.

They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

dnhkng•41m ago
I think the commenter means the Flash vs Terra benchmarks.
dnhkng•42m ago
It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

yms_hi•28m ago
I think it's better than GPT Luna.
throwaw12•21m ago
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well
villish•10m ago
> Terminal Bench: Flash 82.7 vs Terra 78.4

Terra 87.4

https://openai.com/index/gpt-5-6/

NitpickLawyer•55m ago
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

dnhkng•40m ago
Totally! This with DwarfStar delivers usable local AI (I hope!)
spwa4•54m ago
In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
wolttam•49m ago
It runs really well on 2 DGX Sparks - 60t/s
Reubend•24m ago
They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.
nickandbro•22m ago
I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

I admire DeepSeek's openness, but even they have been raising prices after their discounts.

minraws•14m ago
They haven't raised prices though the plan was to increase it with peak hour usage for V4 Pro GA release, they didn't do that, so no price increases there I believe.

As for vision yeah it sucks but Luna is also 2x input and 1.5x output for 1M context...

That's around 0.4 in/1.8 out

DSv4 is wayyy cheaper.

And it's open now you have Luna at home if you have a decent set of GPUs you can run this on 2Sparks or one very expensive Mac or just like 6-8 5090s..

minraws•9m ago
I know most folks can't afford it I am working on making it viable to rent shared hosting the biggest issue is data leak and prompt injection attacks with shared hosting. (Since the server owner connects to your main system via the coding agent)

I guess using a ZDR provider is good enough for now.

dudisubekti•12m ago
Gpt 5.6 Luna cache read is $0.02 per mtok

V4 flash cache read is $0.0028 per mtok

That's not "a bit cheaper", just saying

wolttam•21m ago
Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.
ggcr•14m ago
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

KronisLV•8m ago
Wonder how good the proper version of V4 Pro will be.

I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

Show HN: Slaide, open-source Markdown slides AI writes and PowerPoint opens

https://getslaide.com/
1•ppiper•35s ago•0 comments

Iptv safe: a privacy checklist before you sign up

https://iptvstreams.tv/app
1•sarahgreen•3m ago•0 comments

Musk dismisses report of Tesla's potential China business sale

https://www.reuters.com/business/media-telecom/tesla-weighs-sale-china-business-pave-way-potentia...
1•ulfw•3m ago•0 comments

At the end of the day, my slaves are just a tool

https://www.lesswrong.com/posts/Dfz8dFNtqSei4F2qA/at-the-end-of-the-day-my-slaves-are-just-a-tool
1•jesseduffield•4m ago•0 comments

Medieval Ideas About Food

https://lithub.com/fish-bad-sugar-good-and-other-medieval-ideas-about-food/
1•simplegeek•5m ago•0 comments

Beltrunner: Game Design Postmortem

https://blog.gingerbeardman.com/2026/07/30/beltrunner-game-design-postmortem/
1•tobr•5m ago•0 comments

UK's only all-female chess team promoted – then told they must recruit a man

https://www.theguardian.com/sport/2026/jul/31/uk-female-chess-team-league-rule-men-join-sport
1•beardyw•6m ago•0 comments

Hi, Jinks – a new way to make games

https://jinks.gg/
1•tobr•7m ago•0 comments

The Software Engineering Practice Atlas

https://eng-atlas.dev
1•duyetdev•9m ago•0 comments

From Backtests to Live Crypto Trading Without Rewriting Strategy Code

https://medium.com/@DolphinDB_Inc/from-backtest-to-production-implementing-live-cryptocurrency-tr...
1•Polly_Liu•13m ago•0 comments

Show HN: I stopped babysitting my AI agents by pushing them to Telegram

https://blackflare.dev/
2•TrungTin•16m ago•0 comments

Samsung's chip workers are jumping ship to rival SK Hynix

https://www.technologyreview.com/2026/07/28/1140853/samsung-chip-workers-exodus-sk-hynix/
1•joozio•19m ago•0 comments

Top IPTV providers with free trial: what to evaluate in 2026

https://iptvproviders.online/
1•aramirez454•21m ago•0 comments

Turning screen recordings into structured incident and bug reports

https://tentomushi.to
1•pperner•22m ago•2 comments

Use Twenty Questions Game to Benchmark LLMs

https://mindalyze-com.github.io/deep-20-bench/
1•p-heusser•25m ago•1 comments

Made in the USSR: 6 video games Soviets went crazy over (2020)

https://www.rbth.com/lifestyle/332384-video-games-soviet-russian-tetris
1•downbad_•27m ago•0 comments

What I've Learned in 45 Years in the Software Industry (2021)

https://www.bti360.com/what-ive-learned-in-45-years-in-the-software-industry/
1•downbad_•28m ago•0 comments

Apple's genius play – MacBook Neo

https://monkeylike.substack.com/p/apple-marketing-strategy-genius
1•TIJ•31m ago•0 comments

C64 Demo Effects Explained: Rodents in the Attic [video]

https://www.youtube.com/watch?v=uZ1atMUOUMU
1•robin_reala•32m ago•0 comments

Benchmarking Guardrails for AI Agent Safety

https://blog.mozilla.ai/can-open-source-guardrails-really-protect-ai-agents/
2•TangoBee•34m ago•0 comments

Cold-Bread Rebellion: Human fights back

https://thinkingtoasters.com/2026/07/30/cold-bread-rebellion/
1•yesiot•39m ago•0 comments

HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA

https://github.com/password162156/hexcore-llm
2•XXXGhost•42m ago•0 comments

A Brief History of the Internet's Favorite Scam

https://thereader.mitpress.mit.edu/a-brief-history-of-the-internets-favorite-scam/
2•the-mitr•43m ago•0 comments

Show HN: A Handwritten Blogging Platform

https://handwritten.blog/
2•emilesilvis•44m ago•4 comments

Judge says Trump still lacks evidence for Anthropic supply-chain risk

https://techcrunch.com/2026/07/30/judge-says-trump-admin-still-lacks-evidence-for-anthropic-suppl...
1•lai0602•45m ago•0 comments

How to speed up the Rust compiler in July 2026

https://nnethercote.github.io/2026/07/31/how-to-speed-up-the-rust-compiler-in-july-2026.html
1•throawayonthe•45m ago•0 comments

Best Screen Recording Software for Mac in 2026

https://screensage.pro/blog/best-screen-recording-software-mac
2•auv1107•47m ago•0 comments

Show HN: One page personal calendar with hierarchy

2•eltonlin•48m ago•0 comments

How to plant a nuclear plant in Iran

https://www.digitaldigging.org/p/how-to-plant-a-nuclear-plant-in-iran
3•LordAtlas•48m ago•0 comments

I Migrated a 350k-Line Java/JSP Application to TypeScript in Five Days

https://gokulakrishna.co/2026/07/31/migrated-350000-line-java-jsp-application-typescript-five-days/
2•gkrishna•50m ago•0 comments