frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

The session you cannot take with you

https://earendil.com/posts/session-portability/
242•apitman•4h ago•49 comments

JEP 401: Value Objects (Preview) merged to OpenJDK master

https://github.com/openjdk/jdk/pull/31120
80•mfiguiere•3h ago•23 comments

DeepSeek-V4-Flash Update

https://api-docs.deepseek.com/updates/
151•dnhkng•2h ago•48 comments

Stacked PRs are now live on GitHub

https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview/
620•tomzorz•15h ago•206 comments

Show HN: What should the GUI for AI agents look like?

https://marbleos.com/demo
33•akbabu•2h ago•15 comments

I flagged two research papers for fake authors and both were accepted as orals

https://geospatialml.com/posts/reviewing-ai-slop/
177•volumes94•9h ago•76 comments

Show HN: Gander, an Android file viewer that asks for no permissions at all

https://github.com/mokshablr/gander
27•mokshablr•2h ago•11 comments

Gemini Robotics 2 brings whole body intelligence to robots

https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
547•ai2027•17h ago•436 comments

Where USB Memory Sticks are Born (2013)

https://www.bunniestudios.com/blog/2013/where-usb-memory-sticks-are-born/
42•jacquesm•3d ago•3 comments

The mean means nothing: data visualization to debug a latency problem

https://fzakaria.com/2026/07/27/the-mean-means-nothing
26•fanf2•1d ago•2 comments

Read this before you buy that TV streaming stick

https://krebsonsecurity.com/2026/07/read-this-before-you-buy-that-tv-streaming-stick/
694•speckx•15h ago•397 comments

Simulating TCP loss and congestion in browser using Go/WASM

https://ccsim.fly.dev
9•dilyevsky•2d ago•0 comments

Physicists Solve a Muon Mystery. Now, Old Results Don't Add Up

https://www.quantamagazine.org/physicists-solve-a-muon-mystery-now-old-results-dont-add-up-20260729/
212•ibobev•16h ago•129 comments

The AI Aesthetic

https://blog.jim-nielsen.com/2026/ai-aesthetic/
291•montroser•8h ago•126 comments

The Religion of Speed

https://graybeard.ing/the-religion-of-speed/
127•MobiusHorizons•8h ago•66 comments

The Economic Benefit of Refactoring

https://martinfowler.com/articles/exploring-gen-ai/refactoring-economic-benefit.html
232•javaeeeee•17h ago•96 comments

CodePen 2.0

https://chriscoyier.net/2026/07/30/codepen-2-0/
165•robin_reala•14h ago•47 comments

Rune 1.1: adds Python, an Emacs editor, a symbol index and is now free

https://rune.build/blog/rune-1-1-release
80•ernestrc•10h ago•27 comments

Bad Apple but It's Traceroute

https://jssfr.de/2026-07-27-bad-apple-but-traceroute.html
101•jssfr•3d ago•24 comments

The American Grilled Cheese Sandwich Essay (2024)

https://buttondown.com/theswordandthesandwich/archive/the-best-american-grilled-cheese-sandwich-e...
47•NaOH•3d ago•34 comments

Memo-1: A 6502 computer built from scratch, using a Minitel as its terminal

https://github.com/MemoireMorte/Memo-1
73•sciences44•2d ago•10 comments

GCC steering committee announces AI policy

https://lwn.net/Articles/1086041/
284•arto•20h ago•312 comments

UEFA and its national associations will not participate in FIFA competitions

https://www.uefa.com/news-media/news/02a7-213a92896eb0-54dfbf454e3b-1000--statement-on-behalf-of-...
986•dickfickling•13h ago•530 comments

Show HN: Cubic Doggo 06R: 12-DOF 4-Legged Robot with IMU

https://github.com/SphericalCowww/CubicDoggo_06R
3•SphericalCowww•5d ago•0 comments

The lost civic life of movie rental stores

https://thereader.mitpress.mit.edu/the-lost-civic-life-of-movie-rental-stores/
162•facundo_olano•18h ago•215 comments

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

https://www.ctgt.ai/research/distillation-censorship-transfer
122•cgorlla•14h ago•64 comments

Saber-toothed cats became inbred–and struggled to move–before they went extinct

https://www.science.org/content/article/saber-toothed-cats-became-inbred-and-struggled-move-they-...
49•gmays•10h ago•22 comments

Investigating three real-world incidents in our cybersecurity evaluations

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
176•surprisetalk•9h ago•136 comments

Advancing the price-performance frontier with GPT‑5.6

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
568•tedsanders•15h ago•366 comments

Why is everyone trying to build a solid-state battery?

https://www.construction-physics.com/p/why-is-everyone-trying-to-build-a
189•crescit_eundo•19h ago•239 comments
Open in hackernews

DeepSeek-V4-Flash Update

https://api-docs.deepseek.com/updates/
151•dnhkng•2h ago

Comments

dnhkng•1h ago
DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

Iolaum•1h ago
Since they did this with their own harness I m not sure it's apples to apples comparison.
NitpickLawyer•1h ago
> not sure it's apples to apples comparison.

They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

dnhkng•1h ago
I think the commenter means the Flash vs Terra benchmarks.
dnhkng•1h ago
It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

yms_hi•1h ago
I think it's better than GPT Luna.
throwaw12•1h ago
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well
bayesianbot•4m ago
IIRC GPT 3 was priced at per 1k tokens, had to check, the biggest GPT 3 model from OpenAI was $0.06/1k, so $60 / 1M. gpt-3.5-turbo was the first model after ChatGPT and that was $2 / 1M. And no caching. So not really in the same ballpark
villish•1h ago
> Terminal Bench: Flash 82.7 vs Terra 78.4

Terra 87.4

https://openai.com/index/gpt-5-6/

NitpickLawyer•1h ago
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

dnhkng•1h ago
Totally! This with DwarfStar delivers usable local AI (I hope!)
spwa4•1h ago
In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
wolttam•1h ago
It runs really well on 2 DGX Sparks - 60t/s
arjie•37m ago
Not yet, right? That's the old DS V4 preview release. We're still waiting for the weights to come out.
lukan•29m ago
"it'll barely run on an M5 Max "

The max version I could order now with 128 GB?

If so, the price for local inference would be 12 000 € vs 500 000 € for a B300.

NitpickLawyer•24m ago
There's also the 2x spark way, which should be ~8k eur? Someone down the thread reported ~60tps for 2x sparks. That's totally usable for local inference.

You can also do 2x 6kPRO in a workstation, for ~20k.

Reubend•1h ago
They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.
nickandbro•1h ago
I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

I admire DeepSeek's openness, but even they have been raising prices after their discounts.

minraws•1h ago
They haven't raised prices though the plan was to increase it with peak hour usage for V4 Pro GA release, they didn't do that, so no price increases there I believe.

As for vision yeah it sucks but Luna is also 2x input and 1.5x output for 1M context...

That's around 0.4 in/1.8 out

DSv4 is wayyy cheaper.

And it's open now you have Luna at home if you have a decent set of GPUs you can run this on 2Sparks or one very expensive Mac or just like 6-8 5090s..

minraws•1h ago
I know most folks can't afford it I am working on making it viable to rent shared hosting the biggest issue is data leak and prompt injection attacks with shared hosting. (Since the server owner connects to your main system via the coding agent)

I guess using a ZDR provider is good enough for now.

dudisubekti•1h ago
Gpt 5.6 Luna cache read is $0.02 per mtok

V4 flash cache read is $0.0028 per mtok

That's not "a bit cheaper", just saying

nickandbro•55m ago
That's a good point. Yeah their caching input is insane.
wolttam•1h ago
Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.
ggcr•1h ago
Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

KronisLV•1h ago
Wonder how good the proper version of V4 Pro will be.

I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

arjie•56m ago
Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).
wkcheng•45m ago
If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

Crazy.

egeozcan•44m ago
Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.
flysoft•42m ago
Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.
thirtygeo•42m ago
For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself
PhilippGille•38m ago
The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

try-working•36m ago
DeepSeek themselves called it `deepseek/deepseek-v4-flash`. Pro is still like that.
baalimago•38m ago
Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro?

Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which now becomes even better. Although, in the deepseek APIs, the cost is leaking all information about codebases to China.

try-working•35m ago
Let's see how the market reacts.
kamikazechaser•34m ago
The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.
storywatch•32m ago
How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.
f311a•29m ago
I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

regularfry•13m ago
It's good but (at least on openrouter) it's got an annoyingly tight output token limit. So if it does get stuck in a reasoning pit, it won't work its way out of it in time.

It's replaced the Kimi models for me though.

Goranek•28m ago
Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

Does this make sense?

geek_at•24m ago
it does! I'm always amazed when I use DSV4 flash for coding or server checks and after an hour of working with it my (pure api call) balance is about 30 cents
baalimago•18m ago
Looks like they will release an updated version of deepseek-v4-pro soon, which most likely will beat kimi k3 at both intelligence and cost (judging by the vast improvements to dsv4 flash)
try-working•13m ago
Yep. And you can use my model router to route between them in the background.
miyuru•27m ago
Judging by the openrouter leaderboard ranking for today, it looks like Dv4F us more popular than mimov2.5.

https://openrouter.ai/rankings?view=day#leaderboard-table

These days cost per task is more important, and SOTA models have become expensive.

ra•12m ago
What's the best way to run this on a 64GB M2 Pro?
lionkor•7m ago
I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

- Cost: $4.55USD

- API requests: 3,467

- Tokens: 323,183,886

And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

re-thc•35m ago
> I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng.

The leaked interview has him saying it doesn't matter... as much as open source doesn't matter. There's enough in it for everyone right now and they aren't after everything.

Perspective: DeepSeek doesn't have enough infrastructure to serve their target customers already.

nickandbro•28m ago
Can you post the link to the leaked interview? From what I understand he has been pretty tight lipped for a guy who has a larger stake worth more in his company than Dario does in Anthropic.
Rzor•18m ago
https://www.techflowpost.com/en-US/article/32744