frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

58•nickweb•1h ago
DSeek plans to officially release the V4.1 Flash model around September 10, 2026 (Beijing Time). After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time. In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!

We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.

Comments

neugls•1h ago
Waiting to use it
oefrha•47m ago
Source is apparently a banner announcement on https://platform.deepseek.com/usage. Had me searching for a couple minutes...
nickweb•23m ago
I swear I put that at the start of the post. Must've managed to miss it when copy and pasting!
swiftcoder•33m ago
If they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time
tarruda•27m ago
Hopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
fluoridation•14m ago
Interesting, I had assumed it'd be too large to fit. What quant and context size are you running?
nickweb•20m ago
Via nitter: https://xcancel.com/JustinGorya/status/2097287080128708930

Looks like the new model can be used if summoned via the API but the API won't list it.

igleria•18m ago
v4 pro was decent then a better cheaper faster model comes now?

As a consumer I feel like hansel and gretel combined, deepseek could be the witch.

throwaway473825•6m ago
It's not unprecedented given that GLM 5.3 Flash was better and cheaper than GLM 5.2.
nicce•9m ago
> In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!

Wow. Imagine OpenAI/Google/Anthropic doing this! Nope.

thrownaway561•6m ago
I will continue to be amazed by how much power you get from DeepSeek Flash for the cost. I have let that puppy lose on so many projects and it is has never let me down. It can build and entire Rails app in no time and even do the tests. For most things, I don't get why people pay the money for Claude. DeepSeek Flash is my default agent in Omarchy.
EbNar•4m ago
Since a few months, I almost exclusively use the Chinese "flash" models for my need. They are a joy and they cost pennies per answer. Great job.

DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

64•nickweb•1h ago•13 comments

First HN Post

3•aeroscissorz1•5m ago•0 comments

Ask HN: 3.5 inch diskette read errors, would a period correct drive do better?

32•rietta•1d ago•35 comments

I'm going back to coding by hand

39•trencedamp•5h ago•28 comments

Ask HN: Are others seeing Google's reCAPTCHA rejecting Firefox users?

252•Animats•5d ago•127 comments

Ask HN: Fable hacked my piano, can I release the results?

295•jmpman•3d ago•163 comments

Ask HN: How do you manage skills files?

308•imadtaieber•2d ago•285 comments

Apparently CodePen 2.0 sends data to their servers as you type

112•maxim-fin•2d ago•59 comments

Ask HN: Are you leveraging Spec-Driven / Spec-Anchored development?

4•locusofself•13h ago•3 comments

Ask HN: Who is using MCP in production?

198•sukit•6d ago•200 comments

Ask HN: Is ageism in tech still a problem in 2026?

3•leonagano•11h ago•1 comments

Ask HN: Would you read a statistics textbook?

106•usernametaken29•3d ago•66 comments

Ask HN: Would you hire a new engineer, or get 300k worth of tokens for the team?

5•websap•12h ago•8 comments

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

404•halcdev•5d ago•705 comments

Ask HN: Show your micro-SaaS

22•genekrapivin•2d ago•21 comments

The universal programming language of LLMs

5•senorqa•14h ago•9 comments

Tell HN: OpenAI keeps stealing my money

10•dingdong2026•19h ago•3 comments

Ask HN: Any Software Engineers here who enjoy their AI-native dev workflow?

12•pkos98•2d ago•12 comments

Ask HN: Anyone using dictation with coding agents?

3•TomEleff•22h ago•7 comments

Tell HN: OpenAI brings back 5 hour limit for plus and business standard users

127•spwa4•1d ago•146 comments

Ask HN: Connecting Kubernetes dependencies to application telemetry

6•chipfixer•2d ago•2 comments

Ask HN: UK Rescue Rocket Sheds/Houses Information

24•burnt-resistor•2d ago•6 comments

Tell HN: Both recent GCP outages caused by fiber optic maintenance

20•fastest963•5d ago•1 comments

DeepSeek v4.1 Flash is now available for internal beta testing

21•dares2573•1d ago•8 comments

Ask HN: (Why) Was the LLM breakthrough useful for images, audio, etc.?

4•rogerrogerr•2d ago•5 comments

Tell HN: I want to see the same moon as you

16•UnderABlueMoon•4d ago•16 comments

Ask HN: Best Platforms to Orchestrate Agents

3•venkat971•1d ago•2 comments

Ask HN: How safe are our password managers in face of LLM cyber attacks?

3•muddi900•2d ago•1 comments

Bun rewrite and FLT formalization had nearly identical resource usage

5•Zsfe510asG•1d ago•1 comments

Is OpenAI silently routing GPT‑6 requests to GPT‑4o?

2•skaiuijing•1d ago•4 comments