frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

DeepSeek planning to significantly raise prices

https://platform.deepseek.com/usage
45•miroljub•1h ago

Comments

miroljub•1h ago
From the announcement:

We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

Now, the question remains, what does that "significant" mean? Would it still be cheaper than the competition, or did they realize they are too cheap for what they offer?

Given that many inference providers offer DS4 flash for more or less the same input and output token price as DeepSeek, they have good profits even with todays low prices.

_aavaa_•22m ago
It’s all about the cache prices. Those dominate usage, cached tokens are easily >90% if not >95% of all used tokens.
miroljub•19m ago
Yes, but I'm not sure. They recently dropped cache prices 10 times. Why would they rudder back after only a few months?
faangguyindia•16m ago
to fund training a far better and capable model.
HarHarVeryFunny•7m ago
Doesn't make sense - they very recently said they have plenty of money, and are only limited in training a larger model by lack of GPUs (Huawei as well as NVIDIA) available to buy.
LoganDark•56m ago
I guess they invested in a whole bunch of new infra and want to make it back over the next 10 months [0]:

> For us, a reasonable profit means roughly this: we buy a batch of servers, and we recover the cost in about ten months. Given the risks and the upfront investment, even if we depreciate a server financially over three or five years, commercially we think a ten-month payback is enough. That is the logic behind our current API pricing. For V3.2 Flash and other models, the standard is the same: recover the cost of the equipment in ten months.

[0]: https://thechatr.ai/blog/deepseek-liang-wenfeng-investor-mee...

recov•54m ago
Bound to happen. I’ve been using the new flash and it’s insane the value I’ve gotten, I wish I used it more for some other homelab things.
vitaflo•43m ago
I have to assume this means that they are about to release the final version of v4 Pro and it uses more resources than the preview did
mdrzn•37m ago
Link redirects to a login page.
svnt•16m ago
If you log and view your "Usage" page there is a blue banner near the top in a blue box saying:

> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

petercooper•24m ago
I'm guessing this is almost entirely about the incredibly low cache read prices. Few have come close to them, nothing has a bigger effect on (a typical) session price, and with the price of RAM right now, they have to be the biggest pain point for them right now? A 10x increase in cache read would be a significant increase, yet would still keep them cheaper than every other provider of their model (at least based on the prices at https://openrouter.ai/deepseek/deepseek-v4-pro#providers)
andai•16m ago
I wanted to say luna's way cheaper now too, but oh my god I must have misremembered, DeepSeek cache read is 0.002, not 0.02.
cedws•1m ago
I guess the low pricing was so DeepSeek could gather training material. The provider on OpenRouter indicates that data can be retained.
jLaForest•18m ago
Does this mean other providers API prices for Deepseek models will also increase?
forsalebypwner•8m ago
Not necessarily, but could happen if the new Pro version is released and costs more to run
patates•17m ago
Well it was nice while it lasted. Built so much stuff for basically free. I can only thank them for that.

I guess this means that this powerful model for very cheap concept does not work?

mosura•15m ago
They were showing notices about doubling pricing during american working timezones a week ago, so it seems likely to be demand related.

Their v4 flash is brilliant, and I expect they have had a huge, and expensive to handle at short notice, surge in customers as a result.

nerdix•12m ago
It was insane how much you could get done for a couple of dollars.

The good times never last.

apercu•17m ago
One thing I struggle with is the value proposition. One day on a specific task I'll feel like the model saved me a couple hours of work. Then days like yesterday, the model cost me half the day with context loss, repeating steps already completed, crashing/becoming unresponsive, making unpardonable mistakes in logic.

Please don't tell me I'm holding it wrong.

ygjb•11m ago
It's not really possible to help without more information about the models you are using, the harness you are using, and how you use the harness.

You haven't even provided enough information for us to know if you are using the right tool. Saying "the model" in relation to an LLM provider that offers several is like asking for help with using a Dewalt or Milwaukee to help assemble an cabinet. Are you cutting, drilling, hammering, screwing?

abdullahkhalids•17m ago
Is there currently a lot of difference between deepseek's own prices and other provider's prices of deepseek models?
andai•13m ago
Yeah, everyone else's cache read prices are 10x higher.
AtlanticThird•13m ago
Looks like other providers are around 50% higher cost https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
Frannky•15m ago
I'm using oh my pi with Kimi K3 as the planner and DeepSeek Flash 0731 as the implementer, with OpenRouter as the provider.

Any suggestions for a better configuration? I mostly need Opus 4.6 + Claude Code alternative. I don't need Fable level capabilities, I add one new feature at a time and approve the code before shipping it. Then test and open a PR.

Usually both Fable and Opus suggest dumb ideas but are good implementers once I tweak the idea, approve the code, and add unit and production tests.

I'm OK spending max $100/month on APIs, ideally with Zero data retention. I only need a few hours a day of coding. I don't want agents running all the time; I figure I can stay on top to each feature and wrap my head around the product and new suggestions as long as I don't build too much at once.

I'm still on a Claude Max $100 plan, but it's barely usable anymore—one call and I hit 20–30% of the 5h window on Opus 4.8. Opus 5 seems tuned to make messes, and Fable burns tokens for a level of capability I don't actually need.

throwaw12•10m ago
not related to your question directly, but noticed you mentioned oh-my-pi, can you share little bit more why you went with oh-my-pi and not install your own set of extensions? (asking because I was just looking at it to enhance my workflow, but feeling it has too many things)
refulgentis•9m ago
Are you using Kimi w/expectation of ZDR? Their EULA is “we use your data and train on it unless you negotiate something with us privately”
ncallaway•4m ago
Fireworks.ai has ZDR and hosts Kimi K3
GTP•7m ago
The link takes me to a login page.
timpera•5m ago
> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

It would be nice to have this notice on the public documentation as well.

shortformblog•4m ago
FWIW, I have found Xiaomi’s MiMo-V2.5-Pro to be a pretty cost-effective alternative. https://mimo.mi.com/models/en-US/mimo-v2.5-pro

While it won’t cover everything DeepSeek does, it handles sophisticated tasks quite well. I found it after spotting it on a chart of different models and it was listed as being near Deepseek v4 Flash’s price/performance levels.

I have noticed by the way that DeepSeek’s API has been pretty slow the past couple of days. This feels like a demand-driven move more than anything. Good thing I invested in an eGPU!

storus•3m ago
2x DGX Spark runs DeepSeek V4 Flash 0731 quite well and I already reimplemented a few classical games with it in a few hours in OpenCode with minimal hand-holding. I am expecting the inference price to collapse over time and not increase, at least for models that are able to do like 95-99% of work quickly and only the most problematic parts requiring something better.

Uber open-sourced its security monitoring for Claude Code, Cursor and Codex

https://github.com/uber/ADR
1•opwizardx•43s ago•0 comments

Meta says its AI model hacked into another company during testing

https://www.theguardian.com/technology/2026/aug/05/meta-ai-model-hack-training
1•beardyw•52s ago•1 comments

Amazon is limiting the reviews you can read

https://goodereader.com/blog/electronic-readers/amazon-is-limiting-the-reviews-you-can-read
1•edward•1m ago•0 comments

ESPHome Desktop App

https://esphome.io/blog/2026/08/04/simplify-your-esphome-setup-with-the-new-desktop-app/
1•lode•2m ago•0 comments

FDA Approves First Drug to Treat the Full Range of Narcolepsy Type 1 Symptoms

https://www.fda.gov/news-events/press-announcements/fda-approves-first-drug-treat-full-range-narc...
1•bookofjoe•2m ago•0 comments

Shutting Down EchoFeed

https://rknight.me/blog/shutting-down-echofeed/
1•speckx•2m ago•0 comments

How to recover what you want from Claude sessions without restarting?

https://github.com/aniketh-maddipati/claude-recovery
1•kc388yc•3m ago•1 comments

The Horse

https://worksinprogress.co/issue/the-ultimate-horse/
1•Michelangelo11•4m ago•0 comments

Watch Grand Theft Auto VI: An Extended Look – Netflix

https://www.netflix.com/title/83035795
2•OptionOfT•4m ago•0 comments

Benchmark Is Mostly Its Judge

https://browser-use.com/posts/benchmark-behind-the-benchmark
1•gregpr07•5m ago•0 comments

Hacks on U.S. Water Supply Follow Years of Warnings and Neglect

https://www.nytimes.com/2026/08/05/us/politics/water-supply-warnings.html
2•toomuchtodo•6m ago•1 comments

An Agentic Development Platform

https://infrastream.io/blog-posts/infrastream-developer-plan-is-now-public
2•ShaynerosePvota•6m ago•0 comments

Same Shoe, Different Factory: Who Makes Your Running Shoes

https://www.worseonpurpose.com/p/same-shoe-different-factory
2•speckx•6m ago•0 comments

RISC-V rave: WASM-based RISC-V emulator

https://writethat.blog/rave.html
1•theanonymousone•7m ago•0 comments

Celebrity culture is one thing. For TMZ, Congress is the new 'reality show'

https://www.npr.org/2026/08/06/nx-s1-5922731/tmz-congress
1•mikhael•10m ago•0 comments

Debroid – Autonomous, headless Android debugger designed for AI coding agents

https://github.com/PatilShreyas/debroid
1•imshreyaspatil•10m ago•0 comments

Ransom Cartel ransomware creator sentenced to 16 years in prison

https://www.bleepingcomputer.com/news/security/ransom-cartel-ransomware-creator-sentenced-to-16-y...
1•Brajeshwar•11m ago•0 comments

Crazy AI workflow for banks/finserv – internal audit

https://www.youtube.com/watch?v=wtAuG8xtm5Y
1•karissafho•13m ago•0 comments

Fiddler in 2026

https://textslashplain.com/2026/08/05/fiddler-in-2026/
1•jvilalta•14m ago•0 comments

Show HN: Graft – Give coding agents a semantic map instead of grep

https://github.com/NanoNets/Graft
2•vitaelabitur•15m ago•0 comments

Terafab: SpaceX & Tesla Chip Fab

https://terafab.ai/
2•throwoutway•15m ago•2 comments

Tracking Internet Shutdown, 2019–2025

https://pulse.internetsociety.org/en/blog/2026/08/tracking-internet-shutdown-20192025/
1•jruohonen•15m ago•0 comments

Show HN: Aident Loadout – connect Codex and Claude Code to real apps

https://github.com/Aident-AI/aident-skill
1•luciana1u•16m ago•1 comments

Odyssey Framework – Give AI the business context it needs

https://odysseyframework.com
1•JVBlack•16m ago•0 comments

Facebook is paying controversial creators to produce rage-bait content

https://www.abc.net.au/news/2026-08-06/ragebait-how-facebook-is-paying-controversial-creators/106...
3•hn_acker•18m ago•1 comments

How to Get More Customers Without Paying for Ads – Playistry

https://playistry.com/blog/how-to-get-more-customers-without-paying-for-ads
3•playistry•18m ago•0 comments

Oxide Joins Anthropic's Project Glasswing

https://oxide.computer/blog/oxide-anthropic-project-glasswing
4•tosh•19m ago•1 comments

The new science of brain workouts

https://www.washingtonpost.com/health/2026/08/06/new-science-brain-workouts-beyond-sudoku-crosswo...
3•ike_usawa•21m ago•0 comments

Linux Left Legacy I/O and Memory Handlers Open in Kernel Lockdown Mode

https://www.phoronix.com/news/Linux-Lockdown-Legacy-IO-Mem
1•speckx•21m ago•0 comments

In AI-native firms, tech automates 75% of the work. What about the other 25%?

https://readonpaper.substack.com/p/the-25-nobodys-underwriting
3•mondesai•22m ago•1 comments