frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

DeepSeek V4 Pro 0813

https://openrouter.ai/deepseek/deepseek-v4-pro-0813
204•explosion-s•1h ago•61 comments

Tailscale Traces Database Corruption to 16y/o SQLite WAL-Reset Bug

https://tailscale.com/blog/sqlite-wal-reset-bug
384•ropbear•3h ago•57 comments

DeepSeek V4 Pro 0813 quietly released

https://api-docs.deepseek.com/guides/responses_api/
58•HiPHInch•56m ago•5 comments

2026 Eclipse Webcams

https://jonty.github.io/2026_eclipse_webcams/
377•zoenolan•5h ago•89 comments

Tim King, AmigaDOS developer, has died

https://amiga-news.de/en/news/AN-2026-08-00070-EN.html
112•doener•3h ago•22 comments

Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot

https://knownagents.com/insights
119•gavinhking•3h ago•68 comments

Reflex (YC W23) Is hiring Growth and GTM Roles

https://www.ycombinator.com/companies/reflex/jobs/71x5GFb-growth-engineer
1•apetuskey•33m ago

Why Tiny JPEGs Look Different in Chrome

https://guillaumetech.github.io/posts/jpg-scaling-chrome/
143•gutechh•3h ago•26 comments

Wednesday, August 12: GitHub, Incident with Pull Requests and Issues

https://www.githubstatus.com/incidents/76t89hbfb09h
28•arm32•1h ago•3 comments

License plate reader searches should require a warrant

https://andrewpwheeler.com/2026/08/12/license-plate-reader-searches-should-require-a-warrant/
330•apwheele•2h ago•199 comments

AI is removing the middle class of software engineering

https://blog.florianherrengt.com/ai-removing-middle-class-software-engineering.html
399•florianherrengt•4h ago•331 comments

Glaciers on the Climate Dashboard

https://climate.metoffice.cloud/glaciers.html
19•mooreds•55m ago•2 comments

Qwen/Qwen3.8-2.4T-A95B

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
186•Philpax•2h ago•55 comments

HTML over WebSockets: real-time SPAs with barely any JavaScript

https://en.andros.dev/blog/ef4968f5/html-over-websockets-real-time-spas-with-barely-any-javascript/
11•redbell•42m ago•7 comments

What sort of maths are LLMs good at?

https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/
195•ColinWright•7h ago•90 comments

Bike Bureau: Report Bike Lane Obstructions

https://loudbicycle.com/bb
20•kdmccormick•1h ago•6 comments

Show HN: Woxi - Open-source Mathematica / Wolfram Language reimplementation

https://woxi.ad-si.com
204•adius•7h ago•29 comments

Hax – a minimalist, terminal-native coding agent written in C

https://usehax.dev/
35•OleksandrC•2h ago•7 comments

Shade Map

https://shademap.app
62•fredley•4h ago•16 comments

The Bit Player: My Father with Steve Zissou

https://www.theparisreview.org/blog/2026/07/27/the-bit-player-my-father-with-steve-zissou/
7•Thevet•4d ago•0 comments

Delphi 13 Community Edition Is Now Available

https://blogs.embarcadero.com/delphi-13-community-edition-is-now-available/
118•layer8•6h ago•83 comments

Felix and I

https://jacobfilipp.com/felix/
44•surprisetalk•2d ago•1 comments

Pixel Watch 5

https://blog.google/products-and-platforms/devices/pixel/pixel-watch-5/
11•ortusdux•1h ago•1 comments

Automatic1111 for Apple metal, 40% speed up sd1.5

https://therad.ninja/from-8-10-seconds-to-3-7-teaching-automatic1111-to-speak-metal-on-an-m3-pro/
45•dmikey831•3h ago•18 comments

Solving the Shortest Vector Problem in $2^{0.6039n}$ Time via Mid-Point Hessian

https://arxiv.org/abs/2608.02478
23•sbulaev•1w ago•4 comments

My Agent Setup

https://chad.cm/posts/2026-8-11-my-agent-setup
69•carimura•3h ago•36 comments

High-Res Photo Shows Sand-Capped Butte Rising from Mars Plain of Polygons

https://petapixel.com/2026/08/04/amazing-high-res-photo-shows-a-butte-rising-from-mars/
123•bookofjoe•6d ago•9 comments

AI coding startup Lovable raised $400M at $13.3B valuation up from $6.6B in 2025

https://lovable.dev/blog/series-c
15•thoughtpeddler•1h ago•6 comments

Bigos (Polish Hunter's Stew) Recipe Builder

https://chefsbinge.com/bigos-recipe-builder/
61•doublepg23•5d ago•21 comments

uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook

https://digitalescapetools.com/2026/08/ublock-origin-stops-chasing-facebook-ads.html
137•Markoff•6h ago•106 comments
Open in hackernews

DeepSeek V4 Pro 0813

https://openrouter.ai/deepseek/deepseek-v4-pro-0813
197•explosion-s•1h ago

Comments

aabdi•1h ago
https://api-docs.deepseek.com/quick_start/pricing/

Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.

swiftcoder•57m ago
How does it stack against the updated Deepseek Flash version?
k__•51m ago
Around 5 percentage points better. (E.g., 87% instead of 82%)
Gecko4072•47m ago
So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.
k__•44m ago
I tried the previous Pro model and in the end it was 50% more expensive than the previous Flash.

Wasn't worth it.

saaga•40m ago
Yea that's what I was thinking. Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure.

I was running a session over a couple days and it didnt cross a dollar lol.

networked•33m ago
I haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page (https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. They critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.
trollbridge•6m ago
Interesting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient.

I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.

sparkling•42m ago
deepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.
k__•40m ago
I wouldn't exactly call it snappy, but faster than Pro, yes.
ericd•36m ago
Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.
JacobAsmuth•26m ago
Well sure but you're running on tens of thousands of dollars of hardware.
saaga•40m ago
I feel the same too. I like the speed. I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.
pixelesque•36m ago
I've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev).

Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.

swiftcoder•33m ago
yeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow
surgical_fire•16m ago
I use a plan -> implement wotkflow for this reason.

pro plans, flash implements. I am super happy with how flash behaves like that.

JacobAsmuth•27m ago
Per token. You need to look at pricing per task.
trollbridge•9m ago
... which still comes out cheaper, since DeepSeek caches so much more.

I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.

xynelius•13m ago
If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]:

For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.

Cost per request for V4 Pro: $0.000875 per request.

Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.

[1] https://opencode.ai/docs/go/#usage-limits

indigodaddy•55m ago
@dang - Pls merge this with https://news.ycombinator.com/item?id=49274018
LeonKnst•52m ago
I find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
HawtAds•27m ago
Hacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.
BlackRabbit1•27m ago
A lot of it/infrastructure departments aren't aware that you can use Asian models hosted within the US or even EU.
spacebanana7•24m ago
In an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.
BlackRabbit1•13m ago
There are EU/US providers offering Deepseek/Qwen/Kimi/etc.-as-a-Service. With zero ties of their infrastructure to China.

Fully compatible with the well known Antrophic API.

You only have to replace the URL and your key.

scrlk•52m ago
Benchmarks:

    | Benchmark                | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2   | Kimi-K3   | Opus-4.8  | Fable 5       |
    |                          | 0813      | 0731        | Preview   | Preview     |           |           |           | (w/ fallback) |
    |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
    | HLE (wo/w tools)         | 42.7/60.0 | 37.8/51.5   | 37.7/48.2 | 34.8/45.1   | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0     |
    | Terminal Bench 2.1       | 87.9      | 82.7        | 72.1      | 61.8        | 81.0      | 88.3      | 85.0      | 88.0          |
    | NL2Repo                  | 61.5      | 54.2        | 38.5      | 39.4        | 48.9      | -         | 69.7      | -             |
    | Cybergym                 | 83.3      | 76.7        | 52.7      | 38.7        | -         | 80.0      | 78.3      | 83.1          |
    | DeepSWE                  | 62.7      | 54.4        | 12.8      | 7.3         | 46.2      | 67.5      | 58.0      | 70.0          |
    | Toolathlon-Verified      | 74.1      | 70.3        | 55.9      | 49.7        | 59.9      | 76.5      | 76.2      | 77.9          |
    | Agents' Last Exam        | 25.7      | 25.2        | 16.5      | 15.8        | 23.8      | 27.6      | 25.7      | -             |
    | AutomationBench (Public) | 31.8      | 25.1        | 12.8      | 10.8        | 12.9      | 30.8      | 27.2      | 29.1          |
    | DSBench-FullStack        | 71.1      | 68.7        | 41.8      | 37.0        | 61.8      | 73.7      | 71.6      | 77.2          |
    | DSBench-Hard             | 67.2      | 59.6        | 31.1      | 25.8        | 54.5      | 63.0      | 71.7      | 68.3          |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...
parsimo2010•17m ago
The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence...

For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.

- 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.

- 86.6 on Terminal Bench 2.1. Pro 0813 is better.

- 55.9 on NL2Repo. Pro 0813 is better.

- 27 on Agent's Last Exam. Pro 0813 is a little worse.

- 72.5 on Toolathon-Verified. Pro 0813 is better.

- 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.

- 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.

I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.

Gecko4072•45m ago
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
Jsttan•26m ago
What is the new price through?
Gecko4072•23m ago
https://api-docs.deepseek.com/quick_start/pricing/

edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount

minraws•22m ago
isn't it the same old pricing? did they increase V4 Pro pricing already?
nchmy•20m ago
i dont see any price increase there... what am i missing?
alecsm•4m ago
Right below the pricing it is stated that they plan to increase the prices in the near future.
vdfs
book_mike•21m ago
What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
okamiueru•16m ago
How do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.
Readerium•11m ago
V4 Pro has vision correct?
trollbridge•10m ago
No.
alecsm•9m ago
I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.

I can't wait to try this new one.

yipinwong•2m ago
Worse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)
npn•23m ago
I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.
eli•4m ago
Opus 5 medium to Opus 5 max is only 3 points, if that puts it in context
sinuhe69•22m ago
Well, one reason is that we always have to work with the quirks of each model. So, a know model is often preferred over a new/unknown one because we have to be vigilant again. (Negative) surprises are mentally exhausting in the long run. IMO, you can work much better when you know the model.
ianm218•14m ago
I suspect if you follow dev groups in developing countries people are much more focused on token/ price efficiency.

For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale

trollbridge•10m ago
By that standard, the release of Grok 4.6 was also timed on the same day.

Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing.

somenameforme•6m ago
Their leaks would confirm this sort of attitude. They're not trying to become the top player or anything like that - just working to play their part in pushing LLM tech forward and going from there. It was quite refreshing from the 'here's how we're going to dominate the world' nonsense. It's undoubtedly the same attitude that just let's them shrug and cancel the fund raising round after the leaks came from said funding round.
eli•10m ago
Official pricing only kinda matters for an open weight model, no?
parsimo2010•2m ago
It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at.
maherbeg•7m ago
I mean at the rate of model releases happening, I think a lot of these will collide more often than expected!
bel8•16m ago
So it's a Fable class LLM?

                             DSV4Pro vs Fable5
    HLE w tools              60.0 vs 63.0
    Terminal Bench 2.1       87.9 vs 88.0
    Cybergym                 83.3 vs 83.1
    DeepSWE                  62.7 vs 70.0
    Toolathlon-Verified      74.1 vs 77.9
    AutomationBench (Public) 31.8 vs 29.1
    DSBench-FullStack        71.1 vs 77.2
    DSBench-Hard             67.2 vs 68.3
aftbit•14m ago
Fabble lol
eli•13m ago
Fable's guardrails would never let it do something like Cybergym so at least for that one it's measuring Opus 5
wren6991•7m ago
We have a first-party figure from the system card [1]:

> Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).

So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.

[1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...

goldenarm•14m ago
Geometric mean of all these benchmarks :

* GPT-5.6 Sol: 65.5

* Fable 5 (w/ fallback): 64.5

* Opus 5: 64.0

* DS-V4-Pro 0813: 62.5

* Kimi-K3: 62.3

* DS-V4-Flash 0731: 55.8

* GLM-5.2: 47.3

•
2m ago
It's a big confusion, some[0] say an email was sent about significant price increase, personal I haven't seen anything official

[0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...

nolist_policy•24m ago
DeepSeek V4 Flash is the "too cheap to meter" of AI. And you can run the full unquantized model locally for $8000 (2x DGX Spark) at full 1M context and decent speeds: https://github.com/elsung/dgx-spark-deepseek-v4-flash#-long-...
Eueudhsbsj32•21m ago
What's the new pricing?

The prices on OpenRouter still look the same.

igravious•18m ago
yup :)

i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)

how are you doing it?

am also using Kimi K3 via kimi-code

and also GLM 5.2 via ZCode

happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments