frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

DeepSeek V4 Flash across 14 providers: cost, speed and caching

https://www.inference.academy/benchmarks/serving-tradeoffs
3•batuhanaktas61•55m ago

Comments

batuhanaktas61•55m ago
Inference is all about tradeoffs. Therefore choosing an inference provider means choosing a set of tradeoffs which matters a lot for every use case.

Providers serving the same model can differ in first-token latency, generation speed, caching, cost and reliability. I wanted to understand which differences matter for which workloads. So, I created an open source repo to run experiments for model-provider pairs with ease on either existing benchmarks or with my own traces. https://github.com/aktasbatuhan/compound

I’m also interested in making those choices easier to revisit. With closed models, choosing the intelligence often means choosing the vendor. Open models give us more freedom to choose who serves them, but that freedom becomes useful only when we can measure the alternatives.

I tested DeepSeek V4 Flash across 14 serving routes: 12,795 measured calls, input sizes from 1k to 100k tokens, two output budgets, and cold and warm prompts. Temperature was zero and reasoning was off.

My favorite outcome was how much caching changed across providers. On DeepSeek’s own route, a repeated 100k prompt with a 100-token output budget cost about 15 times less than the cold condition. The cheapest option for fresh prompts wasn’t necessarily the cheapest for repeated ones.

Automatic routing adds another variable. OpenRouter can send requests to different upstreams, while cached prefixes depend on where those requests land. Auto routing still achieved substantial cache hits in this experiment, but its warm-request costs were higher than the cheapest pinned alternatives.

The report compares these tradeoffs with interactive charts and downloadable measurements. I built the experiment with Compound so others can run similar comparisons on their own prompts. The useful question is which provider meets your workload’s latency, cost and reliability requirements.

"Intelligence curse" means promise of abundance and UBI unlikely to be fulfilled

https://intelligence-curse.ai/defining/
1•thoughtpeddler•5m ago•0 comments

Known registered domain count for every TLD

https://namedesk.app/stats/tlds
1•zeppelin_7•8m ago•0 comments

Australian government to chart course away from 'American war machine'

https://www.theguardian.com/australia-news/2026/sep/08/grassroots-labor-pressure-group-albanese-g...
1•Alien1Being•10m ago•0 comments

Train Jazz

https://www.aico.nyc/amtrak
1•bookofjoe•15m ago•0 comments

Actually Real AI

https://actually-real-ai.com/
1•measurablefunc•15m ago•0 comments

DoodleIQ – Rent your idle local-LLM machine to others by the second

https://www.doodleiq.com/
1•rahulbats•16m ago•0 comments

Speeding up a Phoenix LiveView web app with a CDN

https://jola.dev/posts/speeding-up-phoenix-liveview-cdn
1•auraham•16m ago•0 comments

Michigan Scientist Buried Seeds in Glass Bottles in 1879, 142 Yr Later Some Grew

https://timesofindia.indiatimes.com/science/a-michigan-scientist-buried-seeds-in-20-glass-bottles...
1•rmason•17m ago•0 comments

Cool Web Tool – I got tired of paying for multiple SEO tools

https://coolwebtool.com/
1•umit-ozturk•20m ago•0 comments

The Education of a Doomer

https://borretti.me/article/the-education-of-a-doomer
2•f311a•22m ago•0 comments

x86 Evolution for Segmentation and Paging

https://inbox.sourceware.org/binutils/CAKSQd8XJVf3jq8boOJcOE_GzcNvcqMUHeW-whbw3uGnVE2iN_w@mail.gm...
1•st_goliath•22m ago•0 comments

Working on Economics with Fable 5

https://wilsoniumite.com/2026/08/03/working-on-economics-with-fable-5/
2•Wilsoniumite•23m ago•0 comments

One Thing Worth Thinging

2•arkensaw•26m ago•1 comments

Please, take 2 min to tell your representative to regulate AI (US)

https://actonai.org/
1•doitLP•27m ago•0 comments

I built a workspace beside Safari instead of another browser

https://tiledbrowser.app/articles/why-i-built-a-workspace-beside-safari/
1•tiledbrowser•29m ago•0 comments

"code masturbation": getting pleasure out of making something only for yourself

https://www.youtube.com/watch?v=pVMM23kUVH8
1•thebooglebooski•30m ago•1 comments

The smallest edge AI device for local LLMs

https://tiiny.ai/
4•alex-moon•32m ago•0 comments

Own Your Means of Production

https://mcottondesign.com/post/own-your-means-of-production/
2•mcotton•33m ago•0 comments

Show HN: Intercontinental Ballistic Wikipedia Edit Map

https://defamation.dev/icbe/
4•UrineSqueegee•38m ago•0 comments

The purpose of DNS is to spread scams

https://shkspr.mobi/blog/2026/09/the-purpose-of-dns-is-to-spread-scams/
3•pavel_lishin•41m ago•1 comments

Show HN: Rejourney – Finds user "leaks" in apps and websites

https://rejourney.co/demo/leaks
1•mrr7337•45m ago•0 comments

Pixel is a unit of length and area

https://www.nayuki.io/page/pixel-is-a-unit-of-length-and-area
2•speerer•48m ago•2 comments

Beta User for Proven Trips

https://proventrips.com
3•rohitnrl•49m ago•1 comments

Show HN: DNTLS – Decentralized Trust Layer for the Open Internet

https://dntls.net/
2•aliasxneo•50m ago•0 comments

Monthly Updates [Sept]

https://fmhy.net/posts/sept-2026
2•Cider9986•51m ago•0 comments

Girder – an MCP server that gives coding agents a code graph

https://github.com/dhishwasher/Girder
1•dhishwasher•52m ago•0 comments

Ollama replacement 2-4x faster for no extra compute cost

https://github.com/omgitsbase/llmash
1•wowitsbase•55m ago•4 comments

Wabi-sabi

https://en.wikipedia.org/wiki/Wabi-sabi
2•sdeframond•55m ago•0 comments

DeepSeek V4 Flash across 14 providers: cost, speed and caching

https://www.inference.academy/benchmarks/serving-tradeoffs
3•batuhanaktas61•55m ago•1 comments

From Spring Boot to Swift: The Business Logic Was the Cheap Part

https://leosolbach.medium.com/from-spring-boot-to-swift-the-business-logic-was-the-cheap-part-206...
1•frizlab•55m ago•1 comments