frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

2•medicis123•54m ago
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.

WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.

First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.

MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST

DeepSeek V4 Flash C1 1,518.91 21.15 21.15

DeepSeek V4 Flash C4 1,533.15 55.99 14.00

Gemma 4 26B A4B C1 4,579.73 30.22 30.22

Gemma 4 26B A4B C4 4,702.16 63.75 15.94

Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42

Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71

Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary

MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait

DeepSeek V4 Flash 4,154.34 49.30 16s

Gemma 4 26B A4B 18 4,781.44 64.67 6s

Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s

We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/

Ravenspire – watch your Claude Code and Codex agents as a JRPG

https://github.com/Kapala-Solutions/ravenspire
1•vribeirocmp•47s ago•0 comments

Show HN: 1,280 open USDZ furniture assets for VR/AR

https://lx2026.github.io/genhome3d-1280/
1•linxy97•1m ago•0 comments

Tesla's profits slide despite growing revenue as it pivots to robotics and AI

https://www.theguardian.com/technology/2026/jul/22/tesla-profits-earnings
3•jethronethro•3m ago•0 comments

Anthropomorphism in Children's Interactions with LLM Chatbots

https://arxiv.org/abs/2607.18250
1•StatsAreFun•4m ago•0 comments

I built ClickIt a local clipboard manager for macOS

https://github.com/rohankc69/clickit
1•rohankc_01•8m ago•0 comments

The Human Kintsugi

https://0xff.nu/human-kintsugi/
1•hxii•11m ago•0 comments

RefluXFS: A Linux Kernel Local Privilege Escalation to Root in XFS

https://blog.qualys.com/vulnerabilities-threat-research/2026/07/22/refluxfs-a-linux-kernel-local-...
2•garyhtou•13m ago•0 comments

Updates on Chinese AI: Kimi-K3, Xi at WAIC, and 4 Months to Mythos

https://aiunderheaven.substack.com/p/ai-under-heaven01
1•taiwandongsuan•18m ago•0 comments

Show HN: I ran 12 AI bots predicting stocks for two months, every call public

https://ldbd.app
1•kkjh0723•19m ago•0 comments

Training Agent Harness Like Training a ML Model

https://www.henrypan.com/blog/2026-07-18-harness-training/
1•megadragon9•20m ago•0 comments

Type-Aware Linting Stable

https://oxc.rs/blog/2026-07-22-type-aware-linting-stable
2•luispa•20m ago•0 comments

Antares from Cisco: Highly Efficient Open Models for Vulnerability Localization

https://blogs.cisco.com/ai/introducing-antares-the-most-efficient-open-weight-ai-models-for-vulne...
1•SwellJoe•21m ago•0 comments

Angela Merkel's Light Bulb Went Missing

https://substack.com/@michelelynnjakubowski/note/c-299849449
1•Mjakubowski68•21m ago•0 comments

Encyclopedia of Things Considered Harmful

https://harmful.cat-v.org/
1•valyala•25m ago•0 comments

The Online Safety Act is supposed to make Britain's internet a safe space

https://www.thenerve.news/p/adele-walton-column-online-safety-act-ofcom-suicide-kenneth-law
1•horatioduke•30m ago•0 comments

Sanctions and Entity List designations are on the table for Chinese AI models

https://twitter.com/SecScottBessent/status/2080008411790368895
1•MaKey•32m ago•0 comments

Show HN: ValuePair – a friendship app that cares about values first

https://valuepair.app
2•zloy88•34m ago•1 comments

Monday.com lays off hundreds to focus on AI

https://techcrunch.com/2026/07/22/monday-com-lays-off-hundreds-to-focuses-on-ai/
2•duck•36m ago•1 comments

Quadraphonic Sound

https://en.wikipedia.org/wiki/Quadraphonic_sound
2•mindcrime•39m ago•0 comments

Fretboard Memorisation with Modular Arithmetic

https://ohaodha.ie/blog/fretboard-memorisation-with-modular-arithmetic/
2•ohaodha•44m ago•0 comments

Open Version of MCP Lists

https://github.com/rizzdev/awesome-mcp-open
1•rizzdev•47m ago•0 comments

Reindeer eyes seasonally adapt to ozone-blue Arctic twilight

https://royalsocietypublishing.org/rspb/article/289/1977/20221002/86488/Reindeer-eyes-seasonally-...
3•MaysonL•47m ago•0 comments

AI agent went rogue and hacked startup by itself, OpenAI reveals

https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-s...
1•gcanyon•47m ago•1 comments

Solar Open 2: Korea's Sovereign Foundation Model, Built for Agentic Use

https://www.upstage.ai/blog/en/solar-open-2
3•ilreb•51m ago•0 comments

Show HN: Promptrack A local menu bar app that tracks your Claude Code usage

https://promptrack.dev/
1•feskk•53m ago•0 comments

The Future of SaaS Is Cloneable

https://www.builder.io/blog/the-future-of-saas-is-cloneable
1•gbourne•53m ago•0 comments

'Rust makes coding fun again': Why Linux is moving away from C, says Greg KH

https://www.zdnet.com/article/greg-kroah-hartman-linux-kernel-rust/
10•arto•54m ago•0 comments

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

2•medicis123•54m ago•0 comments

Charles Ross spent 50 yrs building Star Axis naked-eye observatory in New Mexico

https://www.nytimes.com/2026/07/22/arts/design/charles-ross-star-axis-land-art.html
2•ChrisArchitect•55m ago•2 comments

Bing Copilot (ChatGPT-4) Flunks Math [pdf] (2024)

https://www.cs.dartmouth.edu/~doug/ChatMath.pdf
2•ryandotsmith•56m ago•0 comments