frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

3•medicis123•2h ago
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.

WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.

First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.

MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST

DeepSeek V4 Flash C1 1,518.91 21.15 21.15

DeepSeek V4 Flash C4 1,533.15 55.99 14.00

Gemma 4 26B A4B C1 4,579.73 30.22 30.22

Gemma 4 26B A4B C4 4,702.16 63.75 15.94

Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42

Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71

Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary

MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait

DeepSeek V4 Flash 4,154.34 49.30 16s

Gemma 4 26B A4B 18 4,781.44 64.67 6s

Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s

We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/

Ask HN: What changed in your life when you started meditating (and how)?

21•dondraper36•4h ago•12 comments

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

3•medicis123•2h ago•0 comments

Ask HN: What websites do you visit every day?

3•b8•27m ago•5 comments

Ask HN: Why can't Google create monopolies anymore?

3•BeForce•3h ago•2 comments

Build a website for AI pics for dating apps in 20 hours, launched ads, failed

2•MaximTsyg•3h ago•2 comments

Ask HN: US Equivalent of Anabin?

13•xqb64•1w ago•5 comments

AWS: Inaccurate Estimated Billing Data – $1.7 billion

1313•nprateem•5d ago•754 comments

Ask HN: How to get started with a sustainable bootstrapped business of your own?

2•romanovtexas•7h ago•2 comments

Ask HN: Anyone working on practical robotics (cobots) for the home?

3•bredren•7h ago•0 comments

Grok is a surprisingly good automated theorem prover

2•henryrobbins00•7h ago•1 comments

Building an AI-orchestrated publishing workflow for a long-form writing project

2•tmuhlestein•8h ago•0 comments

IMDB now automatically creating user accounts when you are signed into Amazon

6•terminalbraid•9h ago•9 comments

Thanks HN for 15 years of support and helping me find my life's work

832•nicholasjbs•5d ago•107 comments

Ask HN: How do people keep track of organizational knowledge?

26•kadhirvelm•1d ago•22 comments

Ask HN: How to you optimize output while not burning out?

3•gszr•10h ago•4 comments

Ask HN: What Are You Working On? (July 2026)

292•david927•1w ago•1141 comments

Ask HN: Alternative Careers in the Age of AI

9•helpfulmandrill•11h ago•10 comments

Ask HN: Where is NFTs as a technology or investment?

4•Otternonsenz•7h ago•5 comments

You've reached the end!