WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.
First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.
MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST
DeepSeek V4 Flash C1 1,518.91 21.15 21.15
DeepSeek V4 Flash C4 1,533.15 55.99 14.00
Gemma 4 26B A4B C1 4,579.73 30.22 30.22
Gemma 4 26B A4B C4 4,702.16 63.75 15.94
Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42
Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71
Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary
MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait
DeepSeek V4 Flash 4,154.34 49.30 16s
Gemma 4 26B A4B 18 4,781.44 64.67 6s
Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s
We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/