Reference benchmarks for DwarfStar 4 (ds4), the MIT-licensed C inference engine written by Redis creator Salvatore Sanfilippo, expose a split that matters to anyone specifying a personal AI workstation: of two 128 GB machines, one ingests prompts far faster and the other hands tokens back more than twice as quickly.
Running DeepSeek V4 Flash at q2 and a 32,768-token context, ds4 generates 34.4 tokens per second on an M5 Max Mac with 128 GB of unified memory, against 14.4 on an NVIDIA DGX Spark with the same 128 GB. Prefill inverts the ranking: the Spark ingests 855.9 tokens per second at that context while the Mac manages 557.0. Stretch to 65,536 tokens and the Spark barely moves, at 823.0, while the Mac falls to 398.5.
Key takeaways
- ds4 is a self-contained, MIT-licensed C inference engine from Salvatore Sanfilippo that targets a short list of specific open-weight models rather than acting as a general GGUF runner.
- At a 32,768-token context on DeepSeek V4 Flash q2, the published rows show 34.4 generated tokens per second on a 128 GB M5 Max against 14.4 on a 128 GB DGX Spark, while the Spark leads prefill 855.9 to 557.0.
- Three binaries ship together: a CLI, an OpenAI- and Anthropic-compatible local server, and a coding agent that runs inference in-process rather than over a socket.
What ds4 deliberately is not
The project documentation is blunt that ds4 is not a general GGUF runner. It only loads the GGUF files the project itself produces, and model support is described as opportunistic, with a model dropped when a better replacement arrives. The engine does not link against GGML, though it credits llama.cpp and retains MIT-licensed quantization layouts, CPU dot-product logic and some kernels from it.
Supported families currently span DeepSeek V4 Flash including an experimental vision checkpoint, DeepSeek V4.1 Flash, DeepSeek V4 PRO, GLM 5.2 and 5.3, GLM 5.3 Flash and Qwen3.8 Flash Next. Metal is the primary target on Macs with 96 GB or more. CUDA work centres on the DGX Spark but extends to multi-GPU rigs including Ada Lovelace cards such as the L40S, and ROCm covers Strix Halo systems like the Framework Desktop.
Fitting a large mixture-of-experts model onto one desk rests on asymmetric quantization: compress the routed experts hard, keep the shared critical paths precise. Qwen3.8's smaller Q2 build lands at 41.73 GiB of main and multi-token-prediction weights, which the project calls the entry option for 64 GB Macs.
Why the agent runs inside the engine
The ds4-server binary speaks OpenAI- and Anthropic-style APIs so tools such as Claude Code, Codex CLI and OpenCode can aim a base URL at the local machine. The separate ds4-agent skips that boundary entirely and drives inference in-process, which makes the coding session the on-disk KV cache itself β long prefixes save to SSD and resume by prompt hash instead of being re-ingested after a restart.
That is the expensive half of agent work. DeepSeek has been pushing the same lever from the model side, where V4.1 Flash cut KV cache to 890 bytes per token.
The caveats are stated out loud
Sanfilippo calls the software beta quality, with a large QA pass before each release but regressions still possible. He also discloses that ds4 was built with strong assistance from AI coding agents, with humans leading the ideas, testing and debugging, and says readers unhappy with AI-developed code should look elsewhere. He frames the result less as a finished product than as a working template owners are expected to extend with their own agents.
The distributed figures temper any instinct to chain machines. Across two M5 Max Macs over Thunderbolt 5 on Flash Q4, prefill scales but generation does not, staying autoregressive and paying roughly a 19 percent penalty against a single box. A denser eight-L40S setup is reported at about 126 tokens per second aggregate generation across 16 concurrent sessions β a multi-user server profile, not a faster single chat.
The numbers come from ds4-bench, which walks a fixed public-domain text to each context frontier in 2,048-token steps and probes 128 greedy tokens, so every row is throughput at that context size rather than a run average. They are published in the project's performance guide and mirrored on the community DwarfStar benchmark table, which is now asking ROCm and generic CUDA owners to submit their own runs.
FAQ
Is ds4 open source?
Yes. ds4 is released under the MIT license, the same license covering the quantization tables and kernels it adapts from llama.cpp. The source sits at github.com/antirez/ds4, and the repository also ships the project's GGUF, imatrix, quality and speed tooling.
Can ds4 run any GGUF model file?
No. The engine is narrow by design and runs only the GGUF files the project builds for its supported families. Generic GGUF files from elsewhere are explicitly not the target, which is the trade it makes to validate each supported layout end to end.
How much memory does DeepSeek V4 Flash need under ds4?
Metal is aimed at Macs with 96 GB or more, and the reference rows were measured on 128 GB machines. Smaller systems can fall back on SSD streaming, and the 41.73 GiB Qwen3.8 Q2 build is positioned as the starting point for 64 GB Macs.






