AI Newsway

antirez's ds4 Runs DeepSeek V4 Flash at 34.4 Tokens/s on an M5 Max, 14.4 on a DGX Spark

The MIT-licensed C engine loads only its own GGUF builds β€” and the reference table inverts the hardware ranking depending on whether you count prefill or generation.

|5 min read0
AI Summary
DwarfStar 4 (ds4) is an MIT-licensed C inference engine from Redis creator Salvatore Sanfilippo for running frontier open-weight models locally. Reference rows put DeepSeek V4 Flash q2 at 34.4 generated tokens per second on a 128 GB M5 Max and 14.4 on a DGX Spark at a 32,768-token context, with the Spark leading prefill 855.9 to 557.0. The engine loads only its own GGUF builds and ships a CLI, a local API server and an in-process coding agent.
Apple's Mac Studio: high-memory Apple Silicon desktops are ds4's primary Metal target for running DeepSeek V4 Flash locally.
Apple's Mac Studio: high-memory Apple Silicon desktops are ds4's primary Metal target for running DeepSeek V4 Flash locally.

Reference benchmarks for DwarfStar 4 (ds4), the MIT-licensed C inference engine written by Redis creator Salvatore Sanfilippo, expose a split that matters to anyone specifying a personal AI workstation: of two 128 GB machines, one ingests prompts far faster and the other hands tokens back more than twice as quickly.

Running DeepSeek V4 Flash at q2 and a 32,768-token context, ds4 generates 34.4 tokens per second on an M5 Max Mac with 128 GB of unified memory, against 14.4 on an NVIDIA DGX Spark with the same 128 GB. Prefill inverts the ranking: the Spark ingests 855.9 tokens per second at that context while the Mac manages 557.0. Stretch to 65,536 tokens and the Spark barely moves, at 823.0, while the Mac falls to 398.5.

Key takeaways

  • ds4 is a self-contained, MIT-licensed C inference engine from Salvatore Sanfilippo that targets a short list of specific open-weight models rather than acting as a general GGUF runner.
  • At a 32,768-token context on DeepSeek V4 Flash q2, the published rows show 34.4 generated tokens per second on a 128 GB M5 Max against 14.4 on a 128 GB DGX Spark, while the Spark leads prefill 855.9 to 557.0.
  • Three binaries ship together: a CLI, an OpenAI- and Anthropic-compatible local server, and a coding agent that runs inference in-process rather than over a socket.

What ds4 deliberately is not

The project documentation is blunt that ds4 is not a general GGUF runner. It only loads the GGUF files the project itself produces, and model support is described as opportunistic, with a model dropped when a better replacement arrives. The engine does not link against GGML, though it credits llama.cpp and retains MIT-licensed quantization layouts, CPU dot-product logic and some kernels from it.

Supported families currently span DeepSeek V4 Flash including an experimental vision checkpoint, DeepSeek V4.1 Flash, DeepSeek V4 PRO, GLM 5.2 and 5.3, GLM 5.3 Flash and Qwen3.8 Flash Next. Metal is the primary target on Macs with 96 GB or more. CUDA work centres on the DGX Spark but extends to multi-GPU rigs including Ada Lovelace cards such as the L40S, and ROCm covers Strix Halo systems like the Framework Desktop.

Fitting a large mixture-of-experts model onto one desk rests on asymmetric quantization: compress the routed experts hard, keep the shared critical paths precise. Qwen3.8's smaller Q2 build lands at 41.73 GiB of main and multi-token-prediction weights, which the project calls the entry option for 64 GB Macs.

Why the agent runs inside the engine

The ds4-server binary speaks OpenAI- and Anthropic-style APIs so tools such as Claude Code, Codex CLI and OpenCode can aim a base URL at the local machine. The separate ds4-agent skips that boundary entirely and drives inference in-process, which makes the coding session the on-disk KV cache itself β€” long prefixes save to SSD and resume by prompt hash instead of being re-ingested after a restart.

That is the expensive half of agent work. DeepSeek has been pushing the same lever from the model side, where V4.1 Flash cut KV cache to 890 bytes per token.

The caveats are stated out loud

Sanfilippo calls the software beta quality, with a large QA pass before each release but regressions still possible. He also discloses that ds4 was built with strong assistance from AI coding agents, with humans leading the ideas, testing and debugging, and says readers unhappy with AI-developed code should look elsewhere. He frames the result less as a finished product than as a working template owners are expected to extend with their own agents.

The distributed figures temper any instinct to chain machines. Across two M5 Max Macs over Thunderbolt 5 on Flash Q4, prefill scales but generation does not, staying autoregressive and paying roughly a 19 percent penalty against a single box. A denser eight-L40S setup is reported at about 126 tokens per second aggregate generation across 16 concurrent sessions β€” a multi-user server profile, not a faster single chat.

The numbers come from ds4-bench, which walks a fixed public-domain text to each context frontier in 2,048-token steps and probes 128 greedy tokens, so every row is throughput at that context size rather than a run average. They are published in the project's performance guide and mirrored on the community DwarfStar benchmark table, which is now asking ROCm and generic CUDA owners to submit their own runs.

FAQ

Is ds4 open source?

Yes. ds4 is released under the MIT license, the same license covering the quantization tables and kernels it adapts from llama.cpp. The source sits at github.com/antirez/ds4, and the repository also ships the project's GGUF, imatrix, quality and speed tooling.

Can ds4 run any GGUF model file?

No. The engine is narrow by design and runs only the GGUF files the project builds for its supported families. Generic GGUF files from elsewhere are explicitly not the target, which is the trade it makes to validate each supported layout end to end.

How much memory does DeepSeek V4 Flash need under ds4?

Metal is aimed at Macs with 96 GB or more, and the reference rows were measured on 128 GB machines. Smaller systems can fall back on SSD streaming, and the 41.73 GiB Qwen3.8 Q2 build is positioned as the starting point for 64 GB Macs.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

A Lab With 'a Couple of GPUs' Says a 125B Model Now Fits on One 24GB Card
Developer Tools

A Lab With 'a Couple of GPUs' Says a 125B Model Now Fits on One 24GB Card

CMU's Tim Dettmers says his lab's framework runs a 125B model on a single 24GB GPU and a 550B model on a 128GB MacBook, ahead of a six-release open-source week.

Seung Jung11 days ago
Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5
Developer Tools

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5

Claude Opus 5.5 is 20% cheaper than Opus 5 and returns HTTP 400 on four request patterns Opus 5 accepted. A fifth change silences agent progress streams.

Seung Jung5 days ago
Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour
Developer Tools

Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour

Docker's new Cloud Sandboxes put its microVM agent isolation on hosted compute, billed by the second from $0.07 an hour, and hand the Kits format to the CNCF.

Seung Jung8 days ago
Survey of 305 Developers Finds 71% Shipped AI Code They Didn't Understand
Developer Tools

Survey of 305 Developers Finds 71% Shipped AI Code They Didn't Understand

Coddy surveyed 305 weekly AI coding tool users. Four in five called it dependence, not advantage, and after-hours coding ranged 36% to 62% by tool.

Seung Jung13 days ago
Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js
Developer Tools

Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js

Vercel traced a reported Next.js RCE to libheif, disabled AVIF platform-wide on August 13, and coordinated fixes across sharp, libvips and libheif by August 25.

Seung Jung14 days ago
Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts
Developer Tools

Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts

OpenAI shipped two mid-tier models on Tuesday and made the pitch almost entirely about money. GPT-6 Sol lists at $2 per million input tokens and $10 per million...

Seung Jung10 days ago