# ds4 dwarfstar [2026 Review]: Redis-Style Local LLM Loops

> ds4 (DwarfStar 4) turns local LLM inference into a Redis-like workflow: one long-lived server, deterministic prompt-keyed KV caching, and fast edit-run-edit loops without a GUI circus.

- Canonical: https://www.kunalganglani.com/blog/ds4-dwarfstar-review
- Author: Kunal Ganglani
- Published: 2026-10-03 · Updated: 2026-10-03
- Category: Developer Tools · Tags: local-llm, inference, open-source, llm-serving, developer-workflow

## TL;DR

ds4 (DwarfStar 4) is a local AI tool that runs certain large language models on your own Mac or GPU box and exposes them through a command line, a local server, and a built-in coding agent. It matters because it treats “warm context” like something you can save and reuse, so repeated prompts don’t feel slow every time you restart. The big idea is a disk cache for prompt prefixes, keyed by a hash, which can make day-to-day coding loops much snappier. If you want a headless, developer-first workflow, ds4 is worth trying. If you want maximum model compatibility, pick a general runner instead.

ds4 (DwarfStar 4) is a deliberately narrow local LLM inference stack that runs specific “project GGUF” models on Metal, CUDA, or ROCm, and exposes the same loaded model through a CLI, an HTTP server, and a native coding agent. The consequence is simple: your model loads once, your prefixes can stay warm across restarts via a disk KV cache, and your local dev loop stops feeling like a science project. This is a **ds4 dwarfstar** review from the perspective that matters: daily loops, not demo screenshots.

Most local runners optimize for “get *a* model talking.” ds4 optimizes for “get *this* workflow tight.” That’s a very antirez move. If you’ve ever loved Redis because it’s boringly predictable under load, ds4 will feel familiar.

Here’s the mental model I keep coming back to. **ds4 treats inference state like a cacheable, persistent asset**, not a disposable byproduct of one prompt.

> If your local model has to cold-start its context on every run, you don’t have a dev loop. You have a latency tax.

## Local frontier inference, narrow on purpose

The first thing to internalize is that ds4 is not trying to be a universal `llama.cpp` replacement. It’s not aiming to run “whatever GGUF I found on Hugging Face at 2am.” It’s aiming to make a small set of models feel *great*.

![Two nvidia titan x graphics cards side by side](https://cdn.sanity.io/images/vzekdneq/production/b6e62bfa053ea230e29467b971c39f8af28e7a26-1200x675.webp)

As [Salvatore Sanfilippo (antirez)](https://github.com/antirez/ds4) explains in the README, ds4 targets a small set of model families and expects **project-produced GGUF files** that are validated end-to-end. Translation: you’re trading breadth for a stack where the engine, quantization choices, and model layouts are all tested together.

I’m a fan of this trade. The local LLM world has too much “works on my machine” energy. People treat quant files like Pokémon cards, swap them constantly, then act shocked when things crash or degrade.

ds4 can make strong claims like “KV cache as a disk citizen” because it’s not trying to be everything to everyone. It’s building a coherent system.

The official site calls out three design choices that are basically the spine of the project:

- **Asymmetric 2-bit quantization** aimed at routed experts (for MoE models)
- **Disk-persisted KV cache** keyed by a hash of the rendered prompt prefix
- **One engine, three interfaces** (`./ds4`, `./ds4-server`, `./ds4-agent`) sharing state
Those aren’t a random grab bag of features. They’re one opinion, expressed three ways.

If you want something that runs every GGUF under the sun, use the thing built for that job. If you want a Redis-style workflow where the hot path is engineered, ds4 is worth your time.

## How the ds4 stack fits together (and why it feels like Redis)

The cleanest way to understand ds4 is as a single model process with multiple front doors:

![Nvidia logo on a green background with abstract spheres](https://cdn.sanity.io/images/vzekdneq/production/9132c5a51531886a26cf2e63c555f3bad87e1d7c-1200x675.webp)

- `ds4` is the interactive CLI.
- `ds4-server` is the HTTP API server (OpenAI + Anthropic style, per the official docs).
- `ds4-agent` is a native coding agent that keeps sessions around.
All three are meant to share:

1. the same model weights (loaded once)
1. the same runtime state
1. the same KV cache
This is where the “Redis-style workflow” analogy stops being cute and starts being useful.

Redis isn’t popular because it has a nice GUI. It’s popular because it behaves like infrastructure:

- run a **long-lived daemon**
- keep hot state **in memory**
- persist the right artifacts to disk
- hit it from any client that can speak the protocol
ds4 is pushing local inference in that direction. Less “open an app and chat.” More “spin up a service and build around it.”

### The KV cache: the real feature

DwarfStar’s landing page spells it out: “The KV cache is keyed by the SHA1 of the rendered prompt prefix and persisted to disk, so a matching prefix is reloaded instead of recomputed.” That’s the whole point.

In practice, if your coding agent has a stable system prompt and a stable “repo context preamble,” ds4 can avoid paying the prefill cost again after restarts.

That changes how you work.

- You restart less often because you’re not terrified of losing warm context.
- You start designing prompts with stable prefixes on purpose.
- You start caring about cache hit rate the way you care about Redis hit rate.
If you’ve been building [AI agents](/pillars/ai-agents) or doing agent orchestration, you already know the pain. Agent loops are only as good as their latency budget.

And if you’re doing local dev to avoid [LLM cost](/blog/ai-agent-cost-per-task-2026), persistent caching isn’t a “nice to have.” It’s the difference between “I’ll use this daily” and “cool demo, never again.”

## Supported hardware (local, streamed, and distributed)

ds4’s hardware story lives in three worlds: Metal, CUDA, and ROCm.

![Two nvidia titan x graphics cards side by side](https://cdn.sanity.io/images/vzekdneq/production/b6e62bfa053ea230e29467b971c39f8af28e7a26-1200x675.webp)

Per the ds4 README by [Salvatore Sanfilippo (antirez)](https://github.com/antirez/ds4), the primary target is **high-memory Macs**. The docs call out **96GB+ Apple Silicon** as the comfortable tier for the biggest supported models, and also mention SSD streaming as an option for “smaller RAM / very large models.”

On the other side:

- NVIDIA CUDA is supported, including multi-GPU serving.
- ROCm is supported with a focus on newer platforms like Strix Halo systems.
If you’re trying to decide between platforms, start with the boring truth I keep repeating in my benchmark work: memory dictates feasibility, throughput dictates happiness.

Based on the benchmark data I maintain at **kunalganglani.com/llm-benchmarks**, unified memory on [Apple Silicon](/blog/apple-silicon-vs-nvidia-for-ai) often makes “will it load?” easy, but “will it feel fast?” is still the real question. Local inference usually fails not because it can’t run, but because it runs at a pace that breaks your flow.

If you’re building a shared box for a team, ds4’s CUDA and multi-user angle matters more than the Mac story. A fast headless server that people can hit over the network beats “everyone runs a laptop stack” in most real orgs.

## Run ds4 in three steps (day 1 to daily driver)

I’m going to keep this practical and reproducible. Not “click this download button” practical. “Can you put it in a `Makefile` and forget about it?” practical.

### Step 1: install ds4 (macOS Metal / Linux CUDA / ROCm)

ds4 is a C project with multiple backends. The canonical source of truth is still the repo README.

Start here:

- [Salvatore Sanfilippo (antirez)](https://github.com/antirez/ds4) for building and usage
- The official site [DwarfStar 4](https://dwarfstar.sh/) for the architecture-level view
If you want a friendlier “tooling wrapper,” the community Go project is worth a look:

- NimbleMarkets ds4go
ds4go is interesting because it treats ds4 as a reusable library via FFI and puts real effort into model download ergonomics.

### Step 2: download/select a supported model (and know where it lives)

ds4’s narrow stance shows up immediately in model management. You’re not pulling arbitrary community quantizations and hoping they behave.

Supported model families called out across the official site and README include:

- DeepSeek V4 Flash and V4.1 Flash
- GLM 5.x (including Flash variants)
- DeepSeek V4 PRO
- Qwen3.8 Flash Next
Support can vary by backend, so don’t assume “it runs on Metal” implies “it runs on ROCm.”

On disk, treat ds4 models the way you treat Docker images or Bazel caches: pin them, version them, and share them intentionally. If you’re doing this as a team, don’t hand-wave it.

A pattern that works:

- Put a `models/` directory on a fast SSD volume.
- Store model version + quant variant in the path.
- Commit a tiny “model manifest” file to your repo.
It’s a small bit of discipline that buys you fewer surprises, faster onboarding, and less “wait, which file are you using?” slack drama.

If you want the broader “how quants behave and where they cliff,” my quant deep-dive is here: [LLM quantization](/blog/llm-quantization-levels-q4-q8-fp16).

### Step 3: run `ds4-server` and connect an OpenAI-compatible client

The official site explicitly positions `ds4-server` as “OpenAI + Anthropic API.” That matters because it means you can often drop ds4 under existing clients without rewriting your whole toolchain.

I’m not going to pretend every client works perfectly. OpenAI-compat is a spectrum. But it covers the 80% case: curl tests, SDKs, and a lot of local tooling.

A minimal sanity check looks like this:

- start `ds4-server` bound to localhost only
- hit a `/v1/...` endpoint with a small prompt
- confirm streaming behavior (if your client needs it)
Once that works, wire it into whatever you actually use.

If your day-to-day flow involves [Claude Code](/blog/claude-code-security-2026) or other agentic coding tools, the question is always the same: can the tool point at an OpenAI-style base URL, and does it behave under retries when the model is busy?

## Everyday use: caching, agent loops, and what actually gets faster

The place ds4 feels different from Ollama or LM Studio is not “tokens per second on a clean benchmark.” It’s what happens after your fifth iteration on the same task.

That’s the reality of coding workflows. You don’t ask once. You ask, tweak, re-run, paste an error, re-run, then realize your prompt template changed and everything got slower again.

### ds4 disk KV cache: what it caches and how to maximize hits

ds4 caches **prefix KV**. The official description is precise: the cache key is the **SHA1 of the rendered prompt prefix**.

That implies three rules that are painfully easy to ignore:

1. **Render deterministically.** If your template injects timestamps, random IDs, or unstable tool metadata above the fold, you’re nuking your own cache.
1. **Stabilize the preamble.** Keep your system prompt, repo summary, and “rules” block identical across runs.
1. **Move volatile context later.** Put “recent diffs,” error logs, and one-off notes after the stable prefix boundary.
If you’ve built RAG systems, this will sound familiar. You’re basically optimizing for a cache-friendly header. Same muscle group.

If you’re doing [RAG](/blog/rag-evaluation-metrics-retrieval-quality) or [retrieval-augmented generation](/blog/fine-tuning-vs-rag-prompt-engineering), you already know the drill. Separate stable context from volatile context.

### Disk growth, invalidation, and SSD reality

A disk KV cache isn’t free.

- It will grow.
- It will create real write activity.
- It will feel terrible on a slow drive.
The upside is still worth it. Just don’t be casual about it. In Redis terms: don’t put your persistence files on a bargain-bin disk and then complain about latency.

If you’re doing this on a laptop, look at your cache directory size weekly, not once a quarter when you’re out of space mid-flight.

### Native coding agent vs “bring your own agent framework”

ds4 ships `ds4-agent`, a native coding agent concept.

Whether you use it depends on how invested you are in your existing stack. If you’re already deep into [agentic AI](/blog/rise-of-agentic-ai) frameworks or you’ve built your own loop, you’ll probably just want `ds4-server`.

If you’re earlier, a native agent that shares model state and cache can be a solid “one stack” experience.

This is the ds4 vibe in one sentence: stop wiring four different tools together for something that should be one process.

## Capability evaluation, speed, and how ds4 compares to Ollama/LM Studio

Let’s talk performance without turning this into benchmark cosplay.

ds4’s README includes a concrete serving example: an **8× NVIDIA L40S** box reaching **~126 tokens/sec aggregate generation** with **16 simultaneous sessions**. That’s straight from [Salvatore Sanfilippo (antirez)](https://github.com/antirez/ds4).

That number is interesting for two reasons:

1. It’s explicitly **aggregate** under concurrency. That’s what real workloads look like.
1. It’s a signal that ds4 cares about multi-session serving, not just single-user desktop chatting.
### Prefill vs decode: what matters for dev loops

For coding workflows, you usually care about **time-to-first-token (TTFT)** and **prefill cost** as much as raw decode throughput.

- Prefill is what hurts when you keep re-sending the same long context.
- Decode throughput matters once you’re already streaming.
This is why ds4’s disk KV cache is such a targeted bet. It’s not trying to be a generic “semantic cache.” It’s saying: stop paying prefill repeatedly for the same prefix.

If you want a deeper mental model on the metrics, my benchmark methodology post is here: [local LLM](/blog/local-llm-benchmark-methodology).

### ds4 vs Ollama vs LM Studio vs llama.cpp (comparison table)

Here’s the comparison I wish more reviews would write. Not “which is best,” but “which breaks your loop less.”

| Dimension | ds4 (DwarfStar 4) | Ollama | LM Studio | llama.cpp |
| --- | --- | --- | --- | --- |
| Scope | Narrow, validated model families + project GGUFs | Broad model catalog, friendly UX | Broad GGUF UX with strong GUI features | Broad engine, DIY integrations |
| Workflow | One engine with CLI/server/agent sharing state | CLI + background service | GUI-first, can serve on LAN | Library/CLI building block |
| Cache across restart | **Yes**: disk KV cache keyed by SHA1 prefix | Usually no persistent prefix KV | Usually no persistent prefix KV | Possible with work, not default |
| Best fit | Tight local dev loops + multi-user serving + specific frontier open weights | “Just run a model” local dev and tooling ecosystem | Power users who want GUI + profiles + LAN serving | Builders who want maximum flexibility |
| When it loses | Unsupported model families, non-project GGUFs | More latency tax on repeated long prefixes | GUI overhead if you want headless-first | More assembly required |

If you’re an LM Studio power user, I wrote my setup notes here: [local LLM](/blog/lm-studio-power-user-setup-2026).

If you’re choosing between Ollama and LM Studio in general, this is the sharper comparison: [local LLM](/blog/ollama-vs-lm-studio-2026).

## Ops guidance: running ds4-server for a team (and basic security)

If you’re going to run ds4-server as a shared internal service, treat it like any other internal service. No hero setups. No “it’s just on the LAN.”

### Run it like a daemon, not a terminal tab

- Use `systemd` on Linux or `launchd` on macOS.
- Write logs to a file and rotate them.
- Pin the model version so “someone updated it” doesn’t become your next incident.
You’ll also want quotas or concurrency limits. ds4 is designed for sessions. Your GPU is not designed for infinite enthusiasm.

If you’re thinking “this is starting to sound like production,” yes. That’s why I treat local inference as part of [production AI](/blog/vllm-production-checklist).

### Bind addresses and authentication

Do not bind an LLM server to `0.0.0.0` on your home network and call it a day. Even on a “trusted” LAN, you’re exposing a service that can be abused.

At minimum:

- bind to `127.0.0.1` by default
- put an auth proxy in front if you need remote access
- restrict egress if you’re mixing local and remote tools
If you want the bigger checklist, I wrote it here: [AI security](/blog/secure-local-llm-inference).

And if you’re using agentic tools that read your repo, internal docs, or tickets, you also need to care about [prompt injection](/blog/prompt-injection-2026-owasp-llm-vulnerability) and broader [AI security](/blog/ai-security-complete-guide).

## Status, upgrades, and when not to use ds4

### Status: fast-moving by design

ds4 is actively evolving. That’s part of its appeal and part of the tax.

As of October 2026, the public narrative is clear:

- model support is expanding (Qwen/vision mentions show up in community chatter)
- APIs are positioned as OpenAI + Anthropic compatible
- multi-user / multi-GPU serving is a first-class target
If you adopt ds4, decide up front: are you pinning versions for stability, or tracking main for features? “We’ll figure it out later” is how you end up with a broken toolchain on a Monday morning.

### How to update ds4 and models safely

My boring recommendation:

- Treat ds4 like a compiler. Upgrade intentionally.
- Keep a known-good binary around.
- Pin model artifacts by version and checksum.
If you want a supply chain mindset for model files, this is worth reading: [LLM security](/blog/verify-gguf-hashes-supply-chain).

### When you should not use ds4

Use ds4 when you agree with its premise. Don’t use it when you want it to be something it’s not.

Skip ds4 if:

- you need broad support for random GGUFs from the internet
- you’re experimenting with lots of niche model families
- you want a GUI-first experience
- you can’t give it fast local SSD for cache + weights
In those cases:

- Ollama is a great default “just run it” tool.
- LM Studio is excellent if you want profiles, quants, and a power-user GUI.
- `llama.cpp` is still the universal building block.
I know “use the right tool” is obvious advice. But teams still waste weeks trying to force a tool to match constraints it was never built for.

The interesting implication of ds4 isn’t that it replaces everything. It’s that it points to the next phase of local inference: **stateful, cache-aware, multi-interface runtimes** that behave more like infrastructure and less like a toy app.

My bet: within a year, “persistent KV cache across restarts” will be table stakes for serious local dev loops. If you’re building developer tooling, treat ds4’s cache design as the baseline. The latency tax is going out of style.

Photo by Marc PEZIN on Unsplash.

## FAQ

### What is ds4 (DwarfStar) and what does it do?

ds4 (DwarfStar 4) is a local inference stack for running specific supported open-weight language models on your own hardware. It includes a CLI, an HTTP server, and a native coding agent that all share the same loaded model state. A key feature is a disk-persisted KV cache so repeated prompt prefixes can be reused across restarts.

### Is ds4 an alternative to Ollama or llama.cpp?

Yes, but only if you want ds4’s narrow, validated approach. Ollama and llama.cpp aim to run a wide range of models with fewer constraints, while ds4 focuses on a small set of model families and “project GGUF” files tested end-to-end. ds4’s persistent prefix KV cache also targets faster repeat workflows, not just one-off chats.

### What is ds4’s disk KV cache and how does it speed up prompts?

ds4 saves the key/value attention cache for a prompt prefix to disk and reuses it later. The cache key is the SHA1 hash of the rendered prompt prefix, so identical prefixes can skip expensive re-prefill work. This is most helpful for workflows with a stable system prompt and repeated context blocks, like coding agents.

### Does ds4 expose an OpenAI-compatible API?

Yes. ds4 includes `ds4-server`, which the official docs describe as supporting OpenAI- and Anthropic-style local APIs. In practice, that means many tools that can point to an OpenAI-like base URL can often talk to ds4 without a custom integration.

### Can ds4 run on multi-GPU setups and serve multiple users?

Yes. The ds4 README describes CUDA support with multi-user serving goals and includes a published example of an 8× NVIDIA L40S setup running 16 sessions concurrently. If you run it as a shared service, you’ll still want to add sensible limits and basic network security like you would for any internal API.

### Where does ds4 store downloaded models and cache files?

Models and cache files live on your local disk, and you should treat them as versioned artifacts. The exact default directories can vary depending on how you install and run ds4, but the practical move is to pick a dedicated fast SSD location and standardize it for your machine or team. That also makes backups, cleanup, and upgrades much safer.
