ds4 dwarfstar [2026 Review]: Redis-Style Local LLM Loops

ds4 (DwarfStar 4) turns local LLM inference into a Redis-like workflow: one long-lived server, deterministic prompt-keyed KV caching, and fast edit-run-edit loops without a GUI circus.

Part of theLLM Hardware & Local AI series
a close up of a server's nameplates on the side of a
Listen to this article
--:--

ds4 (DwarfStar 4) is a deliberately narrow local LLM inference stack that runs specific “project GGUF” models on Metal, CUDA, or ROCm, and exposes the same loaded model through a CLI, an HTTP server, and a native coding agent. The consequence is simple: your model loads once, your prefixes can stay warm across restarts via a disk KV cache, and your local dev loop stops feeling like a science project. This is a ds4 dwarfstar review from the perspective that matters: daily loops, not demo screenshots.

Most local runners optimize for “get a model talking.” ds4 optimizes for “get this workflow tight.” That’s a very antirez move. If you’ve ever loved Redis because it’s boringly predictable under load, ds4 will feel familiar.

Here’s the mental model I keep coming back to. ds4 treats inference state like a cacheable, persistent asset, not a disposable byproduct of one prompt.

If your local model has to cold-start its context on every run, you don’t have a dev loop. You have a latency tax.

Local frontier inference, narrow on purpose

The first thing to internalize is that ds4 is not trying to be a universal llama.cpp replacement. It’s not aiming to run “whatever GGUF I found on Hugging Face at 2am.” It’s aiming to make a small set of models feel great.

Two nvidia titan x graphics cards side by side

As Salvatore Sanfilippo (antirez) explains in the README, ds4 targets a small set of model families and expects project-produced GGUF files that are validated end-to-end. Translation: you’re trading breadth for a stack where the engine, quantization choices, and model layouts are all tested together.

I’m a fan of this trade. The local LLM world has too much “works on my machine” energy. People treat quant files like Pokémon cards, swap them constantly, then act shocked when things crash or degrade.

ds4 can make strong claims like “KV cache as a disk citizen” because it’s not trying to be everything to everyone. It’s building a coherent system.

The official site calls out three design choices that are basically the spine of the project:

  • Asymmetric 2-bit quantization aimed at routed experts (for MoE models)
  • Disk-persisted KV cache keyed by a hash of the rendered prompt prefix
  • One engine, three interfaces (./ds4, ./ds4-server, ./ds4-agent) sharing state

Those aren’t a random grab bag of features. They’re one opinion, expressed three ways.

If you want something that runs every GGUF under the sun, use the thing built for that job. If you want a Redis-style workflow where the hot path is engineered, ds4 is worth your time.

How the ds4 stack fits together (and why it feels like Redis)

The cleanest way to understand ds4 is as a single model process with multiple front doors:

Nvidia logo on a green background with abstract spheres
  • ds4 is the interactive CLI.
  • ds4-server is the HTTP API server (OpenAI + Anthropic style, per the official docs).
  • ds4-agent is a native coding agent that keeps sessions around.

All three are meant to share:

  1. the same model weights (loaded once)
  2. the same runtime state
  3. the same KV cache

This is where the “Redis-style workflow” analogy stops being cute and starts being useful.

Redis isn’t popular because it has a nice GUI. It’s popular because it behaves like infrastructure:

  • run a long-lived daemon
  • keep hot state in memory
  • persist the right artifacts to disk
  • hit it from any client that can speak the protocol

ds4 is pushing local inference in that direction. Less “open an app and chat.” More “spin up a service and build around it.”

The KV cache: the real feature

DwarfStar’s landing page spells it out: “The KV cache is keyed by the SHA1 of the rendered prompt prefix and persisted to disk, so a matching prefix is reloaded instead of recomputed.” That’s the whole point.

In practice, if your coding agent has a stable system prompt and a stable “repo context preamble,” ds4 can avoid paying the prefill cost again after restarts.

That changes how you work.

  • You restart less often because you’re not terrified of losing warm context.
  • You start designing prompts with stable prefixes on purpose.
  • You start caring about cache hit rate the way you care about Redis hit rate.

If you’ve been building AI agents or doing agent orchestration, you already know the pain. Agent loops are only as good as their latency budget.

And if you’re doing local dev to avoid LLM cost, persistent caching isn’t a “nice to have.” It’s the difference between “I’ll use this daily” and “cool demo, never again.”

Supported hardware (local, streamed, and distributed)

ds4’s hardware story lives in three worlds: Metal, CUDA, and ROCm.

Two nvidia titan x graphics cards side by side

Per the ds4 README by Salvatore Sanfilippo (antirez), the primary target is high-memory Macs. The docs call out 96GB+ Apple Silicon as the comfortable tier for the biggest supported models, and also mention SSD streaming as an option for “smaller RAM / very large models.”

On the other side:

  • NVIDIA CUDA is supported, including multi-GPU serving.
  • ROCm is supported with a focus on newer platforms like Strix Halo systems.

If you’re trying to decide between platforms, start with the boring truth I keep repeating in my benchmark work: memory dictates feasibility, throughput dictates happiness.

Based on the benchmark data I maintain at kunalganglani.com/llm-benchmarks, unified memory on Apple Silicon often makes “will it load?” easy, but “will it feel fast?” is still the real question. Local inference usually fails not because it can’t run, but because it runs at a pace that breaks your flow.

If you’re building a shared box for a team, ds4’s CUDA and multi-user angle matters more than the Mac story. A fast headless server that people can hit over the network beats “everyone runs a laptop stack” in most real orgs.

Run ds4 in three steps (day 1 to daily driver)

I’m going to keep this practical and reproducible. Not “click this download button” practical. “Can you put it in a Makefile and forget about it?” practical.

Step 1: install ds4 (macOS Metal / Linux CUDA / ROCm)

ds4 is a C project with multiple backends. The canonical source of truth is still the repo README.

Start here:

If you want a friendlier “tooling wrapper,” the community Go project is worth a look:

  • NimbleMarkets ds4go

ds4go is interesting because it treats ds4 as a reusable library via FFI and puts real effort into model download ergonomics.

Step 2: download/select a supported model (and know where it lives)

ds4’s narrow stance shows up immediately in model management. You’re not pulling arbitrary community quantizations and hoping they behave.

Supported model families called out across the official site and README include:

  • DeepSeek V4 Flash and V4.1 Flash
  • GLM 5.x (including Flash variants)
  • DeepSeek V4 PRO
  • Qwen3.8 Flash Next

Support can vary by backend, so don’t assume “it runs on Metal” implies “it runs on ROCm.”

On disk, treat ds4 models the way you treat Docker images or Bazel caches: pin them, version them, and share them intentionally. If you’re doing this as a team, don’t hand-wave it.

A pattern that works:

  • Put a models/ directory on a fast SSD volume.
  • Store model version + quant variant in the path.
  • Commit a tiny “model manifest” file to your repo.

It’s a small bit of discipline that buys you fewer surprises, faster onboarding, and less “wait, which file are you using?” slack drama.

If you want the broader “how quants behave and where they cliff,” my quant deep-dive is here: LLM quantization.

Step 3: run `ds4-server` and connect an OpenAI-compatible client

The official site explicitly positions ds4-server as “OpenAI + Anthropic API.” That matters because it means you can often drop ds4 under existing clients without rewriting your whole toolchain.

I’m not going to pretend every client works perfectly. OpenAI-compat is a spectrum. But it covers the 80% case: curl tests, SDKs, and a lot of local tooling.

A minimal sanity check looks like this:

  • start ds4-server bound to localhost only
  • hit a /v1/... endpoint with a small prompt
  • confirm streaming behavior (if your client needs it)

Once that works, wire it into whatever you actually use.

If your day-to-day flow involves Claude Code or other agentic coding tools, the question is always the same: can the tool point at an OpenAI-style base URL, and does it behave under retries when the model is busy?

Everyday use: caching, agent loops, and what actually gets faster

The place ds4 feels different from Ollama or LM Studio is not “tokens per second on a clean benchmark.” It’s what happens after your fifth iteration on the same task.

That’s the reality of coding workflows. You don’t ask once. You ask, tweak, re-run, paste an error, re-run, then realize your prompt template changed and everything got slower again.

ds4 disk KV cache: what it caches and how to maximize hits

ds4 caches prefix KV. The official description is precise: the cache key is the SHA1 of the rendered prompt prefix.

That implies three rules that are painfully easy to ignore:

  1. Render deterministically. If your template injects timestamps, random IDs, or unstable tool metadata above the fold, you’re nuking your own cache.
  2. Stabilize the preamble. Keep your system prompt, repo summary, and “rules” block identical across runs.
  3. Move volatile context later. Put “recent diffs,” error logs, and one-off notes after the stable prefix boundary.

If you’ve built RAG systems, this will sound familiar. You’re basically optimizing for a cache-friendly header. Same muscle group.

If you’re doing RAG or retrieval-augmented generation, you already know the drill. Separate stable context from volatile context.

Disk growth, invalidation, and SSD reality

A disk KV cache isn’t free.

  • It will grow.
  • It will create real write activity.
  • It will feel terrible on a slow drive.

The upside is still worth it. Just don’t be casual about it. In Redis terms: don’t put your persistence files on a bargain-bin disk and then complain about latency.

If you’re doing this on a laptop, look at your cache directory size weekly, not once a quarter when you’re out of space mid-flight.

Native coding agent vs “bring your own agent framework”

ds4 ships ds4-agent, a native coding agent concept.

Whether you use it depends on how invested you are in your existing stack. If you’re already deep into agentic AI frameworks or you’ve built your own loop, you’ll probably just want ds4-server.

If you’re earlier, a native agent that shares model state and cache can be a solid “one stack” experience.

This is the ds4 vibe in one sentence: stop wiring four different tools together for something that should be one process.

Capability evaluation, speed, and how ds4 compares to Ollama/LM Studio

Let’s talk performance without turning this into benchmark cosplay.

ds4’s README includes a concrete serving example: an 8× NVIDIA L40S box reaching ~126 tokens/sec aggregate generation with 16 simultaneous sessions. That’s straight from Salvatore Sanfilippo (antirez).

That number is interesting for two reasons:

  1. It’s explicitly aggregate under concurrency. That’s what real workloads look like.
  2. It’s a signal that ds4 cares about multi-session serving, not just single-user desktop chatting.

Prefill vs decode: what matters for dev loops

For coding workflows, you usually care about time-to-first-token (TTFT) and prefill cost as much as raw decode throughput.

  • Prefill is what hurts when you keep re-sending the same long context.
  • Decode throughput matters once you’re already streaming.

This is why ds4’s disk KV cache is such a targeted bet. It’s not trying to be a generic “semantic cache.” It’s saying: stop paying prefill repeatedly for the same prefix.

If you want a deeper mental model on the metrics, my benchmark methodology post is here: local LLM.

ds4 vs Ollama vs LM Studio vs llama.cpp (comparison table)

Here’s the comparison I wish more reviews would write. Not “which is best,” but “which breaks your loop less.”

Dimensionds4 (DwarfStar 4)OllamaLM Studiollama.cpp
ScopeNarrow, validated model families + project GGUFsBroad model catalog, friendly UXBroad GGUF UX with strong GUI featuresBroad engine, DIY integrations
WorkflowOne engine with CLI/server/agent sharing stateCLI + background serviceGUI-first, can serve on LANLibrary/CLI building block
Cache across restart**Yes**: disk KV cache keyed by SHA1 prefixUsually no persistent prefix KVUsually no persistent prefix KVPossible with work, not default
Best fitTight local dev loops + multi-user serving + specific frontier open weights“Just run a model” local dev and tooling ecosystemPower users who want GUI + profiles + LAN servingBuilders who want maximum flexibility
When it losesUnsupported model families, non-project GGUFsMore latency tax on repeated long prefixesGUI overhead if you want headless-firstMore assembly required

If you’re an LM Studio power user, I wrote my setup notes here: local LLM.

If you’re choosing between Ollama and LM Studio in general, this is the sharper comparison: local LLM.

Ops guidance: running ds4-server for a team (and basic security)

If you’re going to run ds4-server as a shared internal service, treat it like any other internal service. No hero setups. No “it’s just on the LAN.”

Run it like a daemon, not a terminal tab

  • Use systemd on Linux or launchd on macOS.
  • Write logs to a file and rotate them.
  • Pin the model version so “someone updated it” doesn’t become your next incident.

You’ll also want quotas or concurrency limits. ds4 is designed for sessions. Your GPU is not designed for infinite enthusiasm.

If you’re thinking “this is starting to sound like production,” yes. That’s why I treat local inference as part of production AI.

Bind addresses and authentication

Do not bind an LLM server to 0.0.0.0 on your home network and call it a day. Even on a “trusted” LAN, you’re exposing a service that can be abused.

At minimum:

  • bind to 127.0.0.1 by default
  • put an auth proxy in front if you need remote access
  • restrict egress if you’re mixing local and remote tools

If you want the bigger checklist, I wrote it here: AI security.

And if you’re using agentic tools that read your repo, internal docs, or tickets, you also need to care about prompt injection and broader AI security.

Status, upgrades, and when not to use ds4

Status: fast-moving by design

ds4 is actively evolving. That’s part of its appeal and part of the tax.

As of October 2026, the public narrative is clear:

  • model support is expanding (Qwen/vision mentions show up in community chatter)
  • APIs are positioned as OpenAI + Anthropic compatible
  • multi-user / multi-GPU serving is a first-class target

If you adopt ds4, decide up front: are you pinning versions for stability, or tracking main for features? “We’ll figure it out later” is how you end up with a broken toolchain on a Monday morning.

How to update ds4 and models safely

My boring recommendation:

  • Treat ds4 like a compiler. Upgrade intentionally.
  • Keep a known-good binary around.
  • Pin model artifacts by version and checksum.

If you want a supply chain mindset for model files, this is worth reading: LLM security.

When you should not use ds4

Use ds4 when you agree with its premise. Don’t use it when you want it to be something it’s not.

Skip ds4 if:

  • you need broad support for random GGUFs from the internet
  • you’re experimenting with lots of niche model families
  • you want a GUI-first experience
  • you can’t give it fast local SSD for cache + weights

In those cases:

  • Ollama is a great default “just run it” tool.
  • LM Studio is excellent if you want profiles, quants, and a power-user GUI.
  • llama.cpp is still the universal building block.

I know “use the right tool” is obvious advice. But teams still waste weeks trying to force a tool to match constraints it was never built for.

The interesting implication of ds4 isn’t that it replaces everything. It’s that it points to the next phase of local inference: stateful, cache-aware, multi-interface runtimes that behave more like infrastructure and less like a toy app.

My bet: within a year, “persistent KV cache across restarts” will be table stakes for serious local dev loops. If you’re building developer tooling, treat ds4’s cache design as the baseline. The latency tax is going out of style.

Photo by Marc PEZIN on Unsplash.

Continue reading

LM Studio app desktop client screen local LLM — illustration for article on LM Studio Power

LM Studio Power User Setup 2026: Profiles, Quants, LAN Serving

My opinionated 2026 LM Studio setup: reproducible model profiles, quant hygiene rules, scripted loads via lms + /api/v1, and a secure multi-user LAN server.

nvidia gpu graphics card close-up — illustration for article on Local LLM Break-Even Math [2026]: Power,

Local LLM Break-Even Math [2026]: Power, Idle, Depreciation

Most “GPU price ÷ tokens” break-even math is wrong. Here’s a spreadsheetable local LLM total-cost model that includes idle power, utilization, depreciation, failures, and opportunity cost.

macbook terminal dark code screen programmer open source — illustration for article on Hermes Agent Desktop

Hermes Agent Desktop Free With Local LLMs: The Claude Code Alternative Nobody's Billing You For [2026]

Hermes Agent runs a full coding agent on your local machine with zero API costs. Here's which models actually work, the hardware you need, and how to set it up.

vLLM vs Ollama 2026: Production Power or Developer Ease?

vLLM vs Ollama 2026: Production Power or Developer Ease?

vLLM wins for high-throughput production deployments where every token/second counts; Ollama wins for local developer workflows where setup speed and portability matter most. Pick wrong and you'll either over-engineer a side project or under-power a real API.

Cite this article
Kunal Ganglani (2026, October 3). ds4 dwarfstar [2026 Review]: Redis-Style Local LLM Loops. Kunal Ganglani. Retrieved October 3, 2026, from https://www.kunalganglani.com/blog/ds4-dwarfstar-review

Frequently Asked Questions

What is ds4 (DwarfStar) and what does it do?

ds4 (DwarfStar 4) is a local inference stack for running specific supported open-weight language models on your own hardware. It includes a CLI, an HTTP server, and a native coding agent that all share the same loaded model state. A key feature is a disk-persisted KV cache so repeated prompt prefixes can be reused across restarts.

Is ds4 an alternative to Ollama or llama.cpp?

Yes, but only if you want ds4’s narrow, validated approach. Ollama and llama.cpp aim to run a wide range of models with fewer constraints, while ds4 focuses on a small set of model families and “project GGUF” files tested end-to-end. ds4’s persistent prefix KV cache also targets faster repeat workflows, not just one-off chats.

What is ds4’s disk KV cache and how does it speed up prompts?

ds4 saves the key/value attention cache for a prompt prefix to disk and reuses it later. The cache key is the SHA1 hash of the rendered prompt prefix, so identical prefixes can skip expensive re-prefill work. This is most helpful for workflows with a stable system prompt and repeated context blocks, like coding agents.

Does ds4 expose an OpenAI-compatible API?

Yes. ds4 includes `ds4-server`, which the official docs describe as supporting OpenAI- and Anthropic-style local APIs. In practice, that means many tools that can point to an OpenAI-like base URL can often talk to ds4 without a custom integration.

Can ds4 run on multi-GPU setups and serve multiple users?

Yes. The ds4 README describes CUDA support with multi-user serving goals and includes a published example of an 8× NVIDIA L40S setup running 16 sessions concurrently. If you run it as a shared service, you’ll still want to add sensible limits and basic network security like you would for any internal API.

Where does ds4 store downloaded models and cache files?

Models and cache files live on your local disk, and you should treat them as versioned artifacts. The exact default directories can vary depending on how you install and run ds4, but the practical move is to pick a dedicated fast SSD location and standardize it for your machine or team. That also makes backups, cleanup, and upgrades much safer.