How to Build a Local Model Registry for GGUF [2026]

Stop downloading random weights from links. Build a local model registry with GGUF versioning, provenance attestations, eval gates, and a real rollback plan.

Part of theLLM Hardware & Local AI series
gguf model files local registry laptop terminal — illustration for article on How to Build a
Listen to this article
--:--

You’re going to end this tutorial with a working local model registry where your team can:

  • publish GGUF weights as versioned artifacts
  • promote them across dev → staging → prod without re-uploading
  • verify hashes/digests every time someone pulls
  • attach SLSA provenance + signatures (so “where did these weights come from?” has a real answer)
  • roll back in minutes when a new quant quietly ruins latency or quality

If Harbor is already running in your environment, budget 60–90 minutes. If it isn’t, this is still the right pattern. You’ll just spend your time on “Harbor Installation and Configuration” first.

And yes. Downloading weights by URL is the new `curl | bash`. It feels fine right up until the day you have to answer: “which exact file was in prod?” and all you have is a half-edited Notion page and a bunch of renamed .gguf files.

This post is opinionated on purpose. If you’re doing local inference with GGUF and you don’t have local model registry gguf versioning provenance nailed down, you’re shipping a supply chain risk with a GPU attached.

What is a local model registry for GGUF/open weights?

A local model registry is a team-owned system for storing, versioning, verifying, and promoting model artifacts (like GGUF weights) so deployments pull by immutable digest, with recorded provenance, evaluation results, and a repeatable rollback process.

Nvidia logo on a green background with abstract spheres

Teams need one for three boring reasons that become urgent overnight:

  1. Reproducibility: “Which exact weights were in prod at 2:17pm?” needs an answer that isn’t Slack archaeology.
  2. Safety + security: without integrity checks and provenance, you cannot distinguish “legit update” from “someone swapped a file.”
  3. Operational control: model incidents are real incidents. You need the equivalent of container image promotion and rollback.

A concrete example: in the order-cancellation microservice I built at Swiggy (which supported millions of deliveries), the biggest failures weren’t raw traffic spikes. They were workflow edge cases and missing compensation paths. Local model rollouts are the same shape. The failure mode is rarely “GPU melted.” It’s “we changed a dependency-like artifact with no escape hatch.”

Also, from maintaining the live pricing data at /llm-prices, I’ve learned a simple lesson: when something becomes operationally important, you stop trusting vibes and start tracking inputs. Model weights deserve the same treatment.

OCI registries, OCI artifacts, and how ORAS works

I keep seeing “model registries” that are really just an NFS folder, a shared drive, or a blessed S3 bucket with some conventions taped on top.

Nvidia logo on a green background with abstract 3D elements

That setup is fine for a weekend project. It falls apart the moment a second team depends on it, or you have to explain an incident to security.

What are OCI Registries?

OCI registries are servers that implement the Open Container Initiative’s registry APIs (the OCI Distribution Spec), which is why tools like docker can push/pull images. ORAS explains this framing explicitly in its docs: registries are evolving into generic artifact stores, not just container image warehouses. See the ORAS project.

Concrete advantage: OCI registries give you RBAC, audit logs, immutability policies, retention policies, and replication in a system your infra team already knows how to run.

What are OCI Artifacts?

OCI artifacts are non-container things stored in an OCI registry without pretending they’re container images. ORAS calls out the old anti-pattern: stuffing random files into image layers. OCI artifacts instead rely on correct mediaType usage so clients can understand what they’re pulling.

For GGUF, this matters because you want:

  • a clear mediaType for .gguf
  • a manifest that can point to related files (tokenizers, templates, eval results)
  • promotion by tagging the same immutable digest

How ORAS works

ORAS is the de facto CLI for working with OCI artifacts. It pushes and pulls blobs plus a manifest to/from the registry. Unlike docker, it treats “media types” as the central primitive.

In practice, your flow becomes:

  1. download upstream weights (Hugging Face or elsewhere)
  2. verify file hash
  3. generate a model manifest
  4. push GGUF + manifest to registry
  5. sign and attach provenance attestations
  6. promote by digest and tags

That’s the spine of this tutorial.

[IMAGE: Diagram of a local GGUF model registry pipeline: upstream weights → build pipeline → Harbor OCI registry → runtimes pulling by digest]

Download upstream safely (Using Git + Faster downloads)

Your registry is only as trustworthy as your ingest pipeline. If your ingest is “somebody grabbed a file from Discord,” your registry is just an internal mirror of chaos.

Two nvidia titan x graphics cards side by side

Using Git

Hugging Face supports downloading models via multiple methods, including Git. That’s useful because a Git commit hash (or tag) is a clean provenance input. Use the official docs: Hugging Face model downloading.

In a build pipeline, I prefer a pattern like:

  • record upstream.repo (e.g. org/model)
  • record upstream.revision (commit SHA or tag)
  • record the exact file name(s) pulled

If you’re downloading GGUF directly from a community repo, you still want an immutable identifier. If the source can’t provide one, you mint your own: compute SHA256 and treat that as “the ID.”

Faster downloads

The same HF docs include operational guidance on faster downloads. At team scale, “faster downloads” stops being a convenience and becomes the difference between “people actually use the system” and “everyone works around it.”

Concrete numbers: GGUF files are commonly 4–20+ GB. One engineer casually re-downloading a 12 GB file across a VPN a few times is enough to make everyone hate local inference.

Practical defaults I recommend:

  • use a shared build runner with a warm HF cache
  • mirror upstream into your registry once, then pull internally forever
  • never promote anything that wasn’t pulled and hashed in CI

Design a GGUF naming + SemVer versioning scheme

If your versioning scheme doesn’t survive a 2am incident, it’s not a scheme. It’s a label maker.

Here’s what I use for teams:

Repository path (Harbor project/repo)

  • models/<family>/<name>
    • example: models/qwen/qwen2.5-coder-32b

SemVer tags

  • vMAJOR.MINOR.PATCH for the “logical model release” your application depends on
    • v1.4.0 means: new capability or measurable behavior change
    • v1.4.1 means: same behavior, packaging/metadata fix

Build metadata tags (optional)

  • v1.4.0+gguf.q4_k_m.llamacpp.2026-09-15

Release channel tags

  • dev, staging, prod

The key is: release channels are just tags that point at the digest you’ve blessed. They’re pointers, not identities.

A concrete reminder from the llama.cpp gguf-py repo: their release tags follow SemVer like gguf-vx.x.x, which is exactly the mental model you want to mirror for model artifacts. See the upstream README in the llama.cpp repo: gguf-py README.

The one rule I’m strict about

Never deploy by mutable tag alone.

Deploy by digest pin, and let the tag be a human-friendly alias.

Create a model manifest (provenance, hashes, evals)

You need a GGUF model manifest schema that answers:

  • What is this artifact?
  • Where did it come from?
  • How was it produced?
  • What files are included and what are their hashes?
  • What eval gates did it pass?

Below is a concrete, copy-pastable schema. Keep it in Git. Validate it in CI. If someone edits it by hand in prod, your process is already broken.

Example: `model-manifest.yaml`

yaml
schemaVersion: 1
kind: gguf.model
name: qwen2.5-coder-32b
version: 1.4.0
channel: dev

artifact:
  filename: qwen2.5-coder-32b.Q4_K_M.gguf
  sizeBytes: 12884901888
  sha256: "<sha256 of the gguf file>"
  mediaType: application/vnd.ggml.gguf

upstream:
  source: huggingface
  repo: Qwen/Qwen2.5-Coder-32B-Instruct
  revision: "<git commit sha or tag>"
  files:
    - filename: model.safetensors
      sha256: "<sha256>"

build:
  conversionTool:
    name: llama.cpp
    version: "<git sha or release tag>"
  quantization:
    scheme: Q4_K_M
    contextLength: 32768

runtime:
  promptTemplate:
    id: chatml
    sha256: "<sha256 of template text>"
  tokenizer:
    filename: tokenizer.json
    sha256: "<sha256>"

license:
  spdx: "<SPDX id if known>"
  notes: "Any internal usage constraints"

evals:
  suite: "local-llm-regression"
  runId: "2026-10-06T02:14:11Z"
  thresholds:
    - metric: "helpfulness.winrate"
      min: 0.52
    - metric: "safety.refusal_rate"
      max: 0.08
    - metric: "latency.ms.p95"
      max: 900
  results:
    - metric: "helpfulness.winrate"
      value: 0.56
    - metric: "latency.ms.p95"
      value: 740

notes:
  - "Promoted after passing evals on RTX 4090 runner"

There are at least 10 numeric anchors in that manifest (size, context length, thresholds). That’s intentional. If your manifest can’t be searched and diffed like real metadata, it’s just documentation cosplay.

Generate the hash + size fields (real, runnable)

This is the part most tutorials skip, then the whole thing turns into vibes.

bash
GGUF_FILE="qwen2.5-coder-32b.Q4_K_M.gguf"

# SHA256 hash
SHA256=$(sha256sum "$GGUF_FILE" | awk '{print $1}')

# Size in bytes (Linux)
SIZE_BYTES=$(stat -c%s "$GGUF_FILE")

echo "sha256=$SHA256"
echo "sizeBytes=$SIZE_BYTES"

On macOS, swap stat -c%s for stat -f%z.

Optional: embed registry pointers inside the GGUF metadata

llama.cpp’s gguf-py includes gguf_dump.py and gguf_set_metadata.py, which let you inspect and change metadata keys. That means you can stamp in:

  • registry.url
  • registry.digest
  • manifest.sha256

The source is explicit about these scripts existing and what they do: gguf-py README.

This isn’t a replacement for registry-side signing. It’s a nice belt-and-suspenders move for the inevitable moment a .gguf gets copied to a random machine “just for testing.”

Store GGUF in Harbor using ORAS (with immutability)

This is the “OCI registry for model artifacts ORAS” part.

Harbor Installation and Configuration

If you don’t have Harbor yet, start with Harbor’s own install docs. They’re clear, and they change over time, so I’m not going to paste a stale mini-guide here: Harbor documentation.

As of the Harbor docs UI I’m looking at (2026), the documentation set is versioned (for example, v2.15.0 is labeled “latest” in their selector). Good. Treat Harbor like any other critical platform. Pin a version, upgrade deliberately, and don’t let “latest” sneak into prod because someone ran the wrong Helm command.

Working with Harbor Projects

Before you push anything:

  • create a Harbor project (e.g. models)
  • decide who gets developer vs maintainer permissions
  • enable project policies you care about (tag immutability, retention)

Again, Harbor documents this under “Working with Harbor Projects” in their docs hub: Harbor documentation.

Push artifacts with ORAS

Pick explicit media types. Don’t throw everything into application/octet-stream and call it “standardized.”

Example layout:

  • config: a tiny JSON config object describing this as a GGUF artifact
  • layer 1: the GGUF file
  • layer 2: your manifest YAML/JSON
  • optional layers: prompt template, tokenizer, eval report

A practical push looks like:

bash
REGISTRY="harbor.internal:443"
REPO="$REGISTRY/models/qwen/qwen2.5-coder-32b"
VERSION_TAG="v1.4.0"

oras login "$REGISTRY"

# Push GGUF + manifest as OCI artifact layers
oras push "$REPO:$VERSION_TAG" \
  --artifact-type application/vnd.ggml.gguf \
  qwen2.5-coder-32b.Q4_K_M.gguf:application/vnd.ggml.gguf \
  model-manifest.yaml:application/vnd.kunalganglani.model.manifest.v1+yaml

Resolve the immutable digest (version-dependent)

Teams get tripped up here because ORAS commands vary by version.

Try oras resolve first. If your installed ORAS doesn’t have it, fall back to fetching the manifest and extracting the digest.

bash
# Option A: resolve (newer ORAS)
oras resolve "$REPO:$VERSION_TAG"

# Option B: fetch manifest (works broadly)
oras manifest fetch "$REPO:$VERSION_TAG" > /tmp/manifest.json

Once you have the digest (sha256:...), do promotions like this:

bash
DIGEST="sha256:<digest>"

# Tag the same digest into channels
docker pull "$REPO@$DIGEST" 2>/dev/null || true
oras tag "$REPO@$DIGEST" dev
oras tag "$REPO@$DIGEST" staging

(If your ORAS version uses a different tagging command, align to your org’s pinned ORAS release. The idea is constant: tags move. Digests don’t.)

Verification and attestations (hashes, SLSA, Cosign)

This is the line between “we have a folder of models” and “we have a registry we can defend.”

Integrity verification strategy

You want integrity checks at three points:

  1. On download: compute SHA256 and compare to an expected hash when possible.
  2. At rest: store the GGUF as a content-addressed blob in the registry. OCI digests are SHA256-based.
  3. At deploy/load time: pull by digest and verify signature/attestation before the runtime uses it.

If you’ve ever had to chase a heisenbug caused by “the artifact changed but the tag didn’t,” you already know this pain.

SLSA provenance (and the predicateType you must not freestyle)

SLSA provenance is defined as “verifiable information about software artifacts describing where, when and how something was produced.” That’s straight from the SLSA spec: SLSA provenance spec.

The same page defines the provenance predicate type in the in-toto format. The exact string is:

  • https://slsa.dev/provenance/v1

It also warns you not to copy whatever is in your browser URL bar for the predicate. Use the literal predicate type string.

Freshness note: the SLSA site also states that v1.2 is the current version, and the v1.0 page is marked “Retired.” The important part for teams is that the predicateType string is designed to resolve to the latest minor version. That’s exactly what you want for long-lived attestations.

Sign and attach with Sigstore Cosign

Cosign is the practical tool most teams already trust for signing artifacts in registries.

Cosign’s signing docs cover workflows for containers and also “Signing Blobs” / “Signing Other Types,” which is what model files effectively are in an OCI registry: Sigstore Cosign signing overview.

In a CI pipeline, a typical pattern is:

  • push artifact to registry
  • sign the digest
  • attach an in-toto attestation containing SLSA provenance

Your policy controller (or your runtime wrapper) should refuse to load unsigned artifacts in prod.

Eval gates + promotion + rollback playbook (the part everyone ignores)

People love talking about “provenance.” They love talking about “local inference.” They love shipping.

They do not love writing the rollback doc.

Then a new quant quietly breaks quality or blows your latency budget, and suddenly everyone is extremely interested in that doc.

What eval gates should run before promotion?

Before you move a digest from staging to prod, require at least:

  1. Quality regression: your task suite win-rate or rubric score. Pick a number. Example: helpfulness.winrate >= 0.52.
  2. Safety regression: refusal rate or jailbreak success rate. Example: safety.refusal_rate <= 0.08.
  3. Perf budget: P95 latency under your SLA on your real hardware. Example: latency.ms.p95 <= 900.
  4. VRAM / RAM cap: ensure the quant actually fits. Example: must run on 16 GB VRAM tier if that’s your fleet.
  5. Prompt injection regression tests if the model is used in an agent or tool-calling flow.

If you’re running AI agents or any production AI system, these gates are not optional. They’re the seatbelt.

For eval harness design, I’ve had good outcomes treating evals like CI. Version the dataset, version the prompts, and record the full run. If you want a deeper eval framework, see how I think about gates in AI engineering evals and safety testing in prompt injection regression testing.

Where do eval results live?

Two places:

  • in your registry as an artifact (JSON report + summary)
  • in your observability system for trend lines

If you’re already using MLflow for tracing, the mental model maps well. I’ve written about using spans as gates in MLflow LLM evaluation tracing.

The rollback playbook: “roll back the weights”

Here’s the playbook I actually want teams to have printed somewhere.

Triggers (pick at least 3 and make them measurable):

  • P95 latency regression > 20% on the same hardware
  • task success rate drop > 5 points on the core eval suite
  • safety regression: jailbreak pass rate increases, or refusal policy breaks
  • memory regressions: OOM on the smallest supported GPU tier

RACI (keep it simple):

  • Responsible: on-call ML/infra engineer
  • Accountable: service owner (who owns the user impact)
  • Consulted: security (if provenance/signing mismatch), product (if user-visible)
  • Informed: incident channel, support

Steps (the muscle memory):

  1. Freeze promotions: block moving any new digest into prod.
  2. Identify current prod digest: your deployment should already record it.
  3. Swap prod tag back to the last known good digest (LKG). This should be one command.
  4. Redeploy/pull by digest: runners should pull the pinned digest, not a tag.
  5. Verify signature/attestation on the rolled-back digest.
  6. Run a post-rollback smoke eval (10–20 prompts). Timebox: 10 minutes.
  7. Write the incident note: what changed (upstream revision, quant scheme, conversion tool), when, who approved.

Automation hook: if you’re using Kubernetes, treat the model digest as a config value in GitOps. Rollbacks become git revert.

This is exactly the same operational philosophy as container rollbacks. The only difference is the artifact is a GGUF file instead of an image.

Governance: retention, access control, audit logs, and licenses

The hard part of “model management” is not the storage. It’s the governance. Storage is easy. Discipline is not.

My baseline rules:

  • Immutability: forbid overwriting tags like v1.4.0. If you need a fix, bump PATCH.
  • Retention: keep at least 3 previous prod digests per model family. Storage is cheaper than incidents.
  • RBAC: only maintainers can move prod tags. Developers can push to dev.
  • Audit logging: promotion events must be traceable. Harbor already has the primitives.
  • License compliance: store license metadata in the manifest. If legal says “no external redistribution,” enforce that at the registry boundary.

If you’re building anything agentic, fold this into your broader AI security posture. Model artifacts are now part of your supply chain, right next to npm and pip.

[IMAGE: Screenshot-style illustration of Harbor projects: models project with repos, tags dev/staging/prod, and retention/immutability policies]

A quick comparison: OCI registry vs file share vs HF cache

If you’re on the fence, this is the decision table I wish more teams would make explicit.

Storage approachVersioning disciplineIntegrity checksRBAC + auditPromotion workflowRollback speed
OCI registry (Harbor) + ORASStrong (tags + digests)Strong (digests + signing)StrongStrong (retag by digest)Minutes
File share / NASWeak (human naming)Manual (if any)WeakWeakHours
Hugging Face cache on dev machinesNoneNoneNoneNonePainful

This isn’t about “enterprise vibes.” It’s about not burning a week on a preventable incident.

Official Harbor walkthrough (optional)

If you’re new to Harbor, here’s an approachable overview video. Watch it for the product concepts, not the exact clicks.

What I’d do next (and my prediction)

If you implement this once for a single GGUF, you’ll quickly want to generalize it to:

  • LoRA adapters
  • prompt templates
  • tool schemas
  • eval datasets
  • the runtime itself (llama.cpp, vLLM, Ollama wrappers)

My prediction for 2026: supply chain controls won’t stay “a container thing.” They’ll become “an everything thing.” Model weights are just the first awkward artifact we’ve had to drag into the grown-up world.

If you’re building local inference for a team and you still deploy models by downloading a file from a URL, stop. Build the registry. Then practice the rollback once. You’ll thank yourself later.

Continue reading

A command line interface showing the text ubuntu@ubuntu:~$ sudo with a blinking cursor

Verify GGUF Model Hashes Supply Chain [2026]: 10 Steps

A team-ready workflow to verify GGUF integrity: compute SHA256, require signed manifests, handle mirrors safely, scan sidecars, and ship updates via an internal registry.

padlock on laptop with light trails

LLM Supply Chain Security Checklist: Lock Down Agents [2026]

A CI-ready checklist to treat models, MCP servers, tool plugins, and prompt packs as real dependencies. Pin, sign, attest, sandbox, and monitor before your agent ships malware.

claude code terminal laptop screen — illustration for article on Claude Code Security [2026]: Risks, Safe

Claude Code Security [2026]: Risks, Safe Setup, Team Policy

Claude Code is safe only if you treat it like a junior engineer with terminal access. Here’s the 2026 playbook: permissions, sandboxing, egress controls, MCP allowlists, retention settings, and incident response.

A close up view of a computer tower

Local LLM Cost vs Cloud API: 2026 Break-Even Math [Calculator]

A workload-specific break-even framework with real per-token math — hardware amortization vs. API spend — for coding, RAG, and batch workloads in 2026.

Cite this article
Kunal Ganglani (2026, October 6). How to Build a Local Model Registry for GGUF [2026]. Kunal Ganglani. Retrieved October 6, 2026, from https://www.kunalganglani.com/blog/local-model-registry-gguf