vLLM Self-Hosted LLM Production Checklist [2026]: Auth + Quotas

A production-first checklist for self-hosting an OpenAI-compatible vLLM endpoint: auth, per-tenant quotas, streaming SSE, queueing, and redaction-safe logs.

Part of theAI in Production series
Rows of black server racks with white logos in a data center

You’re going to end up with a self-hosted vLLM endpoint that looks like a managed LLM API from the outside. OpenAI-compatible paths, streaming responses that don’t buffer, tenant-scoped auth, per-tenant quotas (requests/min, tokens/min, concurrency), and logs you can actually keep without waking up your privacy team.

This is a production checklist for the exact problem most teams hit in week two. The demo works. Then three internal teams point their agents at it, someone pastes a key into a CI log, the SSE stream buffers behind a proxy, and your GPU box turns into a noisy-neighbor fight club.

If you’re here for the keyword: vllm self hosted llm production checklist. That’s what this is.

What is vLLM?

vLLM is an open-source large language model (LLM) inference server that boosts throughput under concurrent load using continuous batching and a memory management technique called PagedAttention.

Nvidia logo on a green background with abstract spheres

If you’ve only used it as “run a container, hit /v1/chat/completions, celebrate,” you’ve seen the fun part. The boring part is what makes it usable by more than one person.

Aleksei Aleinikov nails the positioning: vLLM is about production throughput at concurrency, while tools like Ollama and llama.cpp are optimized for “run locally with almost no setup.” That’s not a value judgement. It’s a routing decision. If you need multi-user capacity, you need the checklist.

Internal context that helps:

  • If you’re still deciding between “local-first UX” and “multi-tenant serving,” start with my local LLM hub and the production angle in AI in production.
  • If you already chose vLLM, you’re in the land of production AI, not hobby setups.

Getting Started With Each (vLLM vs Ollama) — pick the right tool first

This section exists because I’ve watched teams do the wrong thing: they pick an inference server based on the easiest demo, then spend months trying to retrofit multi-tenancy.

Two nvidia titan x graphics cards side by side

Here’s the stance:

  • If you’re serving one developer on one machine, pick Ollama.
  • If you’re serving many users, many services, or agents that will spike concurrency, pick vLLM.

And yes, you can start with Ollama and migrate. But migration is rarely clean because the first thing people do is let clients grow “API surface area” you didn’t design.

Quick comparison (this table is here on purpose. AI Overviews love it, and so do humans):

Problem you will hit in productionvLLM answerOllama answerWhere it belongs
Multi-tenant auth (keys/JWT)Bring your own gatewayBring your own gatewayAPI gateway / reverse proxy
High concurrency throughputContinuous batchingNot the goalInference server
Token-based quotas (TPM)Needs external accountingNeeds external accountingGateway + app-side meter
Streaming SSE under loadWorks, but proxy must be tunedWorks locally, less proxy complexityProxy + client
Kubernetes scalingDesigned for itPossible, awkwardK8s + autoscaling
Developer “it just runs”Not the pointCore productLocal dev

If you want more on choosing runtimes, my post on How to serve a local LLM to multiple users is the “decision tree” version of this.

The Real Business Bottleneck: Reliability Under Load

The business bottleneck is not “can we run the model.” It’s “can we run it predictably when five things go wrong at once.”

Nvidia logo on a green background with abstract 3D elements

Jangwook Kim frames this perfectly in his SLA-oriented writeup: reliability under load is the thing that actually matters when you move from notebook to a service. He points to a public vLLM issue reporting throughput plateauing when concurrency scaled from 4 to 16 on an NVIDIA H100 with vLLM 0.19.1. That’s the part most teams miss.

You can do everything “right” and still hit a scaling wall.

Three practical implications:

1) Benchmark per model, per GPU, per vLLM version. The 0.19.1 vs 0.20.x bump can change behavior. Your model family can change it too.

2) Admit your endpoint is a shared resource. If you let tenants fight for it, your p95 looks like a lie.

3) Your SLOs must be LLM-native. “Request latency” is not enough. You need TTFT, tokens/sec, queue time, and aborted streams.

If you haven’t done this measurement work before, start with my Local LLM benchmark methodology and LLM latency benchmark methodology: streaming UX metrics.

Production Architecture & Code Blueprints (Auth, Quotas, Streaming, Logs)

I’m going to describe the architecture in layers, because that’s how you actually ship it.

Layer 0: vLLM pods

  • Runs the OpenAI-compatible API.
  • You treat it like a dumb engine. No tenancy logic. Minimal logging.

Layer 1: gateway / reverse proxy

  • TLS termination.
  • Auth.
  • Basic rate limiting.
  • Correct SSE streaming behavior.

Layer 2: policy + accounting service (thin control plane)

  • Token metering (TPM) and budgets.
  • Per-tenant concurrency caps.
  • Queueing/fairness.
  • Audit logs with redaction.

If you’re allergic to adding services: I get it. But the alternative is pushing policy into every client, which is worse. A self-hosted endpoint without a control plane becomes a free-for-all within weeks.

1) How do you put authentication in front of vLLM’s OpenAI-compatible API?

You have three realistic patterns.

Pattern A: API keys (fastest to ship)

  • Each tenant gets a key.
  • Gateway validates key and injects x-tenant-id header.
  • Downstream systems never see raw keys.

Pattern B: JWT / OIDC (best for internal multi-team)

  • Your IdP issues JWTs.
  • Gateway validates signature + aud.
  • Map claims (sub, org, entitlements) to tenant.

Pattern C: mTLS + service identity (best for service-to-service)

  • SPIFFE/SPIRE or mesh identity.
  • Usually paired with JWT for humans.

If you’re doing agent tooling, treat auth like you would for an MCP server. I wrote about the failure modes in MCP OAuth security and the safer setup patterns in How to secure MCP servers: auth + authZ.

Concrete rule I use:

  • Humans get OIDC.
  • Services get mTLS or workload identity.
  • Everything gets a tenant header derived at the edge.

Also: do not let the client send x-tenant-id. That’s how you end up with tenant data leakage. You derive tenant identity from auth, not from user-controlled headers. This is basic LLM security hygiene.

2) How do you enforce per-tenant quotas (requests/min, tokens/min, concurrency) for a self-hosted LLM?

You need quotas in three units, because LLM usage is not one-dimensional:

  • RPM (requests per minute): protects your HTTP layer.
  • Concurrent requests: protects GPU scheduling and memory.
  • TPM (tokens per minute): protects cost and fairness.

If you only do RPM, one tenant can send 1 request/minute with max_tokens=8192 and still wreck you.

If you only do concurrency, a tenant can run 1 long request forever.

If you only do TPM, a tenant can DoS your queue with lots of tiny requests.

My recommended enforcement split:

  • Gateway enforces RPM and burst limits.
  • Control plane enforces TPM, max tokens, max context, and concurrency reservations.
  • vLLM enforces hard caps where possible (max model length, max tokens per request). Treat this as last resort.

Token accounting specifics (what actually works):

  • Pre-charge an estimated token cost at admission time using a tokenizer.
  • Reconcile with actual usage after completion.
  • When streaming, update counters periodically (every N tokens) so you can cut off abusive streams.

This ties directly into LLM cost and agent per-task cost calculation. If you don’t meter tokens, you don’t control spend.

Noisy-neighbor prevention (the practical version):

  • Global max concurrency for the whole cluster (protects you).
  • Per-tenant concurrency (protects everyone else).
  • Optional priority tiers (prod > batch).

If you want a mental model, treat it like database connection pooling. The GPU is the database.

3) How should streaming (SSE) be implemented end-to-end to avoid buffering and broken UX?

Streaming is where “it works on my laptop” dies.

A good streaming UX has measurable properties:

  • TTFT p95 stays bounded under load.
  • Streams don’t buffer behind the proxy.
  • Client disconnects stop GPU work quickly.
  • Retries don’t double-bill tokens.

Checklist I use for SSE (proxy + server + client):

  1. Disable proxy buffering for SSE routes. If your gateway buffers, users see nothing for 20 seconds then a wall of text.
  2. Set read timeouts for long streams. Default 60s timeouts are a footgun.
  3. Send heartbeats. A comment line every 10–15 seconds prevents idle timeouts.
  4. Handle client disconnects. When the browser tab closes, stop generation.
  5. Backpressure matters. If the client can’t read fast enough, don’t queue infinite chunks.
  6. Retry semantics must be explicit. If a stream breaks at token 500, do you resume, restart, or fail? Pick one and document it.

You can’t hand-wave this. Streaming is part of your product.

If you’re building agents, streaming also changes tool-call behavior. A half-streamed response that triggers a tool call is how you get weird partial states. Tie this into your agent orchestration and agentic AI decisions.

4) What should you log (and redact) for prompts/responses in production?

Log everything and you will leak everything.

Log nothing and you won’t be able to debug outages, abuse, or customer issues.

The right answer is “log the minimum useful data, redact by default, and retain for a short window.”

My baseline schema:

  • Always log: timestamp, tenant id, request id, model name, prompt token count, completion token count, TTFT, total latency, status code, stream aborted (bool).
  • Sometimes log: prompt/response hashes (not raw), top-level classifier labels, policy decisions (allowed/denied + reason).
  • Rarely log: raw prompts/responses, only for opt-in tenants or sampled debug sessions with tight retention.

Redaction rules:

  • Treat prompts like user data. Because they are.
  • Redact secrets and identifiers before persistence.
  • Never log Authorization headers.

If you want deeper patterns here, read my LLM data leakage playbook and AI agent observability logging schema. They cover retention and correlation IDs in a way that’s actually implementable.

Abuse prevention note: prompt injection shows up in logs. That means logs become sensitive. Tie this into prompt injection and how to stop repo prompt injection in coding agents.

Deploy vLLM on Kubernetes with NVIDIA GPU (and the gotchas that waste days)

You can absolutely run vLLM on Kubernetes. Most production teams do, because the alternative is a pets-not-cattle GPU box that nobody wants to own.

The key is accepting a boring truth: Kubernetes GPU scheduling is a separate system you must get right before you debug vLLM.

A few concrete “this will bite you” details pulled from Vishnu Hari Dadhich’s Minikube/WSL2 guide:

  • Docker Desktop’s built-in Kubernetes on Windows/WSL2 can fail to expose GPUs to pods even when nvidia-smi works on the host.
  • Using Minikube plus the NVIDIA device plugin/toolkit configuration is a reliable workaround.
  • The /dev/shm shared memory gotcha is real. Insufficient shared memory breaks inference in ways that look like random instability.

If you’re running consumer GPUs, also cross-check my ROCm coverage if you’re on AMD, and the broader local AI decisions if you’re still choosing hardware.

One number to keep in your head: Jangwook Kim references an architecture footprint serving across 24 NVIDIA B300 GPUs. You don’t need 24 GPUs to start. But it’s a good reminder that “SLA-grade” means routing, replicas, warmups, and measurement. Not one pod.

Troubleshooting vLLM GPU Support on Docker Desktop WSL2 Kubernetes

This is the fastest way to lose a day: everything works in docker run --gpus all, then your pod shows zero GPUs.

Common failure modes:

  • NVIDIA device plugin not installed or misconfigured.
  • Container runtime mismatch (containerd config not using NVIDIA runtime).
  • Docker Desktop Kubernetes limitations on Windows/WSL2.

Vishnu’s takeaway is pragmatic: Docker Desktop Kubernetes may simply not detect your GPU in WSL2 scenarios, and Minikube with the right setup is more reliable.

Also, don’t ignore shared memory.

If you see weird crashes under load, check /dev/shm sizing before you start blaming “vLLM bugs.”

The 15-item production checklist (auth, quotas, streaming, safety)

Here’s the actual checklist I’d use to declare “this endpoint is production-shaped.” It’s intentionally blunt.

  1. TLS termination at the edge. No plaintext in-cluster unless you have a reason.
  2. Auth in front of vLLM (Authorization: Bearer …). Never expose vLLM directly.
  3. Tenant identity derived from auth, injected as x-tenant-id.
  4. RPM and burst rate limits at the gateway.
  5. Per-tenant concurrency caps (hard).
  6. Per-tenant TPM budgets with token accounting.
  7. Request guardrails: max input tokens, max output tokens, max context.
  8. Global admission control when GPU queue time exceeds SLO.
  9. Streaming SSE configured to not buffer at the proxy.
  10. Client disconnect stops generation.
  11. Structured logs with correlation IDs and tenant ids.
  12. Redaction-first logging. No raw prompts by default.
  13. SLIs: TTFT p95, total latency p95, tokens/sec, 429 rate, queue wait time, aborted streams.
  14. Rollout plan: canary model version, warmup, and rollback within minutes.
  15. Showback: per-tenant token usage reporting, weekly.

If you only do 5 things from this list, do: 2, 4, 5, 9, 13.

A practical rollout strategy that won’t break OpenAI SDK clients

The easiest way to burn trust is breaking client integrations because you swapped a model or changed behavior.

Rules for safe rollouts:

  • Pin compatibility: keep OpenAI-compatible endpoints stable. Add new behavior behind headers or new model IDs.
  • Canary by tenant: route 1 tenant (or 1%) to the new model version first.
  • Warm up: pre-load weights before shifting traffic. Cold loads create fake incidents.
  • Version your policy: quota rules, moderation rules, and default max tokens should be versioned. Otherwise you can’t explain behavior changes.

This is the same kind of discipline I’ve learned running this site’s multi-agent publishing pipeline: deterministic gates catch issues before they become user-visible incidents. In my pipeline, deterministic checks have been more effective than simply “use a bigger model to review the output.” The boring gate wins.

If your clients are agents, add regression tests. Use the same mindset as AI engineering evals: regression gates or agent tool call failure testing.

Here’s the prediction: by mid-2027, self-hosted inference endpoints that don’t ship with token budgets, tenant isolation, and streaming metrics will look as irresponsible as unauthenticated Redis on the public internet. If you’re self-hosting vLLM now, you can get ahead of that expectation. Or you can wait until an internal team turns your GPU into a shared tragedy of the commons.

Here’s the challenge. Pick one thing you’re currently hand-waving. Auth, quotas, streaming, logging, rollouts. Then implement it end-to-end this week. Production doesn’t reward vibes.

Photo by imgix on Unsplash.

Continue reading

inference server rackmount gpu machine — illustration for article on How to Serve a Local LLM

How to Serve a Local LLM to Multiple Users [2026]

If your local LLM server “works” but falls apart at 30–50 concurrent chats, this guide shows the real limiter (KV cache) and the exact knobs in vLLM, SGLang, and TGI that move P99.

vLLM vs Ollama 2026: Production Power or Developer Ease?

vLLM vs Ollama 2026: Production Power or Developer Ease?

vLLM wins for high-throughput production deployments where every token/second counts; Ollama wins for local developer workflows where setup speed and portability matter most. Pick wrong and you'll either over-engineer a side project or under-power a real API.

gpu server rack datacenter nvidia — illustration for article on Mercury 2.5 770 tok/s Benchmark: The

Mercury 2.5 770 tok/s Benchmark: The Production Playbook [2026]

Mercury 2.5 is clocking ~780 tok/s on Artificial Analysis. At that speed, decode stops being your bottleneck. Here’s what actually changes in production: batching, streaming UX, context packing, cost controls, and evals that don’t lie.

A computer monitor sitting on top of a desk

How to Run Qwen 35B on 16GB VRAM [2026]: Flags + Quants

A reproducible 16GB recipe for “Qwen 35B”: which Qwen2.5-32B quants fit, how to budget KV cache, and the exact serving flags that stop OOMs.

Cite this article
Kunal Ganglani (2026, September 28). vLLM Self-Hosted LLM Production Checklist [2026]: Auth + Quotas. Kunal Ganglani. Retrieved September 28, 2026, from https://www.kunalganglani.com/blog/vllm-production-checklist