#vllm
4 posts tagged with #vllm
Every article below is hand-written, technically reviewed, and focused on vllm. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.
Cloud and DevOps vLLM Self-Hosted LLM Production Checklist [2026]: Auth + Quotas
A production-first checklist for self-hosting an OpenAI-compatible vLLM endpoint: auth, per-tenant quotas, streaming SSE, queueing, and redaction-safe logs.
AI and Machine Learning How to Run Qwen 35B on 16GB VRAM [2026]: Flags + Quants
A reproducible 16GB recipe for “Qwen 35B”: which Qwen2.5-32B quants fit, how to budget KV cache, and the exact serving flags that stop OOMs.
Cloud and DevOps How to Serve a Local LLM to Multiple Users [2026]
If your local LLM server “works” but falls apart at 30–50 concurrent chats, this guide shows the real limiter (KV cache) and the exact knobs in vLLM, SGLang, and TGI that move P99.
Developer Tools vLLM vs Ollama 2026: Production Power or Developer Ease?
vLLM wins for high-throughput production deployments where every token/second counts; Ollama wins for local developer workflows where setup speed and portability matter most. Pick wrong and you'll either over-engineer a side project or under-power a real API.