#llm-serving
3 posts tagged with #llm-serving
Every article below is hand-written, technically reviewed, and focused on llm-serving. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.
Cloud and DevOps vLLM Self-Hosted LLM Production Checklist [2026]: Auth + Quotas
A production-first checklist for self-hosting an OpenAI-compatible vLLM endpoint: auth, per-tenant quotas, streaming SSE, queueing, and redaction-safe logs.
Cloud and DevOps How to Serve a Local LLM to Multiple Users [2026]
If your local LLM server “works” but falls apart at 30–50 concurrent chats, this guide shows the real limiter (KV cache) and the exact knobs in vLLM, SGLang, and TGI that move P99.
Developer Tools vLLM vs Ollama 2026: Production Power or Developer Ease?
vLLM wins for high-throughput production deployments where every token/second counts; Ollama wins for local developer workflows where setup speed and portability matter most. Pick wrong and you'll either over-engineer a side project or under-power a real API.