Skip to content
KG.
  • Hire me, or see what I ship.

    • ProjectsCase studies & shipped work
    • ServicesWork with me
    • SponsorSponsor the blog
  • Essays and references on AI engineering.

    • BlogEssays on AI engineering
    • Learning PathsGuided curricula
    • Topic PillarsDeep-dive hubs
    • GlossaryAI & dev terms, defined
    • CheatsheetsQuick references
    • ComparisonsX vs Y, decided
  • Things to click, play, and take apart.

    • Tools28 dev & AI utilities
    • Games16 browser games
    • DemosInteractive explainers
    • ChallengesDaily coding puzzles
  • Me

    • AboutWho I am
    • UsesMy gear & setup
    • ResumeCV (PDF)

    Shelf

    • BookshelfBooks I recommend
    • Reading ListWhat I'm reading
  • Let's talk
  1. Home
  2. ›
  3. Blog
  4. ›
  5. #llm-serving

#llm-serving

3 posts tagged with #llm-serving

Every article below is hand-written, technically reviewed, and focused on llm-serving. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

Rows of black server racks with white logos in a data center Cloud and DevOps

vLLM Self-Hosted LLM Production Checklist [2026]: Auth + Quotas

A production-first checklist for self-hosting an OpenAI-compatible vLLM endpoint: auth, per-tenant quotas, streaming SSE, queueing, and redaction-safe logs.

September 28, 2026 9 min read
Read more
inference server rackmount gpu machine — illustration for article on How to Serve a Local LLM Cloud and DevOps

How to Serve a Local LLM to Multiple Users [2026]

If your local LLM server “works” but falls apart at 30–50 concurrent chats, this guide shows the real limiter (KV cache) and the exact knobs in vLLM, SGLang, and TGI that move P99.

September 16, 2026 9 min read
Read more
vLLM vs Ollama 2026: Production Power or Developer Ease? Developer Tools

vLLM vs Ollama 2026: Production Power or Developer Ease?

vLLM wins for high-throughput production deployments where every token/second counts; Ollama wins for local developer workflows where setup speed and portability matter most. Pick wrong and you'll either over-engineer a side project or under-power a real API.

May 10, 2026 12 min read
Read more

Content

  • Blog
  • Topic Pillars
  • Learning Paths
  • Games
  • Demos

Resources

  • Tools
  • Glossary
  • Cheatsheets
  • Comparisons

About Me

  • About
  • Uses
  • Bookshelf
  • Reading List

Meta

  • Subscribe
  • Changelog
  • Sitemap
  • Privacy
  • Terms
  • RSS
KG

Building intelligent systems. Still chasing those sour icecreams.

Made with coffee and curiosity in Toronto. 2026.