Skip to content
KG.
    • Projects
    • Services
    • Blog
    • Learning Paths
    • Glossary
    • Cheatsheets
    • Topic Pillars
    • Tools
    • Games
    • Demos
    • Challenges
    • About
    • Skills
    • Travel
    • Uses
    • Bookshelf
    • Reading List
  • Let's talk
  1. Home
  2. ›
  3. Blog
  4. ›
  5. #vram

#vram

2 posts tagged with #vram

Every article below is hand-written, technically reviewed, and focused on vram. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

selective focus photography of GEFORCE RTX graphics card AI and Machine Learning

LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]

The practitioner's guide to choosing between Q4_K_M, Q5_K_S, Q8_0, and FP16 quantization for local LLMs — with real perplexity numbers, throughput benchmarks, and per-use-case recommendations.

July 6, 2026 13 min read
Read more
local-llm, ollama, lm-studio, mlx, gpu, vram, apple-silicon, nvidia, ai-hardware, self-hosted-ai Technology

Local LLM Hardware Guide 2026: VRAM, GPUs, and Setup [Tested]

Every VRAM tier, GPU option, and runtime tool mapped out for running local LLMs in 2026 — from budget CPU-only rigs to RTX 5090 and Apple Silicon M5 workstations.

March 1, 2026 16 min read
Read more
© 2026 Kunal Ganglani. Built with coffee and curiosity in Toronto.