Skip to content
KG.
  • Hire me, or see what I ship.

    • ProjectsCase studies & shipped work
    • ServicesWork with me
    • SponsorSponsor the blog
  • Essays and references on AI engineering.

    • BlogEssays on AI engineering
    • Learning PathsGuided curricula
    • Topic PillarsDeep-dive hubs
    • GlossaryAI & dev terms, defined
    • CheatsheetsQuick references
    • ComparisonsX vs Y, decided
  • Things to click, play, and take apart.

    • Tools28 dev & AI utilities
    • Games13 browser games
    • DemosInteractive explainers
    • ChallengesDaily coding puzzles
  • Me

    • AboutWho I am
    • UsesMy gear & setup
    • ResumeCV (PDF)

    Shelf

    • BookshelfBooks I recommend
    • Reading ListWhat I'm reading
  • Let's talk
  1. Home
  2. ›
  3. Blog
  4. ›
  5. #gemma-4

#gemma-4

3 posts tagged with #gemma-4

Every article below is hand-written, technically reviewed, and focused on gemma-4. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

a close up of a cpu chip on a table AI and Machine Learning

Gemma 4 26B CPU Inference Benchmark: 5 tok/s Production Math [2026]

A $300 Xeon from 2013 runs Gemma 4 26B at 5 tok/s with no GPU. Here's the memory bandwidth math, quantization tradeoffs, and production decision framework nobody else is covering.

July 16, 2026 13 min read
Read more
gray and black laptop computer on surface AI and Machine Learning

How to Run Local Agentic AI on Your Mac With MLX After WWDC 2026

Apple's WWDC 2026 MLX session was 13 minutes and skipped the hard parts. Here's the full setup: model selection, MTP speculative decoding, multimodal support, and wiring it all to a coding agent.

June 13, 2026 13 min read
Read more
a black and white photo of an abstract object AI and Machine Learning

Gemma 4 12B vs GPT-4o Mini vs Claude Haiku: Is Google's Local LLM Good Enough to Replace API Calls? [2026]

I ran Gemma's 12B model locally via Ollama and compared it against GPT-4o Mini and Claude Haiku on real dev tasks — here's when the free local model actually beats paid APIs.

June 4, 2026 7 min read
Read more

Content

  • Blog
  • Topic Pillars
  • Learning Paths
  • Games
  • Demos

Resources

  • Tools
  • Glossary
  • Cheatsheets
  • Comparisons

About Me

  • About
  • Uses
  • Bookshelf
  • Reading List

Meta

  • Subscribe
  • Changelog
  • Sitemap
  • Privacy
  • Terms
  • RSS
KG

Building intelligent systems. Still chasing those sour icecreams.

Made with coffee and curiosity in Toronto. 2026.