Skip to content
KG.
  • Hire me, or see what I ship.

    • ProjectsCase studies & shipped work
    • ServicesWork with me
    • SponsorSponsor the blog
  • Essays and references on AI engineering.

    • BlogEssays on AI engineering
    • Learning PathsGuided curricula
    • Topic PillarsDeep-dive hubs
    • GlossaryAI & dev terms, defined
    • CheatsheetsQuick references
    • ComparisonsX vs Y, decided
  • Things to click, play, and take apart.

    • Tools28 dev & AI utilities
    • Games13 browser games
    • DemosInteractive explainers
    • ChallengesDaily coding puzzles
  • Me

    • AboutWho I am
    • UsesMy gear & setup
    • ResumeCV (PDF)

    Shelf

    • BookshelfBooks I recommend
    • Reading ListWhat I'm reading
  • Let's talk
  1. Home
  2. ›
  3. Blog
  4. ›
  5. #latency

#latency

4 posts tagged with #latency

Every article below is hand-written, technically reviewed, and focused on latency. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

White xbox controller and a mechanical keyboard on wood Technology

Xbox Cloud Gaming Pay‑As‑You‑Go Latency: Fix Input Lag Fast [2026]

If Xbox Cloud Gaming goes pay‑as‑you‑go, latency becomes a tax. Here’s the measurement-first workflow and home network settings that actually cut input lag.

September 4, 2026 7 min read
Read more
stock market chart displayed on laptop screen AI and Machine Learning

LLM Latency Benchmark Methodology: Streaming UX Metrics [2026]

A UX-first LLM latency benchmark methodology for streaming chat and agent apps: measure chunk cadence, jitter, tool-call stall time, and end-to-end time-to-usable—not just TTFT.

August 9, 2026 12 min read
Read more
Laptop screen displaying code with colorful lighting. Developer Tools

Rust Allocator: jemalloc vs mimalloc vs tcmalloc for P99 [2026]

Allocator switching can cut P99 latency in Rust services. It can also do absolutely nothing. Here’s how to benchmark it like an adult and tune jemalloc without cargo-culting.

July 20, 2026 10 min read
Read more
llm-api, ai-benchmarks, latency, gpt-4-1, claude-haiku, gemini-flash, time-to-first-token, ai-performance, production-ai, developer-tools Technology

5 LLM APIs Tested for Latency: Real Data [2026]

I benchmarked Claude Haiku 4.5, Claude Sonnet 4, GPT-4.1, GPT-4.1 Mini, and Gemini 2.5 Flash for TTFT, throughput, and end-to-end latency — with a cost-latency decision matrix for production builders.

March 7, 2026 10 min read
Read more

Content

  • Blog
  • Topic Pillars
  • Learning Paths
  • Games
  • Demos

Resources

  • Tools
  • Glossary
  • Cheatsheets
  • Comparisons

About Me

  • About
  • Uses
  • Bookshelf
  • Reading List

Meta

  • Subscribe
  • Changelog
  • Sitemap
  • Privacy
  • Terms
  • RSS
KG

Building intelligent systems. Still chasing those sour icecreams.

Made with coffee and curiosity in Toronto. 2026.