Skip to content
KG.
  • Hire me, or see what I ship.

    • ProjectsCase studies & shipped work
    • ServicesWork with me
    • SponsorSponsor the blog
  • Essays and references on AI engineering.

    • BlogEssays on AI engineering
    • Learning PathsGuided curricula
    • Topic PillarsDeep-dive hubs
    • GlossaryAI & dev terms, defined
    • CheatsheetsQuick references
    • ComparisonsX vs Y, decided
  • Things to click, play, and take apart.

    • Tools28 dev & AI utilities
    • Games13 browser games
    • DemosInteractive explainers
    • ChallengesDaily coding puzzles
  • Me

    • AboutWho I am
    • UsesMy gear & setup
    • ResumeCV (PDF)

    Shelf

    • BookshelfBooks I recommend
    • Reading ListWhat I'm reading
  • Let's talk
  1. Home
  2. ›
  3. Blog
  4. ›
  5. #evaluation

#evaluation

3 posts tagged with #evaluation

Every article below is hand-written, technically reviewed, and focused on evaluation. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

Person wearing headphones using laptop with audio software Cybersecurity

Why AI Voice Detectors Fail [2026]: Codecs, Watermarks, Traps

“AI voice detector” scores swing wildly because the audio channel is the adversary. Here’s how codecs, noise suppression, and bad metrics break detection in the 2026 real world.

September 10, 2026 10 min read
Read more
Abstract grayscale waveform with dark background Cybersecurity

How to Run an AI Voice Detector Accuracy Test [2026 Harness]

Build a repeatable deepfake-voice detector benchmark: datasets, metrics, thresholding by false-positive cost, robustness transforms, and a privacy-safe way to publish results.

August 31, 2026 9 min read
Read more
qa checklist clipboard testing — illustration for article on How to Start an AI Agent Evaluation AI and Machine Learning

How to Start an AI Agent Evaluation Program (5-Task Scorecard)

Ship agents with a regression safety net: pick 5 real tasks, define pass/fail, run evals weekly in CI, and publish a scorecard that forces better decisions.

August 24, 2026 9 min read
Read more

Content

  • Blog
  • Topic Pillars
  • Learning Paths
  • Games
  • Demos

Resources

  • Tools
  • Glossary
  • Cheatsheets
  • Comparisons

About Me

  • About
  • Uses
  • Bookshelf
  • Reading List

Meta

  • Subscribe
  • Changelog
  • Sitemap
  • Privacy
  • Terms
  • RSS
KG

Building intelligent systems. Still chasing those sour icecreams.

Made with coffee and curiosity in Toronto. 2026.