#testing
11 posts tagged with #testing
Every article below is hand-written, technically reviewed, and focused on testing. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.
AI and Machine Learning How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]
A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.
Developer Tools AI Generated Code Quality [2026]: A 0–100 Audit Rubric + CI Gates
AI-generated code quality isn’t a style problem anymore. Here’s a practical 0–100 audit rubric (tests, security, maintainability, rewrite risk) plus CI gates that block hallucinated APIs before they ship.
AI and Machine Learning How to Do Non Deterministic AI System Testing [2026]
A release-gating playbook for LLMs and agents: golden sets, metamorphic tests, variance control, tolerance bands, cohort diffs, and eval-regression postmortems.
Developer Tools Review AI-Generated Code Checklist [2026]: Beat the Bottleneck
AI makes code cheap. Verification is the new tax. Here’s a practical workflow, checklist, and ‘delete & redo’ rubric to keep PRs fast and safe in 2026.
AI and Machine Learning Agent Evaluation Roadmap for Small Teams [2026]: The 30-Min/Week Plan
A pragmatic map of offline vs online vs HITL evals, what to measure beyond “response quality,” and the smallest program that actually prevents agent regressions.
AI and Machine Learning How to Start an AI Agent Evaluation Program (5-Task Scorecard)
Ship agents with a regression safety net: pick 5 real tasks, define pass/fail, run evals weekly in CI, and publish a scorecard that forces better decisions.
AI and Machine Learning How to Do Agent Tool Call Failure Testing [2026 CI Harness]
Build a deterministic CI harness that breaks your agent’s tools on purpose: timeouts, 429s, partial writes, retries, and stuck loops. Then regress it with golden traces.
Cybersecurity How to Do Prompt Injection Regression Testing [2026 CI]
Your prompt-injection evals are probably overfit to cute synthetic prompts. Here’s a CI-ready regression suite that still catches indirect injection: seed corpora, adversarial transforms, canary secrets, and hard fail gates.
AI and Machine Learning AI Engineering Evals: Regression Gates for Prompts, Tools, RAG [2026]
Stop letting prompt tweaks and model upgrades silently break production. Here’s a CI-style regression gate system for prompts, tool calling, and RAG with golden sets, schemas, shadow evals, and failure budgets.
AI and Machine Learning Agent Evaluation Harness [2026]: Replay, Rubrics, CI Gates
Most agent failures aren’t “bad prompts”. They’re multi-step tool cascades. Here’s how I build an agent evaluation harness that actually prevents regressions.
AI and Machine Learning Evaluate AI Agents in Production: 3-Level Framework [2026]
Most AI agent failures trace back to missing evals. Here's the 3-level framework — unit tests, LLM-as-judge, and online evaluation — that actually works in production.