#testing

11 posts tagged with #testing

Every article below is hand-written, technically reviewed, and focused on testing. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

a man using a laptop computer on a table AI and Machine Learning

How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]

A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.

Laptop screen displaying lines of code Developer Tools

AI Generated Code Quality [2026]: A 0–100 Audit Rubric + CI Gates

AI-generated code quality isn’t a style problem anymore. Here’s a practical 0–100 audit rubric (tests, security, maintainability, rewrite risk) plus CI gates that block hallucinated APIs before they ship.

A digital dashboard displaying marketing metrics including CTR and quality score on a screen AI and Machine Learning

How to Do Non Deterministic AI System Testing [2026]

A release-gating playbook for LLMs and agents: golden sets, metamorphic tests, variance control, tolerance bands, cohort diffs, and eval-regression postmortems.

Laptop screen displaying lines of code Developer Tools

Review AI-Generated Code Checklist [2026]: Beat the Bottleneck

AI makes code cheap. Verification is the new tax. Here’s a practical workflow, checklist, and ‘delete & redo’ rubric to keep PRs fast and safe in 2026.

a clipboard with a checklist on it next to a cup of coffee and AI and Machine Learning

Agent Evaluation Roadmap for Small Teams [2026]: The 30-Min/Week Plan

A pragmatic map of offline vs online vs HITL evals, what to measure beyond “response quality,” and the smallest program that actually prevents agent regressions.

qa checklist clipboard testing — illustration for article on How to Start an AI Agent Evaluation AI and Machine Learning

How to Start an AI Agent Evaluation Program (5-Task Scorecard)

Ship agents with a regression safety net: pick 5 real tasks, define pass/fail, run evals weekly in CI, and publish a scorecard that forces better decisions.

a computer screen with a blue background AI and Machine Learning

How to Do Agent Tool Call Failure Testing [2026 CI Harness]

Build a deterministic CI harness that breaks your agent’s tools on purpose: timeouts, 429s, partial writes, retries, and stuck loops. Then regress it with golden traces.

GitHub Actions workflow run laptop screen — illustration for article on How to Do Prompt Injection Cybersecurity

How to Do Prompt Injection Regression Testing [2026 CI]

Your prompt-injection evals are probably overfit to cute synthetic prompts. Here’s a CI-ready regression suite that still catches indirect injection: seed corpora, adversarial transforms, canary secrets, and hard fail gates.

a person typing on a laptop keyboard on a desk AI and Machine Learning

AI Engineering Evals: Regression Gates for Prompts, Tools, RAG [2026]

Stop letting prompt tweaks and model upgrades silently break production. Here’s a CI-style regression gate system for prompts, tool calling, and RAG with golden sets, schemas, shadow evals, and failure budgets.

black hp laptop computer turned on displaying desktop AI and Machine Learning

Agent Evaluation Harness [2026]: Replay, Rubrics, CI Gates

Most agent failures aren’t “bad prompts”. They’re multi-step tool cascades. Here’s how I build an agent evaluation harness that actually prevents regressions.

developer monitoring dashboard laptop screen metrics — illustration for article on Evaluate AI Agents in Production: AI and Machine Learning

Evaluate AI Agents in Production: 3-Level Framework [2026]

Most AI agent failures trace back to missing evals. Here's the 3-level framework — unit tests, LLM-as-judge, and online evaluation — that actually works in production.