#evals

9 posts tagged with #evals

Every article below is hand-written, technically reviewed, and focused on evals. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

a man using a laptop computer on a table AI and Machine Learning

How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]

A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.

MLflow tracking UI experiment dashboard screen — illustration for article on How to Use MLflow LLM AI and Machine Learning

How to Use MLflow LLM Evaluation Tracing [2026] (Spans → Gates)

Turn MLflow traces into living eval datasets, rubric scorecards, and CI/CD regression gates for tool-calling agents. Stop debugging with random logs.

A digital dashboard displaying marketing metrics including CTR and quality score on a screen AI and Machine Learning

How to Do Non Deterministic AI System Testing [2026]

A release-gating playbook for LLMs and agents: golden sets, metamorphic tests, variance control, tolerance bands, cohort diffs, and eval-regression postmortems.

a computer screen with a bunch of data on it AI and Machine Learning

RAG Evaluation Metrics for Retrieval Quality: My Production Playbook

If your RAG app got worse after an embeddings or chunking change, grading answers won’t tell you why. Here’s a retrieval-first eval workflow: recall@k, MRR, citation accuracy, leakage tests, and a frozen offline corpus so regressions are real—not vibes.

a clipboard with a checklist on it next to a cup of coffee and AI and Machine Learning

Agent Evaluation Roadmap for Small Teams [2026]: The 30-Min/Week Plan

A pragmatic map of offline vs online vs HITL evals, what to measure beyond “response quality,” and the smallest program that actually prevents agent regressions.

a person typing on a laptop keyboard on a desk AI and Machine Learning

AI Engineering Evals: Regression Gates for Prompts, Tools, RAG [2026]

Stop letting prompt tweaks and model upgrades silently break production. Here’s a CI-style regression gate system for prompts, tool calling, and RAG with golden sets, schemas, shadow evals, and failure budgets.

a computer screen with a bar chart on it AI and Machine Learning

AI Agent Evaluation Framework 2026: 8 Metrics Beyond Task Success

If your agent eval is just “did it finish the task?”, you’re flying blind. Here’s a 2026-ready scorecard for tool correctness, recovery, safety, and cost-per-success—plus a regression suite blueprint you can actually run in CI.

black hp laptop computer turned on displaying desktop AI and Machine Learning

Agent Evaluation Harness [2026]: Replay, Rubrics, CI Gates

Most agent failures aren’t “bad prompts”. They’re multi-step tool cascades. Here’s how I build an agent evaluation harness that actually prevents regressions.

developer monitoring dashboard laptop screen metrics — illustration for article on Evaluate AI Agents in Production: AI and Machine Learning

Evaluate AI Agents in Production: 3-Level Framework [2026]

Most AI agent failures trace back to missing evals. Here's the 3-level framework — unit tests, LLM-as-judge, and online evaluation — that actually works in production.