#evals
9 posts tagged with #evals
Every article below is hand-written, technically reviewed, and focused on evals. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.
AI and Machine Learning How to Do Synthetic Data for RAG Evaluation [2026 Pipeline]
A leakage-first, adversarial-negative pipeline to generate, label, and validate a synthetic RAG evaluation set without building a benchmark that only tests your generator.
AI and Machine Learning How to Use MLflow LLM Evaluation Tracing [2026] (Spans → Gates)
Turn MLflow traces into living eval datasets, rubric scorecards, and CI/CD regression gates for tool-calling agents. Stop debugging with random logs.
AI and Machine Learning How to Do Non Deterministic AI System Testing [2026]
A release-gating playbook for LLMs and agents: golden sets, metamorphic tests, variance control, tolerance bands, cohort diffs, and eval-regression postmortems.
AI and Machine Learning RAG Evaluation Metrics for Retrieval Quality: My Production Playbook
If your RAG app got worse after an embeddings or chunking change, grading answers won’t tell you why. Here’s a retrieval-first eval workflow: recall@k, MRR, citation accuracy, leakage tests, and a frozen offline corpus so regressions are real—not vibes.
AI and Machine Learning Agent Evaluation Roadmap for Small Teams [2026]: The 30-Min/Week Plan
A pragmatic map of offline vs online vs HITL evals, what to measure beyond “response quality,” and the smallest program that actually prevents agent regressions.
AI and Machine Learning AI Engineering Evals: Regression Gates for Prompts, Tools, RAG [2026]
Stop letting prompt tweaks and model upgrades silently break production. Here’s a CI-style regression gate system for prompts, tool calling, and RAG with golden sets, schemas, shadow evals, and failure budgets.
AI and Machine Learning AI Agent Evaluation Framework 2026: 8 Metrics Beyond Task Success
If your agent eval is just “did it finish the task?”, you’re flying blind. Here’s a 2026-ready scorecard for tool correctness, recovery, safety, and cost-per-success—plus a regression suite blueprint you can actually run in CI.
AI and Machine Learning Agent Evaluation Harness [2026]: Replay, Rubrics, CI Gates
Most agent failures aren’t “bad prompts”. They’re multi-step tool cascades. Here’s how I build an agent evaluation harness that actually prevents regressions.
AI and Machine Learning Evaluate AI Agents in Production: 3-Level Framework [2026]
Most AI agent failures trace back to missing evals. Here's the 3-level framework — unit tests, LLM-as-judge, and online evaluation — that actually works in production.