#testing
2 posts
-
Evals vs Benchmarks vs Certification: What Each One Actually Proves
A mechanism-level explainer of what an eval, a benchmark, and per-decision certification each prove about an AI agent or a model change, why an aggregate pass rate is not per-decision safety, and where the hard, largely-unsolved part still lives.
-
What is agent evaluation?
Agent evaluation: measuring multi-step trajectories, tool use, and open-ended outputs. Why benchmarks alone don't tell you whether an agent works in production.