
from entityprocess
Run AgentV evaluations and optimize agents through eval-driven iteration.
A comprehensive skill for evaluating AI agents and iteratively improving them through data-driven optimization. It provides a structured loop: understand agent behavior, write evaluation test cases (EVAL.yaml/evals.json), run and grade outputs, analyze results for patterns, and apply improvements.
Key Features:
This skill has not been reviewed by our automated audit pipeline yet.