
from eval-layer12
Add a rubric-based evaluation layer to agent projects to measure quality via LLM-as-a-judge scoring and weighted metrics.
This skill enables an agent to implement a rigorous, rubric-based evaluation system for any AI agent project. It transforms qualitative 'vibes' into quantitative metrics by establishing a scoring framework and a programmatic harness to execute it.
Use this skill when you need to:
main.yaml rubrics and seed.yaml test cases.Designed primarily for Claude Code, but framework-agnostic and compatible with any agent capable of writing Python and managing a local workspace (e.g., Cursor, Codex).
This skill has not been reviewed by our automated audit pipeline yet.