SKILL.md packages that extend Claude Code, Cursor, Copilot, and other AI agents.
Tags

auto-improve
An autonomous, judge-gated loop that iteratively improves text artifacts using a rubric and a separate evaluator model to ensure verified gains.

alchaincyf
Autonomous skill optimizer that uses a 9-dimension rubric to evaluate and improve SKILL.md files through an iterative, validation-gated loop.

entityprocess
Run AgentV evaluations and optimize agents through eval-driven iteration.

aeon-bench-pod
Execute end-to-end LLM benchmarks using AEON Bench Pod. From model verification to leaderboard submission.

pjt222
Evaluates pre-trained diffusion models using quality metrics, noise schedule analysis, and latent space probing.

claude-skill-registry
Comprehensive guide for integrating LangSmith to trace, evaluate, and monitor LLM applications and AI agents in production.

de-anthropocentric-research-engine
Multi-stage evaluation pipeline that scores ideas for novelty, assesses feasibility, ranks them, and selects top candidates for further development.

nemotron
Authoritative knowledge-base for Nemotron 3 Nano: architecture, training data, SFT/RL recipes, evaluation and deployment guidance.