
from aeon-bench-pod21
Execute end-to-end LLM benchmarks using AEON Bench Pod. From model verification to leaderboard submission.
This skill provides the operational blueprint for running verified LLM evaluations using the AEON Bench Pod. It allows an agent to deploy a local benchmarking appliance, verify model weights against Hugging Face, execute a comprehensive suite (text, agentic, vision, audio, video, and performance), and submit attested results to the public leaderboard.
Use this skill when asked to benchmark, score, or evaluate an LLM's performance. It is the primary path for generating 'attested' results that rank on the AEON leaderboard, moving beyond simple smoke tests to a full-scale verified evaluation.
aeon_pod_mcp.py for headless agent control.Any AI agent capable of executing shell commands (Docker) and utilizing MCP tools.
This skill has not been reviewed by our automated audit pipeline yet.