Build, test, and improve reliable agent skills.
agent skill development
Creates, tests, benchmarks, and iteratively improves agent skills and their triggering descriptions.
When to use it
Use when creating or improving a skill, evaluating its performance, or optimizing its triggering accuracy.
Tell it what skill you want; it interviews you, drafts and tests the skill, then iterates with your feedback.
This skill
Claude
Sends evals to Claude
web search results
Searches docs and similar skills
eval viewer server
Starts a local review server
/tmp/eval_review_<skill-name>.html
Writes a temporary review page
Python is required to run the benchmark aggregation and review-viewer commands in the evaluation workflow.
Description optimization requires an API key loaded from environment variables, possibly by sourcing a .env file.