See which skills help and where your agent falls short.
agent evaluation
Grades an agent setup from local conversations, proposes evidence-based skill edits, and produces a shareable report.
When to use it
Use when the user wants their agent setup graded from conversation history or wants to know which installed skills work.
Choose the conversation and skill scopes; it grades the sessions and saves report artifacts in a fresh scratch directory.
What you provide
This skill
$REPORT_DIR/transcripts/
Copies transcripts to scratch storage
Warp conversation databases
Reads local Warp conversations
Claude Code project-history JSONL
Reads Claude Code history
Codex rollout JSONL
Reads local Codex rollouts
Python 3 runs the session collector and report renderer.