Build and run evaluators across traces and experiments.
LLM evaluation
Designs, creates, and runs LLM-as-judge and deterministic code evaluators on Arize.
When to use it
Use when creating or improving evaluators, scoring spans or experiments, configuring mappings, or setting up continuous monitoring.
Ask it to manage or run an evaluator; it gathers needed choices, then creates evaluators, tasks, or evaluation runs in Arize.
This skill
Arize account
Creates evaluators and tasks
Arize evaluation runs
Starts Arize evaluation runs
Arize profile
Updates your Arize profile
Requires the ax CLI.
Requires a configured Arize profile with access to the intended space.
An AI Integration with stored provider credentials is required for LLM-as-judge evaluators but not code evaluators.
Delegates AI integration creation to this skill when no platform-managed provider credentials exist.