Turn AI failures into rerunnable evals and targeted fixes.
LLM evaluation
Builds a rerunnable evaluation suite for an AI application and improves the application from evaluation failures.
When to use it
Use when evaluating or improving an AI agent, chatbot, RAG pipeline, or other LLM application with DeepEval.
Work through intake choices with it; it creates and runs a project eval suite, then iterates on failures.
What you provide
This skill
Confident AI
Saves results to Confident AI
Requires Python 3.9 or newer.
The target project must have the deepeval package installed.
Model credentials are required for evaluation metrics or synthetic dataset generation.
A Confident AI login is required only for reporting, hosted traces, and online evaluations.